Security Notes
System Design

Design a Metrics Monitoring System

5 min read 7 sections

DifficultyHard | HelloInterview: problem breakdown


Problem Statement

Design a metrics monitoring and alerting system like Datadog, Prometheus, or CloudWatch. Collect time-series metrics from thousands of hosts, store them efficiently, visualize trends, and alert on anomalies.

πŸ“–

Real-world: The make-or-break insight is that time-series data is special enough to deserve its own database. You're writing millions of mostly-incrementing numbers tagged with labels, almost never updating, and querying by time range β€” so a purpose-built TSDB (Prometheus, InfluxDB, Facebook's Gorilla, the basis of much of this field) crushes a general database. Gorilla's famous trick: delta-of-delta timestamp encoding + XOR float compression shrank data ~12Γ— and kept it in memory for fast alerting. Two themes this design hinges on: downsampling/rollups (keep raw data for hours, then aggregate to coarser resolution for long-term retention β€” you don't store per-second data for a year), and cardinality explosion β€” the classic footgun where a high-cardinality label like user_id or request_id multiplies your series into the millions and melts the system. Security relevance: this is the same architecture a SIEM/detection pipeline uses β€” high-ingest, time-windowed, alert-on-threshold β€” so the monitoring design and the detection design are cousins, and "alert fatigue" / hysteresis tuning applies to both.


Requirements

Functional

  • Collect metrics from services (CPU, memory, request rate, error rate, custom metrics)
  • Store metrics with high write throughput
  • Query: avg(cpu_usage) by host WHERE service=api LAST 1h
  • Alerting: notify when metric crosses a threshold
  • Dashboard visualization

Non-Functional

  • 1000 services Γ— 100 metrics each = 100k metric series
  • Each metric reports every 10s β†’ 10k writes/s average
  • Query latency: < 5s for dashboards
  • Retention: 1 year; older data downsampled

Core Design

Time-Series Data Model

A metric is a named, labeled time-series:

metric_name{label1=value1, label2=value2} value timestamp

cpu_usage{host="web-01", region="us-east-1"} 0.72 1709812800

Key property: write-heavy, append-only. You almost never update or delete individual data points. You do bulk deletes for retention.

Time-Series Database (TSDB)

Don't use a general-purpose DB. TSDBs are purpose-built:

TSDBNotes
PrometheusPull-based; local storage; excellent for short retention
InfluxDBPush-based; columnar storage; good query language (Flux)
TimescaleDBPostgreSQL extension; SQL queries; good for long retention
Cortex / ThanosHorizontally-scaled Prometheus; long-term S3 storage
VictoriaMetricsHigh-performance open-source; Prometheus-compatible
CloudWatch (AWS)Managed; integrates with AWS services

How TSDBs compress efficiently

  • Delta encoding: store differences between timestamps (4096, 4096+10, 4096+20 β†’ 0, 10, 10)
  • XOR compression for floats: consecutive similar values share most bits (Gorilla compression)
  • Result: 1.37 bytes per sample vs 16 bytes naΓ―vely

Pull vs Push

Pull (Prometheus model): server scrapes each target on a schedule. Target exposes /metrics endpoint.

  • Pros: simpler to reason about; server controls scrape rate; health check built-in (if scrape fails, target is down)
  • Cons: doesn't work well for short-lived jobs or serverless

Push (StatsD, InfluxDB model): clients push metrics to a collector.

  • Pros: works for ephemeral processes; lower coupling
  • Cons: client can flood the collector; harder to detect down hosts

Production systems often do both: push through an agent (Datadog agent), agent aggregates locally and pushes to backend.


Architecture

Services / Hosts
    ↓ (push to agent or expose /metrics)
Metrics Agent (per host, local aggregation)
    ↓ (batch send every 10s)
Ingestion Gateway (load balanced)
    ↓
Message Queue (Kafka) β€” partitioned by metric name
    ↓
TSDB Writers (parallel)
    ↓
Time-Series DB (VictoriaMetrics / Cortex)
    |                  |
Query Service     Alerting Service
    |                  |
Dashboards        Alert Manager β†’ PagerDuty / Slack

Downsampling for Retention

Storing full resolution for a year is expensive:

10s resolution (raw):     keep for 7 days
1-minute resolution:      keep for 30 days (aggregated from 10s)
5-minute resolution:      keep for 90 days
1-hour resolution:        keep for 1 year

Background compaction job rolls up raw data into lower-resolution aggregates, then deletes raw data after retention period.

Alerting

Alert rule:
  condition: avg(cpu_usage{service="api"}) LAST 5m > 0.90
  for: 3m     (must stay above threshold for 3 minutes β†’ avoid flapping)
  severity: critical
  notify: pagerduty

Alert evaluation: run every 30s; check last N minutes of TSDB
Alert state machine: OK β†’ PENDING (threshold exceeded) β†’ FIRING (for >= 3m)
Deduplication: don't alert multiple times for same condition
Routing: route by severity/team to correct on-call channel

Key Design Decisions

DecisionChoiceReason
StoragePurpose-built TSDBDelta/XOR compression; fast range queries
IngestKafka bufferAbsorb write spikes; decouple from storage
RetentionDownsamplingBalance query speed vs storage cost
Alert evaluationPull from TSDB on scheduleStateless alerter; easy to scale
AgentLocal pre-aggregationReduce network; handle microbursts

Security Considerations

ThreatMitigation
Metrics exfiltrationAuth on metrics query API; metrics may contain sensitive data (user counts, revenue)
Alert flooding (DoS)Rate limit alerts per channel; deduplication; alert fatigue prevention
Metric injectionValidate metric names and label values; reject excessively high cardinality labels
Cardinality explosionLimit unique label combinations per metric; reject metrics with high-cardinality labels (e.g., user_id)
Scrape endpoint exposure/metrics should be on internal network only; not publicly accessible

Interview Tips

TSDB is not just any database

explain why compression is critical (delta encoding, Gorilla) and how TSDBs are optimized for append-only time-series.

Downsampling

is key to cost management β€” show you understand the trade-off between resolution and retention cost.

Cardinality explosion

is a common production problem β€” mentioning it shows real experience with metrics systems.

Alerting state machine

with for duration prevents flapping β€” a subtle but important detail.