Design a Metrics Monitoring System
DifficultyHard | HelloInterview: problem breakdown
Problem Statement
Design a metrics monitoring and alerting system like Datadog, Prometheus, or CloudWatch. Collect time-series metrics from thousands of hosts, store them efficiently, visualize trends, and alert on anomalies.
πReal-world: The make-or-break insight is that time-series data is special enough to deserve its own database. You're writing millions of mostly-incrementing numbers tagged with labels, almost never updating, and querying by time range β so a purpose-built TSDB (Prometheus, InfluxDB, Facebook's Gorilla, the basis of much of this field) crushes a general database. Gorilla's famous trick: delta-of-delta timestamp encoding + XOR float compression shrank data ~12Γ and kept it in memory for fast alerting. Two themes this design hinges on: downsampling/rollups (keep raw data for hours, then aggregate to coarser resolution for long-term retention β you don't store per-second data for a year), and cardinality explosion β the classic footgun where a high-cardinality label like
user_idorrequest_idmultiplies your series into the millions and melts the system. Security relevance: this is the same architecture a SIEM/detection pipeline uses β high-ingest, time-windowed, alert-on-threshold β so the monitoring design and the detection design are cousins, and "alert fatigue" / hysteresis tuning applies to both.
Requirements
Functional
- Collect metrics from services (CPU, memory, request rate, error rate, custom metrics)
- Store metrics with high write throughput
- Query:
avg(cpu_usage) by host WHERE service=api LAST 1h - Alerting: notify when metric crosses a threshold
- Dashboard visualization
Non-Functional
- 1000 services Γ 100 metrics each = 100k metric series
- Each metric reports every 10s β 10k writes/s average
- Query latency: < 5s for dashboards
- Retention: 1 year; older data downsampled
Core Design
Time-Series Data Model
A metric is a named, labeled time-series:
metric_name{label1=value1, label2=value2} value timestamp
cpu_usage{host="web-01", region="us-east-1"} 0.72 1709812800Key property: write-heavy, append-only. You almost never update or delete individual data points. You do bulk deletes for retention.
Time-Series Database (TSDB)
Don't use a general-purpose DB. TSDBs are purpose-built:
| TSDB | Notes |
|---|---|
| Prometheus | Pull-based; local storage; excellent for short retention |
| InfluxDB | Push-based; columnar storage; good query language (Flux) |
| TimescaleDB | PostgreSQL extension; SQL queries; good for long retention |
| Cortex / Thanos | Horizontally-scaled Prometheus; long-term S3 storage |
| VictoriaMetrics | High-performance open-source; Prometheus-compatible |
| CloudWatch (AWS) | Managed; integrates with AWS services |
How TSDBs compress efficiently
- Delta encoding: store differences between timestamps (4096, 4096+10, 4096+20 β 0, 10, 10)
- XOR compression for floats: consecutive similar values share most bits (Gorilla compression)
- Result: 1.37 bytes per sample vs 16 bytes naΓ―vely
Pull vs Push
Pull (Prometheus model): server scrapes each target on a schedule. Target exposes /metrics endpoint.
- Pros: simpler to reason about; server controls scrape rate; health check built-in (if scrape fails, target is down)
- Cons: doesn't work well for short-lived jobs or serverless
Push (StatsD, InfluxDB model): clients push metrics to a collector.
- Pros: works for ephemeral processes; lower coupling
- Cons: client can flood the collector; harder to detect down hosts
Production systems often do both: push through an agent (Datadog agent), agent aggregates locally and pushes to backend.
Architecture
Services / Hosts
β (push to agent or expose /metrics)
Metrics Agent (per host, local aggregation)
β (batch send every 10s)
Ingestion Gateway (load balanced)
β
Message Queue (Kafka) β partitioned by metric name
β
TSDB Writers (parallel)
β
Time-Series DB (VictoriaMetrics / Cortex)
| |
Query Service Alerting Service
| |
Dashboards Alert Manager β PagerDuty / SlackDownsampling for Retention
Storing full resolution for a year is expensive:
10s resolution (raw): keep for 7 days
1-minute resolution: keep for 30 days (aggregated from 10s)
5-minute resolution: keep for 90 days
1-hour resolution: keep for 1 yearBackground compaction job rolls up raw data into lower-resolution aggregates, then deletes raw data after retention period.
Alerting
Alert rule:
condition: avg(cpu_usage{service="api"}) LAST 5m > 0.90
for: 3m (must stay above threshold for 3 minutes β avoid flapping)
severity: critical
notify: pagerduty
Alert evaluation: run every 30s; check last N minutes of TSDB
Alert state machine: OK β PENDING (threshold exceeded) β FIRING (for >= 3m)
Deduplication: don't alert multiple times for same condition
Routing: route by severity/team to correct on-call channelKey Design Decisions
| Decision | Choice | Reason |
|---|---|---|
| Storage | Purpose-built TSDB | Delta/XOR compression; fast range queries |
| Ingest | Kafka buffer | Absorb write spikes; decouple from storage |
| Retention | Downsampling | Balance query speed vs storage cost |
| Alert evaluation | Pull from TSDB on schedule | Stateless alerter; easy to scale |
| Agent | Local pre-aggregation | Reduce network; handle microbursts |
Security Considerations
| Threat | Mitigation |
|---|---|
| Metrics exfiltration | Auth on metrics query API; metrics may contain sensitive data (user counts, revenue) |
| Alert flooding (DoS) | Rate limit alerts per channel; deduplication; alert fatigue prevention |
| Metric injection | Validate metric names and label values; reject excessively high cardinality labels |
| Cardinality explosion | Limit unique label combinations per metric; reject metrics with high-cardinality labels (e.g., user_id) |
| Scrape endpoint exposure | /metrics should be on internal network only; not publicly accessible |
Interview Tips
explain why compression is critical (delta encoding, Gorilla) and how TSDBs are optimized for append-only time-series.
is key to cost management β show you understand the trade-off between resolution and retention cost.
is a common production problem β mentioning it shows real experience with metrics systems.
with for duration prevents flapping β a subtle but important detail.