SIEM & the Detection Data Pipeline
A detection is only as good as the data feeding it. This file covers how telemetry flows from a log source to an alert: collection, normalization, enrichment, storage, and querying — plus the modern "data-lake / SIEM-less" architecture that large consumer-scale companies (running on AWS + GCP) increasingly use instead of a traditional SIEM.
Last verified2026-06
What Is a SIEM, Really?
SIEM stands for Security Information and Event Management. Strip the marketing and it is three things bolted together:
that collects logs from everywhere (endpoints, cloud, network, apps, identity).
that holds those logs in a searchable form.
that runs rules against the logs and notifies humans.
Everything else (dashboards, case management, compliance reports) is built on top of those three.
Why the Pipeline Matters More Than the Tool
New detection engineers obsess over the SIEM product. Experienced ones obsess over the data getting into it correctly. A perfect detection rule fails silently if:
- the log source isn't onboarded (no data → no alert, and nobody notices)
- the timestamp is parsed wrong (events arrive "in the future" or out of order)
- a field got renamed in an agent update (the rule references a field that no longer exists)
- logs are dropped under load (sampling kicks in and the malicious event is the one dropped)
The blind-spot principlethe most dangerous gap is the one you don't know about. A SIEM that shows "0 alerts" looks identical whether you're secure or whether the log source died three weeks ago. Pipeline health monitoring is a core detection-engineering responsibility.
The End-to-End Pipeline
[ Log Sources ] [ Collection ] [ Processing ] [ Store + Detect ] endpoints (EDR/Sysmon) ──► agents / forwarders ─► parse ─► normalize ─► storage (hot/cold) cloud (CloudTrail, GCP ──► API pulls / streaming ▲ ▲ ├─ detection engine audit logs) (Kinesis, Pub/Sub) │ │ ├─ search / hunt network (Zeek, firewall)──► syslog / NetFlow enrich schema (OCSF) └─ dashboards/alerts identity (Okta, IdP) ──► webhooks / API (GeoIP, asset, applications ──► SDK / structured logs identity, threat-intel)
1. Collection
How logs leave the source and reach you:
software on the host pushes logs (Splunk Universal Forwarder, Elastic Beats, Fluent Bit, the Vector agent).
you poll a cloud provider's API (AWS CloudTrail, GCP audit logs, Okta System Log) on a schedule.
the source pushes to a message bus (AWS Kinesis, GCP Pub/Sub, Kafka) and you consume it. This is the dominant pattern at scale because it decouples producers from consumers and absorbs bursts.
network gear and SaaS apps push events directly.
2. Parsing
Raw logs arrive in wildly different formats — JSON, key-value, CEF, free-text syslog, CSV. Parsing extracts structured fields (source_ip, user, event_type) from each. Bad parsing is the #1 cause of broken detections: if source_ip doesn't get extracted, every rule referencing it silently fails.
3. Normalization (schemas)
Different sources call the same thing different names: src_ip, sourceIPAddress, client.ip, id.orig_h. Normalization maps them all to a single canonical field so one detection works across sources. The two dominant schemas:
vendor-neutral, backed by AWS/Splunk and a broad consortium; rapidly becoming the standard for security data lakes.
Elastic's schema, widely used in the Elastic ecosystem.
Why this matters in interviewsif you can write one detection for "suspicious login" that works whether the log came from Okta, AWS, or an SSH server, you've understood normalization. The schema is what makes that possible.
4. Enrichment
Adding context to an event before it hits detection, so the rule (and the analyst) has what it needs:
turn an IP into a country/ASN (detect impossible-travel logins).
is this a domain controller or a test box? Severity depends on it.
map an account to a human, department, and privilege level.
is this IP/domain/hash on a known-bad list?
Enrichment moves work left (do it once at ingest) instead of right (every analyst doing manual lookups during triage).
5. Storage: hot vs. cold
Security data is huge and most of it is never queried. The cost tradeoff:
fast, expensive, recent (last 30–90 days). Where detections run and active investigation happens.
cheap object storage (S3, GCS) for older data. Slower to query, but you need it for retrospective hunts and breach investigations that look back a year.
The retention tension: detections need speed (hot), but incident response and compliance need history (cold). You tier to balance cost against the reality that a breach discovered today may have started 200 days ago.
Traditional SIEM vs. Modern Data-Lake ("SIEM-less")
The architecture you'll meet at a large, cloud-native company differs from the classic enterprise SIEM.
| Traditional SIEM | Data-Lake / SIEM-less | |
|---|---|---|
| Examples | Splunk, QRadar, ArcSight, Sentinel | Panther, Matano, detections on BigQuery / Snowflake / Athena |
| Storage | Proprietary index | Open formats in object storage (S3/GCS, Parquet) |
| Cost model | License by data volume (GB/day) — expensive at scale | Pay for storage + compute separately; cheaper at huge volume |
| Detection | Vendor rule language, GUI | Detection-as-code (Python/SQL), Git-based |
| Scale ceiling | Strained at petabyte/consumer scale | Built for it (separates storage from compute) |
| Lock-in | High | Low (your data stays in open formats you own) |
Why consumer-scale shops (large consumer apps running on GCP + AWS) lean SIEM-less: at billions of events per day, traditional per-GB SIEM licensing becomes financially absurd, and the data volume exceeds what a single indexed product handles gracefully. Running detections directly on a data lake (e.g., scheduled SQL over BigQuery, or Python detections in Panther reading from S3) decouples storage cost from compute and keeps the data in formats the company already operates at scale.
The tradeoffdata lakes give you scale and cost control but less out-of-the-box polish — you build more of the detection, enrichment, and case-management layers yourself. That's exactly the work a detection engineer does.
Detection Latency: Streaming vs. Batch
When a detection runs shapes what it can catch:
the detection evaluates each event as it arrives (seconds of latency). Needed for high-urgency signals: active credential abuse, ransomware behavior. More expensive to run.
the detection runs every N minutes/hours over accumulated data. Fine for slower-burn signals: a user accessing an unusual volume of records over a day. Cheaper; enables complex aggregations across large windows.
Most programs run a hybrid: a small set of streaming detections for the things that need instant response, and a larger set of scheduled detections for everything else.
Pipeline Health & Detection Reliability
Because a dead pipeline looks like a quiet one, you actively monitor the plumbing:
alert when a log source's event volume drops to zero or deviates sharply from baseline ("Sysmon from finance subnet stopped reporting").
alert when expected fields go missing or change type after a source update.
track the delay between event time and ingest time; growing lag means detections are evaluating stale data.
capture events that failed to parse so they're visible and fixable, not silently dropped.
Interview gold"a SIEM showing zero alerts is ambiguous — it could mean you're secure or it could mean the data stopped flowing. That's why I monitor source health as a first-class detection." This single insight separates engineers from operators.
Interview Questions
It's generated at the source (say an endpoint), shipped by an agent or streamed onto a bus like Kinesis or Pub/Sub. On the way in it's parsed — structured fields extracted from the raw format — then normalized to a common schema like OCSF so a field like source IP has one canonical name regardless of source. It's enriched with context (GeoIP, asset criticality, identity, threat-intel matches), written to a hot store for fast querying, and the detection engine evaluates rules against it — either in real time for urgent signals or on a schedule for slower aggregations. A match generates an alert, already carrying the enrichment context the analyst needs to triage. Older data ages out to cheap cold storage for retrospective hunts.
Cost and scale. Traditional SIEMs license by data volume, so at billions of events a day the bill becomes untenable, and a single indexed product strains at that volume. A data-lake approach — detections running on BigQuery, Snowflake, or Athena, or a platform like Panther reading from S3 — separates storage from compute, keeps data in open formats the company already operates at scale, and avoids per-GB lock-in. The tradeoff is you build more of the detection and case-management layers yourself instead of getting them out of the box, but for a cloud-native shop that's an acceptable trade for the scale and cost control.
Different log sources name the same concept differently — source IP might be src_ip, sourceIPAddress, or client.ip. Normalization maps all of them to a single canonical field defined by a schema like OCSF or ECS. It matters because it lets you write one detection that works across every source: a "suspicious login" rule fires whether the event came from Okta, AWS, or SSH, because they all normalized to the same fields. Without it, you'd rewrite every detection per source and silently miss events whose fields didn't get extracted.
Most likely the pipeline broke, not the rule: the log source stopped reporting, an agent update renamed a field the rule depends on, or events are being dropped under load. The dangerous part is it's silent — zero alerts looks identical to "all clear." Prevention is treating pipeline health as a first-class detection: heartbeat monitoring that alerts when a source's volume drops to zero or deviates from baseline, schema-drift alerts when expected fields disappear, ingest-lag tracking, and dead-letter queues for parse failures. Plus continuous detection validation with Atomic Red Team so I'd catch a dark detection by testing rather than waiting for a real attack.
Real-time/streaming for signals where minutes matter — active credential abuse, ransomware file-encryption behavior, a root account doing something dangerous — because the response window is short and worth the higher compute cost. Scheduled/batch for slower-burn or aggregate signals — a user accessing an unusual volume of records over a day, low-and-slow exfiltration — where you need to aggregate across a large time window and instant latency adds no value. Most programs are hybrid: a lean set of streaming detections for the urgent few, and a larger set of scheduled ones for everything else, to control cost.