Detection Engineering
Detection engineering is the discipline of building, testing, tuning, and maintaining the logic that turns raw telemetry into actionable alerts. It is the spine of any "Detection & Response" role. This file covers the lifecycle, the rule languages, how to measure quality, and how interviewers probe whether you actually think like a detection engineer.
Last verified2026-06
What Is Detection Engineering?
A SOC analyst responds to alerts. A detection engineer creates and improves the alerts. The job is software engineering applied to threat detection: you write detection logic as code, test it, version-control it, measure its quality, and retire it when it stops earning its keep.
The mental shift that separates a detection engineer from a tool operator: you don't ask "what does my SIEM alert on?" — you ask "what attacker behavior do I want to catch, what evidence does it leave, and how do I write logic that fires on that evidence without drowning in false positives?"
Why It Matters
Out-of-the-box rules catch commodity attacks. Targeted intrusions evade them. Custom detections close the gap between "what the vendor ships" and "what your environment actually looks like."
A detection that fires 500 times a day with a 2% true-positive rate trains analysts to ignore it. Detection engineering is as much about suppressing noise as it is about catching threats.
Detection engineering lets you answer "which attacker techniques can we actually detect?" with data, not vibes — using frameworks like MITRE ATT&CK.
The Detection Engineering Lifecycle
A detection is not write-once. It moves through a repeatable lifecycle, just like a software feature.
1. HYPOTHESIZE "Attackers using technique X leave evidence Y in log source Z."
│
2. RESEARCH Find the data source. Confirm Y actually appears in Z.
│ Generate the behavior in a lab (Atomic Red Team).
│
3. BUILD Write the detection logic (Sigma/SPL/KQL/YARA-L).
│ Define severity, ATT&CK mapping, and response runbook.
│
4. TEST Run against (a) known-malicious data → does it fire?
│ (b) 30 days of production data → how noisy is it?
│
5. TUNE Add exclusions for known-good. Adjust thresholds.
│ Target a true-positive rate you can live with.
│
6. DEPLOY Ship as code (PR review, CI validation). Document it.
│
7. MONITOR Track fire rate, TP/FP ratio, and analyst feedback.
│
8. RETIRE If the log source dies, the threat is gone, or a better
detection supersedes it — remove it. Dead rules are debt.Interview signalwhen handed a detection idea, walk through this loop out loud. Don't jump straight to "I'd write a rule" — name the hypothesis, the data source, how you'd test for noise, and how you'd tune.
Detection-as-Code
Modern detection teams treat detections like application code:
detections live in Git. Every change is a reviewed pull request.
a pipeline lints the rule syntax, runs it against test data (unit tests with known-malicious and known-benign samples), and blocks merges that break.
shared functions for common enrichment (GeoIP, asset criticality, identity lookups).
another engineer reviews logic and false-positive risk before deploy.
Platforms built around this model: Panther (Python detections-as-code), Matano (open-source, OCSF-native), Sigma + converters, Elastic detection rules (YAML in Git), Chronicle/Google SecOps (YARA-L rules in repos).
# Example: a Panther-style detection (Python). The `rule` function returns
# True when the event is suspicious. This is testable like any unit.
def rule(event):
# Hypothesis: AWS root account usage is rare and high-risk.
return (
event.get("userIdentity", {}).get("type") == "Root"
and event.get("eventName") not in ALLOWED_ROOT_EVENTS # tuned exclusion
)
def title(event): # the alert headline analysts see
return f"AWS Root activity: {event.get('eventName')} from {event.get('sourceIPAddress')}"
def severity(event):
return "HIGH"Detection Rule Languages
You will be expected to read — and ideally write — at least one of these. Know the shape of each; you can look up exact syntax.
Sigma — the vendor-neutral standard
Sigma is YAML that describes a detection abstractly, then converts to whatever SIEM you run (Splunk, Elastic, Chronicle, Sentinel). Write once, deploy anywhere.
title: Suspicious PowerShell Download Cradle
logsource:
product: windows
category: process_creation
detection:
selection:
Image|endswith: '\powershell.exe'
CommandLine|contains:
- 'DownloadString'
- 'IEX'
- 'Net.WebClient'
condition: selection # fire when 'selection' matches
falsepositives:
- Legitimate admin scripts
level: high
tags:
- attack.execution
- attack.t1059.001 # ATT&CK technique IDSplunk SPL — search-driven
index=windows EventCode=4688 NewProcessName="*\\powershell.exe"
| where match(CommandLine, "(?i)(downloadstring|iex|net\.webclient)")
| stats count by host, User, CommandLine
| where count > 0KQL — Microsoft Sentinel / Defender
DeviceProcessEvents
| where FileName == "powershell.exe"
| where ProcessCommandLine has_any ("DownloadString", "IEX", "Net.WebClient")
| summarize count() by DeviceName, AccountName, ProcessCommandLineYARA-L — Google SecOps (Chronicle)
rule suspicious_powershell_download {
events:
$e.metadata.event_type = "PROCESS_LAUNCH"
$e.target.process.file.full_path = /powershell\.exe$/ nocase
$e.target.process.command_line = /downloadstring|iex|net\.webclient/ nocase
condition:
$e
}The common shapefilter to the right log source → match the suspicious field values → group/aggregate → apply a threshold. Once you see that pattern, the languages are interchangeable.
Mapping to MITRE ATT&CK
ATT&CK is the shared vocabulary for what attackers do. Every detection should map to one or more technique IDs (e.g., T1059.001 = PowerShell). This lets you:
which techniques can you detect, partially detect, or not at all?
build detections for techniques your actual threat actors use (threat-informed defense).
"we have 3 detections for credential dumping (T1003)" is a sentence leadership and other teams understand.
Toolsthe ATT&CK Navigator colors a heatmap of your coverage. DeTT&CT scores data-source quality and detection coverage per technique. The honest output is a map with green (solid coverage), yellow (partial/noisy), and red (blind spots) — and red is where your next sprint goes.
Coverage caveat"we have a detection for T1059" does not mean you detect all PowerShell abuse. Coverage is about depth and evasion-resistance, not a checkbox. A detection that any attacker trivially bypasses is yellow at best.
Detection Quality: The Metrics
The two failure modes are false positives (noise → alert fatigue) and false negatives (misses → breaches). Detection quality is the balance.
| Metric | What it measures | Why it matters |
|---|---|---|
| True Positive Rate (precision) | of alerts that fired, how many were real | Low precision = analyst burnout, ignored alerts |
| False Positive Rate | benign events that wrongly fired | The primary tuning target |
| Recall / coverage | of real attacks, how many we caught | High FN = silent breaches |
| MTTD (Mean Time To Detect) | dwell time before detection | Directly tied to breach cost |
| MTTR (Mean Time To Respond) | detect → contained | Measures the whole D&R pipeline |
| Alert volume / analyst / shift | sustainable load | If unsustainable, the SOC is broken regardless of coverage |
The precision/recall tradeoffyou can catch everything by alerting on everything (perfect recall, useless precision), or alert on nothing (perfect precision, zero recall). Good detection engineering finds the threshold where a rule is sensitive enough to catch the threat but specific enough that analysts trust it. State this tradeoff explicitly in interviews.
Detection Validation & Testing
You cannot claim a detection works until you've seen it fire on real malicious activity and stay quiet on benign activity.
a library of small, mapped-to-ATT&CK tests that safely execute attacker techniques so you can confirm your detection fires. The fastest way to test a single technique.
red team executes a technique, blue team confirms detection (or finds the gap), in a tight feedback loop. The single highest-value detection-improvement activity.
automated, continuous tools (e.g., commercial BAS platforms) that re-run attack scenarios on a schedule to catch detection drift.
store known-malicious sample events as test fixtures; CI replays them on every rule change to ensure you didn't break a detection.
The drift problema detection that worked at deploy can silently break — a log source changes format, a field is renamed, an agent stops reporting. Continuous validation catches "detection went dark" before an attacker does.
The Pyramid of Pain (why TTP detections beat IOC detections)
David Bianco's Pyramid of Pain ranks indicators by how much it hurts the attacker when you detect on them:
▲ TTPs ← hardest to change; detecting these forces
│ Tools the attacker to retool. Most durable.
│ Network/Host Artifacts
│ Domain Names
│ IP Addresses
▼ Hash Values ← trivial to change; detecting these annoys
the attacker for minutes. Most volatile.A hash-based detection breaks the moment the attacker recompiles. A behavior-based detection ("process spawned from Office that then spawns PowerShell with an encoded command") forces them to change how they operate — far more expensive. Build toward the top of the pyramid. This is the single most cited concept in detection interviews.
Interview Questions
I start with a hypothesis tied to attacker behavior, e.g. "adversaries dumping LSASS leave a process-access event where a non-system process opens lsass.exe with specific access rights." I confirm the data source actually captures that (Sysmon Event ID 10, or EDR telemetry), then generate the behavior safely with Atomic Red Team to get a known-malicious sample. I write the logic, test it against that sample (does it fire?) and against 30 days of production data (how noisy is it?). I tune with exclusions for known-good callers, map it to ATT&CK (T1003.001), set severity and a response runbook, then ship it as a reviewed PR. After deploy I monitor fire rate and analyst feedback and tune further.
First I'd confirm whether it's catching anything real — pull the historical true-positive rate. If it's near zero, the rule is pure noise and either needs aggressive tuning or retirement. I'd analyze the false positives for a common pattern (a specific service account, scanner, or admin tool) and add targeted exclusions rather than broadening the rule. If the underlying behavior is genuinely too common to alert on directly, I'd convert it from an alert into an enrichment or a hunting input — context that supports other detections rather than a standalone page. The goal is to protect analyst trust; a rule nobody acts on is worse than no rule, because it adds noise and creates a blind spot people assume is covered.
Because of the Pyramid of Pain. IOCs like hashes and IPs sit at the bottom — the attacker changes them in seconds by recompiling or rotating infrastructure, so an IOC detection has a very short shelf life. Behavioral detections target TTPs at the top of the pyramid — the way an attacker operates, like Office spawning a script interpreter that makes a network connection. Changing that forces the attacker to fundamentally retool, which is expensive and slow. IOC detections are still useful for fast, cheap blocking of known-bad, but durable coverage comes from behavior.
I'd look at it from two directions. Coverage: map detections to MITRE ATT&CK and score them with something like DeTT&CT, weighted toward the techniques our actual threat actors use, and track the red/yellow/green map shrinking its red over time. Quality: per-detection precision (true-positive rate), overall alert volume per analyst per shift (is the load sustainable?), and pipeline outcomes — MTTD and MTTR. Critically, I'd validate continuously with purple-team exercises and Atomic Red Team so I know detections still fire and haven't drifted. Coverage without quality just means a lot of noisy rules; both have to move together.
Detection-as-code means treating detections like software: they live in Git, every change is a peer-reviewed pull request, and a CI pipeline lints the syntax and runs the rule against known-malicious and known-benign test fixtures before merge. You adopt it for the same reasons you version application code — change history and rollback, peer review to catch false-positive risk before production, automated regression testing so a rule edit doesn't silently break another, and reproducibility across environments. Platforms like Panther and Matano are built around this; Sigma plus a Git repo achieves it vendor-neutrally.
You can't — and recognizing that is the point. The first step in detection engineering is data-source validation: confirming the telemetry that would reveal the behavior actually exists and is being collected. If it doesn't, the detection task becomes a visibility task: onboard the missing log source (enable Sysmon, turn on CloudTrail data events, deploy an EDR sensor), or find a proxy signal elsewhere in the kill chain. I'd flag the blind spot explicitly on the ATT&CK coverage map as red so it's a visible, prioritized gap rather than a silent assumption that we're covered.