AWS Detection, Logging & Response Services
This file covers the AWS services that sit between raw logs and a closed incident: where logs go, how you query them, why they go missing, how you investigate, and how you automate and rehearse the response. It complements AWS fundamentals (what GuardDuty, Security Hub, Config and Macie are) and the AWS IR playbooks (what to do per scenario).
Last verified2026-10 — AWS adds and renames security features often; check current service documentation before relying on limits or feature names.
Logging Architecture — Where Every Log Should Land
What is this?
AWS produces many separate logs: CloudTrail records API calls, VPC Flow Logs record network connections, Route 53 Resolver query logs record DNS lookups, and applications write their own logs. A logging architecture decides which of these you switch on, where they are stored, who can delete them, and how you query them.
Why it matters
An investigation can only answer questions the logs can answer. The most common failure in a cloud incident is not a missing detection, it is discovering that data events, DNS logs or the affected region were never logged. Attackers who get admin also try to delete logs, so where logs live matters as much as whether they exist.
How it works
The standard multi-account pattern:
Member accounts (prod, dev, …) Log Archive account (Security OU)
┌───────────────────────────────┐ ┌──────────────────────────────────┐
│ Organisation CloudTrail ──────┼──────────────►│ S3 bucket: Object Lock, KMS, │
│ VPC Flow Logs ────────────────┼──────────────►│ bucket policy denies delete, │
│ Route 53 Resolver query logs ─┼──────────────►│ only the org trail may write │
│ CloudWatch Logs (apps, agent) ┼─ subscription►│ (via Firehose) │
└───────────────────────────────┘ └──────────────┬───────────────────┘
│ read-only access
Security Tooling account (delegated
admin: GuardDuty, Security Hub,
Detective, Security Lake, Athena)Key log sources and what each is for:
control-plane API calls (CreateUser, RunInstances). On by default for 90 days in Event history; a trail is needed for long retention.
object- and function-level calls (s3:GetObject, lambda:Invoke). Off by default and billed per event, but they are the only proof of what data was read. Use advanced event selectors to log only sensitive buckets.
created once in the management account (or a delegated admin), it logs every member account, and member accounts cannot switch it off.
5-tuple connection records per network interface, subnet or VPC. They do not capture traffic to the Amazon DNS resolver, the instance metadata service (169.254.169.254), Amazon Time Sync or DHCP, so a metadata-credential theft never shows up there.
the same idea for traffic crossing a transit gateway between VPCs and on-prem; needed when east–west traffic never touches a single VPC's interfaces.
every DNS query made from a VPC. This is where DNS tunnelling, beaconing and lookups of known-bad domains show up. Pair with Route 53 Resolver DNS Firewall to block domains.
ships OS and application logs and metrics from EC2 and on-prem servers into CloudWatch Logs.
ALB/NLB access logs, CloudFront logs, API Gateway access and execution logs, WAF logs, S3 server access logs, EKS control-plane audit logs. Each is switched on per resource.
Amazon Security Lake
Security Lake is a managed data lake that collects AWS and third-party security logs from every account and region, converts them to OCSF (Open Cybersecurity Schema Framework), and stores them as Parquet in S3 in a delegated admin account.
- Native sources include CloudTrail management and data events, VPC Flow Logs, Route 53 Resolver query logs, Security Hub findings, EKS audit logs and WAF logs; custom sources can write OCSF too.
- Subscribers (your SIEM, Athena, OpenSearch, a third-party tool) get either query access through Lake Formation or a notification when new objects land.
- Why OCSF matters: a field like "source IP" has the same name whatever product produced it, so one detection query works across sources. See SIEM normalisation.
Querying logs
- CloudWatch Logs Insights — interactive queries over CloudWatch log groups:
fields @timestamp, userIdentity.arn, eventName, sourceIPAddress
| filter eventSource = "iam.amazonaws.com" and eventName like /Create|Attach|Put/
| stats count(*) as calls by userIdentity.arn, bin(1h)
| sort calls desc
| limit 20SQL over logs in S3 (CloudTrail, Flow Logs, Security Lake). Cheapest for large, historical hunts; partition by date and account to control cost.
search and dashboards with near-real-time ingestion; used when you need correlation and visualisation. Lambda functions are commonly used to parse and enrich logs on the way in.
turn a log pattern (for example a root console login) into a CloudWatch metric and alarm. This is how the CIS AWS Foundations benchmark alarms are built.
In practiceWhat reconnaissance with a stolen key looks like in CloudTrail — the event most cloud investigations start from:
{ "eventTime": "2026-10-03T02:14:09Z", "eventSource": "iam.amazonaws.com", "eventName": "ListAttachedUserPolicies",
"userIdentity": { "type": "IAMUser", "userName": "ci-deploy", "accessKeyId": "AKIA4EXAMPLE7Q" }, ← long-lived key
"sourceIPAddress": "185.220.101.44", ← not the CI runner this user normally calls from
"userAgent": "aws-cli/2.15.0 Python/3.11 Linux/6.5",
"errorCode": "AccessDenied" } ← recon often shows as a burst of denied calls🎯On the job. The highest-value CloudTrail alerts are cheap: root sign-in,
StopLogging/DeleteTrail/PutEventSelectors, new access keys for existing users,CreateUseroutside your identity pipeline, and bursts ofAccessDeniedfrom one principal. GuardDuty'sStealth:IAMUser/CloudTrailLoggingDisabledfinding covers the most important one.
Security angle
- Protect the log bucket from the people being logged: separate account, Object Lock, deny
s3:DeleteObjectandPutBucketPolicyto everyone but a break-glass role, and an SCP denyingcloudtrail:StopLogging,DeleteTrailandUpdateTrail. - Turn on CloudTrail log file integrity validation so you can prove logs were not altered (SHA-256 digests signed hourly).
- Log all regions, including ones you "don't use"; attackers launch miners in the regions nobody watches.
Troubleshooting Missing Logs
What is this?
A checklist for the common exam and real-world scenario: "logging is configured, but nothing arrives."
Why it matters
Silent logging failures are worse than no logging, because everyone believes coverage exists. Detection engineers treat pipeline health as a detection in its own right (see pipeline health).
How it works
Almost every case is a permission, a setting that was never switched on, or a network path:
the bucket policy must allow the cloudtrail.amazonaws.com service principal to s3:PutObject (scoped with aws:SourceArn to the trail). If the trail uses a KMS key, the key policy must let CloudTrail call kms:GenerateDataKey*.
the instance role lacks CloudWatchAgentServerPolicy; the agent config file points at the wrong log path; the agent is not running; or a private subnet has no route to CloudWatch (needs a NAT gateway or a logs VPC endpoint).
the execution role lacks logs:CreateLogGroup, logs:CreateLogStream and logs:PutLogEvents (the AWSLambdaBasicExecutionRole managed policy).
REST APIs need an account-level CloudWatch Logs role ARN set in API Gateway settings, then logging enabled per stage. Access logs are a separate per-stage setting.
standard logging must be enabled per distribution and the destination must accept it; real-time logs go through Kinesis Data Streams.
check the excluded traffic list above, the capture filter (ACCEPT, REJECT or ALL), and the delivery role's permissions.
the reader has S3 access but not kms:Decrypt on the bucket's key.
In practiceTwo commands answer "is the trail actually working?":
$ aws cloudtrail get-trail-status --name org-trail --query "{logging:IsLogging,lastDelivery:LatestDeliveryTime,error:LatestDeliveryError}"
{ "logging": true, "lastDelivery": "2026-10-11T09:55:02Z", "error": "AccessDenied" } ← "on", but delivery is failing
$ aws cloudtrail validate-logs --trail-arn arn:aws:cloudtrail:…:trail/org-trail --start-time 2026-10-10T00:00:00Z
Results: 24/24 digest files valid, 1302/1302 log files valid ← no tamperingSecurity angle
An attacker with enough permissions may change a bucket policy or key policy rather than stop the trail, so logging breaks without a StopLogging event. Alert on gaps: no CloudTrail delivery for N minutes, PutBucketPolicy or PutKeyPolicy on logging resources, and CloudWatch agent heartbeat loss.
Investigation — Amazon Detective
What is this?
Detective is an investigation tool. It builds a behaviour graph from CloudTrail, VPC Flow Logs, GuardDuty findings and EKS audit logs, keeping up to a year of history, and lets you pivot: "this role → which IPs used it → what else those IPs touched → what the baseline for this role looks like."
Why it matters
GuardDuty tells you something happened; Detective helps answer how far it went and what is normal for this entity. The exam phrase "root cause analysis" or "scope of an event" usually points to Detective.
How it works
- A GuardDuty finding fires, for example credential use from an unusual IP.
- From the finding, open Detective: it shows the role's API volume over time, new geolocations, and new user agents.
- Pivot to the IP: which other principals and instances it interacted with.
- Detective finding groups cluster related findings into one incident, so a single attack sequence is not triaged as ten tickets.
Detective requires GuardDuty to be enabled and is managed through a delegated admin account.
🎯On the job. A good Detective session answers three questions you'll write in the incident report: what is normal for this role or instance (the baseline), when did behaviour first deviate (the start of the incident), and what else did the same IPs or sessions touch (the scope).
Security angle
Detective does not detect or block anything; it is for analysts. For automated cross-signal correlation, GuardDuty Extended Threat Detection produces "attack sequence" findings across multiple events.
Incident Response Preparation, Automation and Testing
What is this?
The services that turn a written IR plan into something that runs: runbooks that execute steps, workflows that chain them, fault-injection to rehearse, and backups to recover from.
Why it matters
SCS-C03 domain 2 asks you to design and test a plan, not just respond. Interviewers ask "how would you make this repeatable?": the answer is automation with a human approval step for destructive actions.
How it works
GuardDuty / Security Hub finding
│
▼
EventBridge rule (match finding type + severity)
│
├──► SNS / chat alert to on-call
│
└──► Step Functions workflow (or Systems Manager Automation runbook)
├─ 1. tag the resource "under-investigation"
├─ 2. snapshot EBS volumes, capture memory (forensics account)
├─ 3. swap security group to an isolation SG (no inbound/outbound)
├─ 4. detach instance profile / revoke role sessions
└─ 5. open an OpsItem / ticket with evidence linksrunbooks (YAML documents) with steps such as aws:executeAwsApi and aws:approve. AWS publishes ready-made runbooks for isolating instances and collecting diagnostics.
aggregates operational issues as OpsItems, with related resources and runbooks attached.
keeps instances in a defined state on a schedule (agent installed, configuration applied); used for regular assessments.
orchestrates multi-step responses with branching, retries and human approval.
an AWS solution that, on a Security Hub finding, isolates the instance and captures disk snapshots and memory into a dedicated forensics account.
routing controls and zonal shift to move traffic away from an impaired region or Availability Zone.
runs controlled experiments (stop instances, throttle APIs, cut network) to prove the response and recovery actually work.
assesses an application against target RTO and RPO (recovery time and recovery point objectives) and recommends fixes.
central backup plans across services; restores are the "recover" step after ransomware or destructive attacks. Details in data protection.
Preparation that pays off at 3 a.m.: pre-provisioned IR roles in every account (assumable only from the security account), a forensics account with a pre-shared KMS key, an isolation security group template, Shield Advanced proactive engagement for internet-facing apps, and game days that run the playbooks for real.
In practiceThe EventBridge rule that wires a finding to automation:
{ "source": ["aws.guardduty"], "detail-type": ["GuardDuty Finding"],
"detail": { "type": [ { "prefix": "UnauthorizedAccess:IAMUser/InstanceCredentialExfiltration" } ],
"severity": [ { "numeric": [">=", 7] } ] } }Its target is a Step Functions workflow that snapshots, captures memory, isolates and opens a ticket, with a manual approval before anything destructive.
Security angle
- Automation runs with powerful roles; scope them to the minimum (tag, snapshot, change security group) and protect who can edit runbooks.
- Auto-containment can tip off an attacker or destroy volatile evidence. Capture memory before isolation or termination, and require approval for destructive steps. See analyst OPSEC.
Interview Questions
I'd create an organisation trail with management events plus data events for sensitive buckets, and send it, VPC Flow Logs and Route 53 Resolver query logs to an S3 bucket in a dedicated log-archive account with Object Lock, KMS and a bucket policy that only lets the trail write and nobody delete. An SCP stops anyone turning CloudTrail off, and a separate security-tooling account gets read access and runs GuardDuty, Security Hub and Athena or Security Lake as delegated admin. The point is that even full admin in a workload account can't erase the evidence.
Because Flow Logs deliberately exclude traffic to the metadata service at 169.254.169.254, along with Amazon DNS, Time Sync and DHCP. So the request that stole the role credentials is invisible at the network layer. You catch it instead through GuardDuty's instance-credential-exfiltration findings, CloudTrail showing the role used from an IP outside AWS, and by enforcing IMDSv2 so SSRF can't reach it in the first place.
Security Lake is a managed data lake that pulls CloudTrail, Flow Logs, DNS logs, Security Hub findings and third-party sources from every account and region into S3, normalised to the Open Cybersecurity Schema Framework. OCSF means a source IP or user name has the same field name regardless of which product produced it, so one detection query works across all sources. It replaces a lot of custom parsing and lets you plug in any SIEM as a subscriber.
First the bucket policy, which must let the CloudTrail service principal put objects, ideally scoped by source ARN to that trail. Then, if the trail encrypts with KMS, the key policy must let CloudTrail generate data keys. After that I'd check the trail is actually logging and pointing at the right bucket and prefix. And I'd treat a sudden stop as a possible attack, because changing a bucket or key policy is a quieter way to blind you than calling StopLogging.
When I need scope and root cause rather than detection. Detective keeps up to a year of behaviour-graph data from CloudTrail, Flow Logs and GuardDuty, so I can see whether this role normally calls these APIs from these IPs, what else the suspicious IP touched, and group related findings into one incident. GuardDuty says "something is wrong"; Detective answers "how far did it go and what's normal here."
An EventBridge rule on the GuardDuty finding starts a Step Functions or Systems Manager Automation workflow that first tags the instance and captures memory and EBS snapshots into a forensics account, and only then swaps it to an isolation security group and revokes the role's active sessions. Destructive steps like termination sit behind a manual approval. Ordering matters: isolating or stopping first can wipe the volatile evidence you need for root cause.
Run it. Tabletop exercises test the people and decisions, and game days execute the actual runbooks in a non-production account. I'd use AWS Fault Injection Service to inject failures and Resilience Hub to check recovery against our RTO and RPO targets. Then I'd measure time to detect, time to contain and whether the logs we needed existed, and feed the gaps back into detections and runbooks.