Scenario: Turning On the Lights — Logging, Alerting, and Hunting in AWS
The ir-scenario lab is the break-in: an over-privileged role, a
forgotten static key, a victim EC2, an SSRF that steals credentials off the metadata
service. The whole time, a defender with no logging sees none of it. This lab is
that defender catching up — and it's the difference between "we think someone got in"
and "here is every call they made, in order, with timestamps."
Last verified2026-06
Read this once, then run the exercises with the ir-scenario lab deployed alongside.
First: what CloudTrail actually is (and isn't)
Plain EnglishCloudTrail is the answering-machine tape for the AWS API. Almost everything in AWS — the console, the CLI, an SDK, one service calling another — is an API call to the AWS control plane. CloudTrail writes down who made each call, when, from what IP, with what parameters, and whether it succeeded. That's it. That single fact is the backbone of nearly every cloud investigation.
What it records (an event) looks like this — the fields you'll actually grep for:
{
"eventTime": "2026-06-17T10:22:31Z",
"eventName": "GetParameter", // the API action
"eventSource": "ssm.amazonaws.com", // the service
"awsRegion": "us-east-1",
"sourceIPAddress":"203.0.113.9", // where it came from
"userIdentity": { // WHO — the most important block
"type": "AssumedRole",
"arn": "arn:aws:sts::1234:assumed-role/ir-lab-prod-app/i-0abc...",
"accessKeyId": "ASIA...", // ASIA = temporary (STS) creds
"sessionContext": { "attributes": { "mfaAuthenticated": "false" } }
},
"errorCode": "AccessDenied" // present ONLY when the call failed
}What it is NOT
"Who called s3:GetObject" is a management/data event CloudTrail
can log; "what bytes were inside the object" is never in CloudTrail. Reading a row out
of a database, the contents of a file — that's the data plane, invisible here.
It records only what happens after the trail exists. This is why you stand logging up on day zero, not after the incident.
Events reach CloudWatch Logs in ~5–15 min and the S3 archive a little later. Good enough to alert on; not a real-time packet capture.
Memory hookOne-linerCloudTrail answers "who did what, where, and did it work?" for every AWS API call — it's the control-plane audit log, not a data-plane wiretap.
The three layers this lab builds
One trail feeds three different jobs. Knowing which layer answers which question is the whole point — and a classic interview question in itself.
┌─────────────────────────────────────────────┐
every AWS API call ──►│ CloudTrail (one trail) │
└───────┬─────────────────────────┬───────────┘
│ │
CloudWatch Logs │ │ S3 (durable archive)
▼ ▼
metric filters + alarms Athena (SQL over history)
= "is it happening NOW?" = "reconstruct what happened"
(real-time, seconds-to-minutes) (forensics / threat hunting)
── and in parallel ──► GuardDuty = "does AWS think this matches a known attack?"
(managed; reads CloudTrail + VPC Flow + DNS for you)| Layer | Question it answers | Latency | In this lab |
|---|---|---|---|
| CloudWatch metric filters + alarms | "Is a dangerous action happening right now?" | minutes | 5 detections → SNS |
| Athena over S3 | "Show me everything X did between these times" | query-time | saved hunt queries |
| GuardDuty | "Does this look like a known attack pattern?" | ~15 min | one detector |
The detections (real-time layer)
These are written as detections-as-code in detections.tf: a CloudWatch Logs
metric filter (a pattern that adds 1 to a counter on each matching event) wired to an
alarm that fires when the count goes above zero and publishes to SNS. They map to the
CIS AWS Benchmark monitoring controls — the canonical "you have CloudTrail, now what do
you alert on?" list.
| Detection | What it catches | Why it matters |
|---|---|---|
root-account-usage | Any action by the root identity | Root should be sealed away with MFA. Any use is an incident. |
unauthorized-api-calls | Bursts of AccessDenied / UnauthorizedOperation | The fingerprint of enumeration — an attacker probing what their stolen identity can do. |
console-login-no-mfa | An IAM user signing in without MFA | Either a misconfigured user or a password-only credential being abused. |
iam-policy-changes | Create/attach/detach/delete of IAM policies | Privilege escalation and persistence happen by editing IAM. |
cloudtrail-tampering | StopLogging / DeleteTrail / UpdateTrail / PutEventSelectors | The adversary's first move is often to blind the logging. Alerting on it is your tripwire. |
The deep point — anti-tamperThe single most important alarm is the last one. A smart attacker who lands admin will try to turn CloudTrail off so the rest of their activity goes unrecorded. But
StopLoggingis itself an API call — so it lands in CloudWatch Logs and trips the alarm before the trail stops feeding S3. You detect the blinding by the act of blinding. (Defense in depth: in a real org the trail lives in a separate logging account the workload identities can't touch at all.)
How an attack appears in CloudTrail (tie-back to ir-scenario)
Walk the ir-scenario attack chain and name the event each step leaves behind. This is
exactly the "trace the intrusion" exercise an IR interview runs you through.
ir-scenario step CloudTrail event you'd hunt for
─────────────────────────────────────────────────────────────────────────────────────
1. SSRF makes the app hit IMDS, steal creds (none — IMDS is link-local, not an API call)
2. attacker uses the role's temp creds calls with userIdentity.accessKeyId = ASIA…,
arn = assumed-role/ir-lab-prod-app/i-…
3. recon: aws iam list-users, s3 ls ListUsers, ListBuckets (often the FIRST tells)
4. read the secret: ssm get-parameter GetParameter on /ir-lab/prod/db-password
5. read state bucket: s3 cp …tfstate GetObject (only if S3 DATA events are on)
6. use the static CI key (persistence) calls with accessKeyId = AKIA… (AKIA = permanent!)
7. attacker disables logging StopLogging / DeleteTrail → tampering alarmThe credential tell
userIdentity.accessKeyIdstartingASIAis a temporary STS credential (a role);AKIAis a permanent IAM-user key. SeeingAKIAmake API calls from a workload — or from a new IP — is the signature of a stolen static key, exactly their-lab-prod-ci-servicecredential from the other lab. The prefix tells you in one glance whether you're chasing a role session or a leaked long-lived key.
A subtle catch from step 1: the credential theft itself is invisible to CloudTrail —
IMDS lives at a link-local address and isn't an AWS API. You don't catch the theft; you
catch the use of the stolen creds (and from a new source IP, which is the heart of
the GuardDuty InstanceCredentialExfiltration finding — role creds showing up off the
instance they were minted for).
Exercises
Run these with the
ir-scenariolab deployed and this lab applied first (so the trail is recording). Events lag ~5–15 min into CloudWatch Logs.
A. Watch the audit log live
# [laptop, in this lab's dir] tail the trail in one terminal:
aws logs tail "$(terraform output -raw log_group)" --follow
# [another terminal] do something benign and watch it appear:
aws sts get-caller-identity
aws s3 ls
# → within a few minutes the GetCallerIdentity / ListBuckets events stream past.Lessonevery CLI command you run is an API call, and it's all being written down. Internalising that is what makes you assume — correctly — that the attacker's actions are recoverable.
B. Trip a detection on purpose
# [laptop] generate a burst of AccessDenied as a low-priv identity (the enumeration
# fingerprint). Assume the read-only IR role from ir-scenario, then try to WRITE:
cd ../ir-scenario/01-iam-foundation
creds=$(aws sts assume-role --role-arn "$(terraform output -raw ir_readonly_role_arn)" \
--role-session-name probe --query Credentials --output json)
export AWS_ACCESS_KEY_ID=$(echo "$creds" | jq -r .AccessKeyId)
export AWS_SECRET_ACCESS_KEY=$(echo "$creds" | jq -r .SecretAccessKey)
export AWS_SESSION_TOKEN=$(echo "$creds" | jq -r .SessionToken)
for i in $(seq 1 5); do aws iam create-user --user-name probe-$i 2>/dev/null; done
# → all AccessDenied (read-only can't write). That's 5 denials in one window.
unset AWS_ACCESS_KEY_ID AWS_SECRET_ACCESS_KEY AWS_SESSION_TOKEN
# [laptop] ~a few minutes later, check the alarm flipped to ALARM:
aws cloudwatch describe-alarms --alarm-name-prefix ir-lab-unauthorized-api-calls \
--query 'MetricAlarms[].{name:AlarmName,state:StateValue}'
# → "state": "ALARM" (and an SNS email if you set alert_email)Lessona low-privilege stolen identity is noisy — it can't help generating
AccessDenied as it maps out what it can do. That noise is one of your most reliable
early signals. (To reset: the alarm self-clears back to OK after a quiet period.)
C. Replay the real attack, then hunt it in Athena
# [victim] run the ir-scenario Exercise A recon on the VICTIM EC2 (so it's the
# prod-app role doing it, from the instance) — see ir-scenario/SCENARIO.md.
# aws ssm get-parameter --name /ir-lab/prod/db-password --with-decryption ...
# aws iam list-users ; aws s3 ls
# [laptop] Athena console → pick the workgroup, run "00 - create ... table" ONCE,
# then run "01 - who assumed which role" and "02 - access-denied by principal".
# To trace one identity end-to-end, edit "03 - everything done by one access key":
# - put the prod-app session's ASIA… key id in, OR
# - put the static AKIA… key id (from: aws iam list-access-keys --user-name ir-lab-prod-ci-service)Lessonthe alarms told you something happened; Athena reconstructs the whole session — every action, in order, from which IP, success or failure. That timeline is what goes in the incident report.
D. Detect the blinding
# [laptop, as admin] simulate the attacker turning logging off:
aws cloudtrail stop-logging --name "$(terraform output -raw trail_name)"
# → trips the cloudtrail-tampering alarm. Turn it back on:
aws cloudtrail start-logging --name "$(terraform output -raw trail_name)"Lessonyou catch the attacker disabling logging with the logging, because
StopLogging is itself a recorded API call. The lag between "they stopped it" and "S3
stops filling" is your window — which is why high-value orgs ship the trail to a separate
account the workload can't reach.
Interview Q&A
It gives you a who-did-what-where record of every control-plane API call — the backbone of any cloud investigation. The blind spots: it's not a data-plane log (it tells you someone called GetObject, never the object's contents, and only logs data events if you explicitly turn them on and pay), it's not retroactive (only events after the trail existed), and it's not instant (5–15 minute lag). So it's an audit log, not a wiretap, and it has to be on before the incident to be worth anything.
They're the real-time layer versus the forensic layer of the same source. CloudWatch metric filters fire alarms within minutes — good for "is a dangerous action happening now?" like root usage or someone stopping the trail. Athena runs SQL over the full S3 archive — good for "reconstruct everything this principal did last Tuesday." You alert with one and investigate with the other; in this lab they read off the one trail.
Because StopLogging is itself an API call, it lands in CloudWatch Logs and trips a tampering alarm before the trail stops feeding S3 — you detect the blinding by the act of blinding. The durable fix is architectural: ship the organization trail to a separate logging account that the workload and even most admins can't touch, so turning it off isn't in their blast radius in the first place.
accessKeyId starts with ASIA in one record and AKIA in another. So what?ASIA is a temporary STS credential — a role session, expires in hours; AKIA is a permanent IAM-user access key that never expires on its own. Seeing AKIA keys making calls from a workload host or a new IP is the signature of a stolen long-lived key — the forgotten service-account credential — and the only way to kill it is to deactivate the key, since it won't expire by itself.
No — the theft itself is invisible, because the metadata service is a link-local endpoint, not an AWS API, so hitting 169.254.169.254 leaves no CloudTrail event. What you catch is the use of the stolen role credentials afterward, especially from a source IP that isn't the instance — which is exactly the heuristic behind GuardDuty's InstanceCredentialExfiltration finding.
aws:SourceArn?Multi-region because attackers deliberately operate in regions you might not be watching, and global services like IAM and STS only log to us-east-1 — a single-region trail has gaps. Log-file validation writes signed digests so you can later prove the archive wasn't tampered with — chain of custody. And the aws:SourceArn condition on the bucket policy stops the S3 "confused deputy": without it, another account could name your bucket as their trail's destination and pollute your logs.