Security Notes
Hands-on Labs

Scenario: Turning On the Lights — Logging, Alerting, and Hunting in AWS

The ir-scenario lab is the break-in: an over-privileged role, a forgotten static key, a victim EC2, an SSRF that steals credentials off the metadata service. The whole time, a defender with no logging sees none of it. This lab is that defender catching up — and it's the difference between "we think someone got in" and "here is every call they made, in order, with timestamps."

10 min read 6 sections 6 model answers verified 2026-06

Last verified2026-06

Read this once, then run the exercises with the ir-scenario lab deployed alongside.


First: what CloudTrail actually is (and isn't)

Plain EnglishCloudTrail is the answering-machine tape for the AWS API. Almost everything in AWS — the console, the CLI, an SDK, one service calling another — is an API call to the AWS control plane. CloudTrail writes down who made each call, when, from what IP, with what parameters, and whether it succeeded. That's it. That single fact is the backbone of nearly every cloud investigation.

What it records (an event) looks like this — the fields you'll actually grep for:

json
{
  "eventTime":      "2026-06-17T10:22:31Z",
  "eventName":      "GetParameter",                  // the API action
  "eventSource":    "ssm.amazonaws.com",             // the service
  "awsRegion":      "us-east-1",
  "sourceIPAddress":"203.0.113.9",                   // where it came from
  "userIdentity": {                                   // WHO — the most important block
     "type": "AssumedRole",
     "arn":  "arn:aws:sts::1234:assumed-role/ir-lab-prod-app/i-0abc...",
     "accessKeyId": "ASIA...",                        // ASIA = temporary (STS) creds
     "sessionContext": { "attributes": { "mfaAuthenticated": "false" } }
  },
  "errorCode":      "AccessDenied"                    // present ONLY when the call failed
}

What it is NOT

Not data-plane.

"Who called s3:GetObject" is a management/data event CloudTrail can log; "what bytes were inside the object" is never in CloudTrail. Reading a row out of a database, the contents of a file — that's the data plane, invisible here.

Not retroactive.

It records only what happens after the trail exists. This is why you stand logging up on day zero, not after the incident.

Not instant.

Events reach CloudWatch Logs in ~5–15 min and the S3 archive a little later. Good enough to alert on; not a real-time packet capture.

Memory hook

One-linerCloudTrail answers "who did what, where, and did it work?" for every AWS API call — it's the control-plane audit log, not a data-plane wiretap.


The three layers this lab builds

One trail feeds three different jobs. Knowing which layer answers which question is the whole point — and a classic interview question in itself.

                         ┌─────────────────────────────────────────────┐
   every AWS API call ──►│            CloudTrail (one trail)            │
                         └───────┬─────────────────────────┬───────────┘
                                 │                          │
                 CloudWatch Logs │                          │ S3 (durable archive)
                                 ▼                          ▼
                  metric filters + alarms            Athena (SQL over history)
                  = "is it happening NOW?"           = "reconstruct what happened"
                  (real-time, seconds-to-minutes)    (forensics / threat hunting)

   ── and in parallel ──►  GuardDuty  = "does AWS think this matches a known attack?"
                           (managed; reads CloudTrail + VPC Flow + DNS for you)
LayerQuestion it answersLatencyIn this lab
CloudWatch metric filters + alarms"Is a dangerous action happening right now?"minutes5 detections → SNS
Athena over S3"Show me everything X did between these times"query-timesaved hunt queries
GuardDuty"Does this look like a known attack pattern?"~15 minone detector

The detections (real-time layer)

These are written as detections-as-code in detections.tf: a CloudWatch Logs metric filter (a pattern that adds 1 to a counter on each matching event) wired to an alarm that fires when the count goes above zero and publishes to SNS. They map to the CIS AWS Benchmark monitoring controls — the canonical "you have CloudTrail, now what do you alert on?" list.

DetectionWhat it catchesWhy it matters
root-account-usageAny action by the root identityRoot should be sealed away with MFA. Any use is an incident.
unauthorized-api-callsBursts of AccessDenied / UnauthorizedOperationThe fingerprint of enumeration — an attacker probing what their stolen identity can do.
console-login-no-mfaAn IAM user signing in without MFAEither a misconfigured user or a password-only credential being abused.
iam-policy-changesCreate/attach/detach/delete of IAM policiesPrivilege escalation and persistence happen by editing IAM.
cloudtrail-tamperingStopLogging / DeleteTrail / UpdateTrail / PutEventSelectorsThe adversary's first move is often to blind the logging. Alerting on it is your tripwire.

The deep point — anti-tamperThe single most important alarm is the last one. A smart attacker who lands admin will try to turn CloudTrail off so the rest of their activity goes unrecorded. But StopLogging is itself an API call — so it lands in CloudWatch Logs and trips the alarm before the trail stops feeding S3. You detect the blinding by the act of blinding. (Defense in depth: in a real org the trail lives in a separate logging account the workload identities can't touch at all.)


How an attack appears in CloudTrail (tie-back to ir-scenario)

Walk the ir-scenario attack chain and name the event each step leaves behind. This is exactly the "trace the intrusion" exercise an IR interview runs you through.

ir-scenario step                                  CloudTrail event you'd hunt for
─────────────────────────────────────────────────────────────────────────────────────
1. SSRF makes the app hit IMDS, steal creds       (none — IMDS is link-local, not an API call)
2. attacker uses the role's temp creds            calls with userIdentity.accessKeyId = ASIA…,
                                                   arn = assumed-role/ir-lab-prod-app/i-…
3. recon: aws iam list-users, s3 ls                ListUsers, ListBuckets  (often the FIRST tells)
4. read the secret: ssm get-parameter             GetParameter on /ir-lab/prod/db-password
5. read state bucket: s3 cp …tfstate              GetObject (only if S3 DATA events are on)
6. use the static CI key (persistence)             calls with accessKeyId = AKIA… (AKIA = permanent!)
7. attacker disables logging                       StopLogging / DeleteTrail  → tampering alarm

The credential telluserIdentity.accessKeyId starting ASIA is a temporary STS credential (a role); AKIA is a permanent IAM-user key. Seeing AKIA make API calls from a workload — or from a new IP — is the signature of a stolen static key, exactly the ir-lab-prod-ci-service credential from the other lab. The prefix tells you in one glance whether you're chasing a role session or a leaked long-lived key.

A subtle catch from step 1: the credential theft itself is invisible to CloudTrail — IMDS lives at a link-local address and isn't an AWS API. You don't catch the theft; you catch the use of the stolen creds (and from a new source IP, which is the heart of the GuardDuty InstanceCredentialExfiltration finding — role creds showing up off the instance they were minted for).


Exercises

Run these with the ir-scenario lab deployed and this lab applied first (so the trail is recording). Events lag ~5–15 min into CloudWatch Logs.

A. Watch the audit log live

bash
# [laptop, in this lab's dir]  tail the trail in one terminal:
aws logs tail "$(terraform output -raw log_group)" --follow

# [another terminal]  do something benign and watch it appear:
aws sts get-caller-identity
aws s3 ls
# → within a few minutes the GetCallerIdentity / ListBuckets events stream past.

Lessonevery CLI command you run is an API call, and it's all being written down. Internalising that is what makes you assume — correctly — that the attacker's actions are recoverable.

B. Trip a detection on purpose

bash
# [laptop]  generate a burst of AccessDenied as a low-priv identity (the enumeration
# fingerprint). Assume the read-only IR role from ir-scenario, then try to WRITE:
cd ../ir-scenario/01-iam-foundation
creds=$(aws sts assume-role --role-arn "$(terraform output -raw ir_readonly_role_arn)" \
  --role-session-name probe --query Credentials --output json)
export AWS_ACCESS_KEY_ID=$(echo "$creds" | jq -r .AccessKeyId)
export AWS_SECRET_ACCESS_KEY=$(echo "$creds" | jq -r .SecretAccessKey)
export AWS_SESSION_TOKEN=$(echo "$creds" | jq -r .SessionToken)

for i in $(seq 1 5); do aws iam create-user --user-name probe-$i 2>/dev/null; done
#   → all AccessDenied (read-only can't write). That's 5 denials in one window.
unset AWS_ACCESS_KEY_ID AWS_SECRET_ACCESS_KEY AWS_SESSION_TOKEN

# [laptop]  ~a few minutes later, check the alarm flipped to ALARM:
aws cloudwatch describe-alarms --alarm-name-prefix ir-lab-unauthorized-api-calls \
  --query 'MetricAlarms[].{name:AlarmName,state:StateValue}'
#   → "state": "ALARM"   (and an SNS email if you set alert_email)

Lessona low-privilege stolen identity is noisy — it can't help generating AccessDenied as it maps out what it can do. That noise is one of your most reliable early signals. (To reset: the alarm self-clears back to OK after a quiet period.)

C. Replay the real attack, then hunt it in Athena

bash
# [victim]  run the ir-scenario Exercise A recon on the VICTIM EC2 (so it's the
# prod-app role doing it, from the instance) — see ir-scenario/SCENARIO.md.
#   aws ssm get-parameter --name /ir-lab/prod/db-password --with-decryption ...
#   aws iam list-users ; aws s3 ls

# [laptop]  Athena console → pick the workgroup, run "00 - create ... table" ONCE,
# then run "01 - who assumed which role" and "02 - access-denied by principal".
# To trace one identity end-to-end, edit "03 - everything done by one access key":
#   - put the prod-app session's ASIA… key id in, OR
#   - put the static AKIA… key id (from: aws iam list-access-keys --user-name ir-lab-prod-ci-service)

Lessonthe alarms told you something happened; Athena reconstructs the whole session — every action, in order, from which IP, success or failure. That timeline is what goes in the incident report.

D. Detect the blinding

bash
# [laptop, as admin]  simulate the attacker turning logging off:
aws cloudtrail stop-logging --name "$(terraform output -raw trail_name)"
# → trips the cloudtrail-tampering alarm. Turn it back on:
aws cloudtrail start-logging --name "$(terraform output -raw trail_name)"

Lessonyou catch the attacker disabling logging with the logging, because StopLogging is itself a recorded API call. The lag between "they stopped it" and "S3 stops filling" is your window — which is why high-value orgs ship the trail to a separate account the workload can't reach.


Interview Q&A

Q
You've been told "we have CloudTrail enabled." What does that actually buy you, and what are its three big blind spots?
Model answer

It gives you a who-did-what-where record of every control-plane API call — the backbone of any cloud investigation. The blind spots: it's not a data-plane log (it tells you someone called GetObject, never the object's contents, and only logs data events if you explicitly turn them on and pay), it's not retroactive (only events after the trail existed), and it's not instant (5–15 minute lag). So it's an audit log, not a wiretap, and it has to be on before the incident to be worth anything.

Q
What's the difference between using CloudWatch alarms versus Athena over the same CloudTrail data?
Model answer

They're the real-time layer versus the forensic layer of the same source. CloudWatch metric filters fire alarms within minutes — good for "is a dangerous action happening now?" like root usage or someone stopping the trail. Athena runs SQL over the full S3 archive — good for "reconstruct everything this principal did last Tuesday." You alert with one and investigate with the other; in this lab they read off the one trail.

Q
A smart attacker with admin will try to disable CloudTrail. How do you still catch them?
Model answer

Because StopLogging is itself an API call, it lands in CloudWatch Logs and trips a tampering alarm before the trail stops feeding S3 — you detect the blinding by the act of blinding. The durable fix is architectural: ship the organization trail to a separate logging account that the workload and even most admins can't touch, so turning it off isn't in their blast radius in the first place.

Q
In a CloudTrail event, the accessKeyId starts with ASIA in one record and AKIA in another. So what?
Model answer

ASIA is a temporary STS credential — a role session, expires in hours; AKIA is a permanent IAM-user access key that never expires on its own. Seeing AKIA keys making calls from a workload host or a new IP is the signature of a stolen long-lived key — the forgotten service-account credential — and the only way to kill it is to deactivate the key, since it won't expire by itself.

Q
The metadata-service credential theft (the SSRF in the sibling lab) — will CloudTrail show it?
Model answer

No — the theft itself is invisible, because the metadata service is a link-local endpoint, not an AWS API, so hitting 169.254.169.254 leaves no CloudTrail event. What you catch is the use of the stolen role credentials afterward, especially from a source IP that isn't the instance — which is exactly the heuristic behind GuardDuty's InstanceCredentialExfiltration finding.

Q
Why a multi-region trail with log-file validation, and why scope the S3 bucket policy with aws:SourceArn?
Model answer

Multi-region because attackers deliberately operate in regions you might not be watching, and global services like IAM and STS only log to us-east-1 — a single-region trail has gaps. Log-file validation writes signed digests so you can later prove the archive wasn't tampered with — chain of custody. And the aws:SourceArn condition on the bucket policy stops the S3 "confused deputy": without it, another account could name your bucket as their trail's destination and pollute your logs.