Walkthrough & Concepts — a complete run of the IR scenario
A read-through writeup of an actual end-to-end run of this lab: the commands, the real outputs, and the concepts that came up along the way. Use it to revise away from the keyboard. For the design of the identities and the deep-dive sections (identity chain, STS, SSRF), see SCENARIO.md; for setup/teardown mechanics see README.md.
Last verified2026-06 — resource IDs below are from one specific run and will differ on yours; the commands and lessons do not.
The arc in one picture
A. RECON — land on the popped EC2, abuse its over-privileged role
B. PERSIST — find & reuse the forgotten static-key "service account"
C. RESPOND — assume IR roles; investigate read-only, then isolate (IAM + network)
D. ACQUIRE — snapshot the disk to tamper-evident, tagged evidenceEvery step is "run a command as a specific identity and watch what it can/can't do." The identity is the lesson.
A — Recon as the compromised workload
SSH'd to the victim, then acted as the box. The metadata service handed over the instance role's credentials with no secret on disk:
[ec2-user@victim]$ TOKEN=$(curl -sX PUT http://169.254.169.254/latest/api/token \
-H "X-aws-ec2-metadata-token-ttl-seconds: 300")
[ec2-user@victim]$ curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
http://169.254.169.254/latest/meta-data/iam/security-credentials/
ir-lab-prod-app
[ec2-user@victim]$ aws sts get-caller-identity
→ "Arn": "arn:aws:sts::123456789012:assumed-role/ir-lab-prod-app/i-017ca7bf99d667de4"The over-privileged role (s3:*, ssm:GetParameter, iam:Get*/List* on *) then
read a production secret, enumerated IAM, and listed the state bucket:
[ec2-user@victim]$ aws ssm get-parameter --name /ir-lab/prod/db-password \
--with-decryption --query Parameter.Value --output text
L4b-Pl@ceholder-NotReal
[ec2-user@victim]$ aws iam list-users → admin + ir-lab-prod-ci-service
[ec2-user@victim]$ aws s3 ls → ir-lab-tfstate-123456789012Self-enumeration worked because the role can read its own IAM — a gift to the attacker:
[ec2-user@victim]$ aws iam get-role-policy --role-name ir-lab-prod-app --policy-name overbroad
→ Action [sts:GetCallerIdentity, ssm:GetParameter*, s3:*, iam:Get*/List*, ec2:Describe*], Resource "*"LessonOne popped box + an over-privileged workload role = account-wide read,
including the Terraform state bucket, where every other secret sits in plaintext.
A web app needed one parameter and one S3 prefix; it got s3:* + iam:Get*/List*.
The "trap" we hit: running the recon on the laptop (as
admin) "works" and proves nothing. The finding only counts run on the victim, asprod-app.
B — The forgotten static-key service account
[laptop]$ aws iam get-access-key-last-used --access-key-id $(terraform output -raw prod_ci_access_key_id)
[laptop]$ AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... aws sts get-caller-identity
→ "Arn": ".../user/ir-lab-prod-ci-service"LessonLong-lived AKIA… keys with PowerUserAccess are durable persistence —
they survive instance termination and password resets. The only kill is
aws iam update-access-key --status Inactive.
C — Respond: investigate read-only, then isolate
Assumed the read-only role, confirmed it can look but not touch:
[laptop, 01-iam-foundation]$ creds=$(aws sts assume-role \
--role-arn $(terraform output -raw ir_readonly_role_arn) --role-session-name investigate \
--query Credentials --output json)
# export AccessKeyId / SecretAccessKey / SessionToken ...
$ aws ec2 describe-instances ... → works (read)
$ aws ec2 stop-instances --instance-ids i-017ca7bf99d667de4
→ AccessDenied ← separation of duties: investigators don't change stateSwitched to the forensic role and isolated the box at two independent layers:
# IAM layer — swap to the deny-all quarantine role:
$ ASSOC=$(aws ec2 describe-iam-instance-profile-associations \
--filters Name=instance-id,Values=i-017ca7bf99d667de4 \
--query 'IamInstanceProfileAssociations[0].AssociationId' --output text)
$ aws ec2 replace-iam-instance-profile-association \
--association-id $ASSOC --iam-instance-profile Name=ir-lab-quarantine
# Network layer — move to the no-rules quarantine SG:
$ aws ec2 modify-instance-attribute --instance-id i-017ca7bf99d667de4 \
--groups sg-0a08b15c15b2bf15eA telling deny: even the forensic role can't StopInstances:
UnauthorizedOperation ... ir-lab-ir-forensic/acquire is not authorized to perform: ec2:StopInstancesThat's deliberate — you never power off a box during acquisition (it wipes RAM and can trip dead-man switches). The forensic role can isolate and snapshot, not destroy.
The two-layer model (the core takeaway)
| Action | Layer | Blocks SSH? | Effect |
|---|---|---|---|
swap to deny-all quarantine role | IAM | No | kills the box's AWS API power; you keep your shell to investigate |
move to quarantine SG (no rules) | Network | Yes | cuts all traffic incl. SSH; box goes fully dark |
After the SG swap, SSH hangs and times out — confirmed live. This is why the lab uses SSH, not Session Manager: SSH is OS-level and survives the IAM deny, so the responder keeps access right up until they choose to cut the network. Sequencing matters: triage volatile data (or keep a forensic-only port-22 rule) before the network cut, or you lock yourself out.
D — Acquire: snapshot to evidence
[forensic role]$ aws ec2 create-snapshot --volume-id vol-05db8cec60e78e03b \
--description "IR acquisition: victim i-017ca7bf99d667de4 root disk" \
--tag-specifications 'ResourceType=snapshot,Tags=[{Key=Case,Value=ir-lab-2026-06-17},
{Key=SourceInstance,Value=i-017ca7bf99d667de4},{Key=AcquiredBy,Value=ir-forensic}]'
→ snap-07cbb006e85b41126 (later: State=completed, 100%)
$ aws ec2 create-volume --availability-zone us-east-1a --snapshot-id snap-07cbb006e85b41126
→ vol-046c840e49f59276eRemaining (Option B, not run here): launch a clean forensic instance in us-east-1a
(admin — the forensic role lacks RunInstances), attach the volume, and
mount -o ro,noexec,nosuid,nodev,norecovery. Full steps:
../../../../forensics/digital-forensics.md.
LessonSnapshot = immutable, timestamped, tagged for chain of custody. You examine a copy, never the original, and mount it read-only so investigating can't alter evidence or execute anything off the suspect disk.
Concepts that came up (revision notes)
How an EC2 gets its identity
Instance → instance profile (the plug) → IAM role (the identity) → policies
(the permissions). No keys on disk: the SDK pulls temporary credentials from IMDS
on demand. Audit a role with list-role-policies / get-role-policy /
list-attached-role-policies, or simulate-principal-policy. Deep dive in SCENARIO.md.
IMDSv1 vs IMDSv2
Both serve the same link-local 169.254.169.254, reachable only from on the
instance — neither is routable off-box. The difference is anti-theft ceremony, not
locality:
| IMDSv1 | IMDSv2 | |
|---|---|---|
| Get creds | single GET | PUT for a session token first, then GET with the token header |
| Anti-SSRF | none | a plain "fetch this URL" SSRF can't do the PUT / set the header |
| Response TTL | normal | hop limit 1 — response can't be proxied off-box |
This victim set http_tokens = "required", so a tokenless v1 GET is refused with
HTTP 401, even from on the box. If the instance allowed v1 (optional), a plain
GET would work — but still only from on the instance.
STS — the temporary-credential minter
AssumeRole (and AssumeRoleWithWebIdentity for OIDC/IRSA) returns a credential
triple: AccessKeyId (ASIA…, vs AKIA… for permanent user keys) +
SecretAccessKey + SessionToken, with an expiry (default 1h — which is why a stale
session made describe-volumes come back empty mid-run). Two gates: the role's
trust policy must allow you, and you need sts:AssumeRole. IMDS runs an
AssumeRole for the instance under the hood — hence the assumed-role/…/i-017… ARN.
SSRF — how creds get stolen without a shell
Trick the server into fetching a URL for you; point it at IMDS
(http://169.254.169.254/latest/meta-data/iam/security-credentials/<role>) and it
returns the role's creds in the HTTP response. This is the 2019 Capital One pattern.
IMDSv2's PUT+token+hop-limit blunts the classic one-shot GET. Full example in
SCENARIO.md.
Where EBS lives (came up at the create-volume step)
An EBS volume is network-attached block storage living in one Availability
Zone (replicated across devices in that AZ) — so it only attaches to instances in
the same AZ. A snapshot lives in AWS-managed S3, region-wide, block-level and
incremental — so it's portable: copy-snapshot cross-region/cross-account to ship
evidence. "snapshot → create-volume in an AZ" = pull the regional evidence master down
into a live disk wherever you need to examine it.
Teardown (what actually has to be deleted)
terraform destroy only removes Terraform-managed resources. The snapshot and the
evidence volume were created by hand and must be deleted separately. And destroy
needs admin — drop any assumed-role creds first.
# 0. back to admin (forensic/readonly creds can't destroy or delete)
unset AWS_ACCESS_KEY_ID AWS_SECRET_ACCESS_KEY AWS_SESSION_TOKEN
aws sts get-caller-identity # → user/admin
# 1. manual forensic artifacts (Terraform doesn't track these)
aws ec2 delete-volume --volume-id vol-046c840e49f59276e
aws ec2 delete-snapshot --snapshot-id snap-07cbb006e85b41126
# 2-4. Terraform stacks, REVERSE order; state bucket LAST
cd 02-billable && terraform destroy && rm -f ir-lab-key.pem # RDS takes minutes
cd ../01-iam-foundation && terraform destroy
cd ../00-bootstrap && terraform destroy # the state bucket- Step 2 shows drift — the destroy plan lists the victim with
iam_instance_profile = "ir-lab-quarantine"andsecurity_groups = ["ir-lab-quarantine-sg"](your out-of-band isolation). Terraform terminates it anyway. - The state bucket has versioning and
force_destroy = true, so step 4 tears it down cleanly even though it still holds state objects + versions. - Confirm clean in the EC2 / RDS / IAM consoles and the Budgets dashboard.
Verified run (2026-06-17)manual volume + snapshot deleted, then
02-billable → 13 destroyed (RDS ~2 min), 01-iam-foundation → 18 destroyed,
00-bootstrap → 4 destroyed. Account clean, billing meter off.
Five things to be able to say in an interview
instance profile → role → policy, delivered as short-lived creds via IMDS; nothing on disk. Over-privileged workload role = account-wide blast radius.
isn't "more local" than v1 — it's anti-SSRF (PUT + token header + hop
limit). http_tokens=required refuses v1.
vends temporary ASIA… creds (triple + expiry) gated by the role's trust
policy and the caller's sts:AssumeRole.
IAM (kill cloud power, keep your shell) and network (cut everything). Snapshot/triage before you cut the network.
immutable, tagged snapshot; examine a read-only copy, never the original; never power the box off (you'd lose volatile memory).