Security Notes
Hands-on Labs

Walkthrough & Concepts — a complete run of the IR scenario

A read-through writeup of an actual end-to-end run of this lab: the commands, the real outputs, and the concepts that came up along the way. Use it to revise away from the keyboard. For the design of the identities and the deep-dive sections (identity chain, STS, SSRF), see SCENARIO.md; for setup/teardown mechanics see README.md.

7 min read 8 sections verified 2026-06

Last verified2026-06 — resource IDs below are from one specific run and will differ on yours; the commands and lessons do not.


The arc in one picture

A. RECON   — land on the popped EC2, abuse its over-privileged role
B. PERSIST — find & reuse the forgotten static-key "service account"
C. RESPOND — assume IR roles; investigate read-only, then isolate (IAM + network)
D. ACQUIRE — snapshot the disk to tamper-evident, tagged evidence

Every step is "run a command as a specific identity and watch what it can/can't do." The identity is the lesson.


A — Recon as the compromised workload

SSH'd to the victim, then acted as the box. The metadata service handed over the instance role's credentials with no secret on disk:

bash
[ec2-user@victim]$ TOKEN=$(curl -sX PUT http://169.254.169.254/latest/api/token \
                     -H "X-aws-ec2-metadata-token-ttl-seconds: 300")
[ec2-user@victim]$ curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
                     http://169.254.169.254/latest/meta-data/iam/security-credentials/
ir-lab-prod-app
[ec2-user@victim]$ aws sts get-caller-identity
  → "Arn": "arn:aws:sts::123456789012:assumed-role/ir-lab-prod-app/i-017ca7bf99d667de4"

The over-privileged role (s3:*, ssm:GetParameter, iam:Get*/List* on *) then read a production secret, enumerated IAM, and listed the state bucket:

bash
[ec2-user@victim]$ aws ssm get-parameter --name /ir-lab/prod/db-password \
                     --with-decryption --query Parameter.Value --output text
L4b-Pl@ceholder-NotReal
[ec2-user@victim]$ aws iam list-users        → admin + ir-lab-prod-ci-service
[ec2-user@victim]$ aws s3 ls                 → ir-lab-tfstate-123456789012

Self-enumeration worked because the role can read its own IAM — a gift to the attacker:

bash
[ec2-user@victim]$ aws iam get-role-policy --role-name ir-lab-prod-app --policy-name overbroad
  → Action [sts:GetCallerIdentity, ssm:GetParameter*, s3:*, iam:Get*/List*, ec2:Describe*], Resource "*"

LessonOne popped box + an over-privileged workload role = account-wide read, including the Terraform state bucket, where every other secret sits in plaintext. A web app needed one parameter and one S3 prefix; it got s3:* + iam:Get*/List*.

The "trap" we hit: running the recon on the laptop (as admin) "works" and proves nothing. The finding only counts run on the victim, as prod-app.


B — The forgotten static-key service account

bash
[laptop]$ aws iam get-access-key-last-used --access-key-id $(terraform output -raw prod_ci_access_key_id)
[laptop]$ AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... aws sts get-caller-identity
  → "Arn": ".../user/ir-lab-prod-ci-service"

LessonLong-lived AKIA… keys with PowerUserAccess are durable persistence — they survive instance termination and password resets. The only kill is aws iam update-access-key --status Inactive.


C — Respond: investigate read-only, then isolate

Assumed the read-only role, confirmed it can look but not touch:

bash
[laptop, 01-iam-foundation]$ creds=$(aws sts assume-role \
    --role-arn $(terraform output -raw ir_readonly_role_arn) --role-session-name investigate \
    --query Credentials --output json)
# export AccessKeyId / SecretAccessKey / SessionToken ...
$ aws ec2 describe-instances ...          → works (read)
$ aws ec2 stop-instances --instance-ids i-017ca7bf99d667de4
  → AccessDenied                          ← separation of duties: investigators don't change state

Switched to the forensic role and isolated the box at two independent layers:

bash
# IAM layer — swap to the deny-all quarantine role:
$ ASSOC=$(aws ec2 describe-iam-instance-profile-associations \
    --filters Name=instance-id,Values=i-017ca7bf99d667de4 \
    --query 'IamInstanceProfileAssociations[0].AssociationId' --output text)
$ aws ec2 replace-iam-instance-profile-association \
    --association-id $ASSOC --iam-instance-profile Name=ir-lab-quarantine

# Network layer — move to the no-rules quarantine SG:
$ aws ec2 modify-instance-attribute --instance-id i-017ca7bf99d667de4 \
    --groups sg-0a08b15c15b2bf15e

A telling deny: even the forensic role can't StopInstances:

UnauthorizedOperation ... ir-lab-ir-forensic/acquire is not authorized to perform: ec2:StopInstances

That's deliberate — you never power off a box during acquisition (it wipes RAM and can trip dead-man switches). The forensic role can isolate and snapshot, not destroy.

The two-layer model (the core takeaway)

ActionLayerBlocks SSH?Effect
swap to deny-all quarantine roleIAMNokills the box's AWS API power; you keep your shell to investigate
move to quarantine SG (no rules)NetworkYescuts all traffic incl. SSH; box goes fully dark

After the SG swap, SSH hangs and times out — confirmed live. This is why the lab uses SSH, not Session Manager: SSH is OS-level and survives the IAM deny, so the responder keeps access right up until they choose to cut the network. Sequencing matters: triage volatile data (or keep a forensic-only port-22 rule) before the network cut, or you lock yourself out.


D — Acquire: snapshot to evidence

bash
[forensic role]$ aws ec2 create-snapshot --volume-id vol-05db8cec60e78e03b \
    --description "IR acquisition: victim i-017ca7bf99d667de4 root disk" \
    --tag-specifications 'ResourceType=snapshot,Tags=[{Key=Case,Value=ir-lab-2026-06-17},
        {Key=SourceInstance,Value=i-017ca7bf99d667de4},{Key=AcquiredBy,Value=ir-forensic}]'
  → snap-07cbb006e85b41126   (later: State=completed, 100%)

$ aws ec2 create-volume --availability-zone us-east-1a --snapshot-id snap-07cbb006e85b41126
  → vol-046c840e49f59276e

Remaining (Option B, not run here): launch a clean forensic instance in us-east-1a (admin — the forensic role lacks RunInstances), attach the volume, and mount -o ro,noexec,nosuid,nodev,norecovery. Full steps: ../../../../forensics/digital-forensics.md.

LessonSnapshot = immutable, timestamped, tagged for chain of custody. You examine a copy, never the original, and mount it read-only so investigating can't alter evidence or execute anything off the suspect disk.


Concepts that came up (revision notes)

How an EC2 gets its identity

Instance → instance profile (the plug) → IAM role (the identity) → policies (the permissions). No keys on disk: the SDK pulls temporary credentials from IMDS on demand. Audit a role with list-role-policies / get-role-policy / list-attached-role-policies, or simulate-principal-policy. Deep dive in SCENARIO.md.

IMDSv1 vs IMDSv2

Both serve the same link-local 169.254.169.254, reachable only from on the instance — neither is routable off-box. The difference is anti-theft ceremony, not locality:

IMDSv1IMDSv2
Get credssingle GETPUT for a session token first, then GET with the token header
Anti-SSRFnonea plain "fetch this URL" SSRF can't do the PUT / set the header
Response TTLnormalhop limit 1 — response can't be proxied off-box

This victim set http_tokens = "required", so a tokenless v1 GET is refused with HTTP 401, even from on the box. If the instance allowed v1 (optional), a plain GET would work — but still only from on the instance.

STS — the temporary-credential minter

AssumeRole (and AssumeRoleWithWebIdentity for OIDC/IRSA) returns a credential triple: AccessKeyId (ASIA…, vs AKIA… for permanent user keys) + SecretAccessKey + SessionToken, with an expiry (default 1h — which is why a stale session made describe-volumes come back empty mid-run). Two gates: the role's trust policy must allow you, and you need sts:AssumeRole. IMDS runs an AssumeRole for the instance under the hood — hence the assumed-role/…/i-017… ARN.

SSRF — how creds get stolen without a shell

Trick the server into fetching a URL for you; point it at IMDS (http://169.254.169.254/latest/meta-data/iam/security-credentials/<role>) and it returns the role's creds in the HTTP response. This is the 2019 Capital One pattern. IMDSv2's PUT+token+hop-limit blunts the classic one-shot GET. Full example in SCENARIO.md.

Where EBS lives (came up at the create-volume step)

An EBS volume is network-attached block storage living in one Availability Zone (replicated across devices in that AZ) — so it only attaches to instances in the same AZ. A snapshot lives in AWS-managed S3, region-wide, block-level and incremental — so it's portable: copy-snapshot cross-region/cross-account to ship evidence. "snapshot → create-volume in an AZ" = pull the regional evidence master down into a live disk wherever you need to examine it.


Teardown (what actually has to be deleted)

terraform destroy only removes Terraform-managed resources. The snapshot and the evidence volume were created by hand and must be deleted separately. And destroy needs admin — drop any assumed-role creds first.

bash
# 0. back to admin (forensic/readonly creds can't destroy or delete)
unset AWS_ACCESS_KEY_ID AWS_SECRET_ACCESS_KEY AWS_SESSION_TOKEN
aws sts get-caller-identity            # → user/admin

# 1. manual forensic artifacts (Terraform doesn't track these)
aws ec2 delete-volume   --volume-id   vol-046c840e49f59276e
aws ec2 delete-snapshot --snapshot-id snap-07cbb006e85b41126

# 2-4. Terraform stacks, REVERSE order; state bucket LAST
cd 02-billable        && terraform destroy && rm -f ir-lab-key.pem   # RDS takes minutes
cd ../01-iam-foundation && terraform destroy
cd ../00-bootstrap      && terraform destroy                          # the state bucket
  • Step 2 shows drift — the destroy plan lists the victim with iam_instance_profile = "ir-lab-quarantine" and security_groups = ["ir-lab-quarantine-sg"] (your out-of-band isolation). Terraform terminates it anyway.
  • The state bucket has versioning and force_destroy = true, so step 4 tears it down cleanly even though it still holds state objects + versions.
  • Confirm clean in the EC2 / RDS / IAM consoles and the Budgets dashboard.

Verified run (2026-06-17)manual volume + snapshot deleted, then 02-billable → 13 destroyed (RDS ~2 min), 01-iam-foundation → 18 destroyed, 00-bootstrap → 4 destroyed. Account clean, billing meter off.


Five things to be able to say in an interview

EC2 identity

instance profile → role → policy, delivered as short-lived creds via IMDS; nothing on disk. Over-privileged workload role = account-wide blast radius.

IMDSv2

isn't "more local" than v1 — it's anti-SSRF (PUT + token header + hop limit). http_tokens=required refuses v1.

STS

vends temporary ASIA… creds (triple + expiry) gated by the role's trust policy and the caller's sts:AssumeRole.

Containment is two layers

IAM (kill cloud power, keep your shell) and network (cut everything). Snapshot/triage before you cut the network.

Acquisition

immutable, tagged snapshot; examine a read-only copy, never the original; never power the box off (you'd lose volatile memory).