Scenario: Identities, the "Default Service Account" Trap, and the Exercises
This is the story the lab tells. Read it once, then run the exercises.
First: "default service accounts" — AWS vs GCP vs Kubernetes
Your instinct ("there are default service accounts that, if you don't lock them down, an attacker can just use") is correct — but it's mostly a GCP / Kubernetes thing, not AWS. Worth getting straight, because the mental model differs per cloud:
- GCP really does this: every project gets a default Compute Engine service account that is granted the broad Editor role and is attached to VMs by default. Compromise a VM that's using it → you inherit Editor on the whole project. That's the "default SA you forgot to lock down" you're picturing.
- Kubernetes does it too: every namespace has a
defaultServiceAccount, and its token is auto-mounted into pods unless you opt out — so a popped pod often has an API token sitting at/var/run/secrets/...for free. - AWS has no such auto-attached "default service account." An EC2 instance has
no identity at all unless you attach an instance profile. So in AWS the
equivalent risks show up differently, and this lab emulates the two that matter:Over-privileged roles attached to compute
the
prod-approle here can read every secret and enumerate IAM. That's the AWS version of "the VM's SA can do too much."Long-lived IAM user access keys that were created once and never rotatedthe
prod-ci-serviceuser here. Static keys are the credential attackers prize: they don't expire, they're often over-privileged, and nobody notices them. This is the closest AWS analog to "a default credential left active."
Memory hookOne-linerGCP/K8s hand you an over-privileged identity by default; in AWS you hand it to yourself by attaching a fat role or leaving static keys lying around.
The identities this lab creates
Production (what an attacker lands in) — risky on purpose
| Identity | Type | Why it's risky (the lesson) |
|---|---|---|
…-prod-app | EC2 role (instance profile) | Over-privileged workload role: s3:*, ssm:GetParameter, iam:List*/Get*. A compromised instance inherits all of it. |
…-prod-ci-service | IAM user + static access key | Long-lived "service account" keys with PowerUserAccess — the forgotten, never-rotated credential. |
/ir-lab/prod/db-password | SSM Parameter Store (SecureString) | The juicy target the over-privileged role can read — proves "blast radius." (SSM instead of Secrets Manager so Section 1 stays free.) |
Incident response (what you respond WITH)
| Identity | Type | Purpose |
|---|---|---|
…-ir-readonly | role (ReadOnlyAccess) | Investigator — read everything, change nothing. |
…-ir-forensic | role (scoped) | Acquisition — snapshot, copy, isolate (modify SG), swap instance profile, SSM. Least-privilege for the job. |
…-break-glass | role (AdministratorAccess) | Emergency admin — normally unused; in real life you alarm on any use. |
…-quarantine | instance profile (deny-all) | The isolation identity you swap onto a compromised instance. |
All three IR roles are assumable by principals in your account:
aws sts assume-role --role-arn <arn> --role-session-name ir.
Where to run
terraform output: identity outputs (role ARNs, service-account keys,prod_param_name) come from01-iam-foundation/; the instance/Lambda/RDS outputs (ssh_command, etc.) come from02-billable/.cdinto the right module first. Exercises A–C need Section 2 applied (you need a running victim).
How an EC2 instance gets its identity (and how to audit it)
The single most important AWS concept in this lab: a process on an EC2 instance never holds AWS keys on disk — it borrows a role through the metadata service.
EC2 instance i-017ca7bf...
└─ attached: Instance Profile "ir-lab-prod-app" (a thin wrapper/plug)
└─ contains exactly ONE: IAM Role "ir-lab-prod-app" (the identity)
└─ policies on the role = what it's ALLOWED to do:
• inline policy "overbroad" → s3:* , ssm:GetParameter , iam:Get*/List* , ...- The instance profile is the plug; the role is the identity; the policies are the permissions. An instance profile holds exactly one role.
- When any SDK/CLI on the box makes an AWS call, it fetches temporary credentials
for that role from the Instance Metadata Service (IMDS) at
169.254.169.254. Those creds auto-rotate every few hours. Nothing is stored on disk — which is why, on a popped box, you can't "find the keys": the box's power is exactly the attached role's policies, delivered live.
Auditing "what can this box do?" — three ways
1. From the box itself (attacker self-enumeration). This role has iam:Get*/List*,
so it can read its own permissions — a gift to an attacker:
ROLE=ir-lab-prod-app
aws iam list-role-policies --role-name $ROLE
# → "PolicyNames": ["overbroad"]
aws iam get-role-policy --role-name $ROLE --policy-name overbroad
# → the statement: Action [s3:*, ssm:GetParameter*, iam:Get*/List*, ec2:Describe*], Resource "*"
aws iam list-attached-role-policies --role-name $ROLE
# → [] (no AWS-managed policies here; the whole grant is the inline "overbroad")
aws iam get-instance-profile --instance-profile-name ir-lab-prod-app
# → shows the profile → role link, and the role's trust policy (who may assume it:
# Principal Service ec2.amazonaws.com — i.e. "any EC2 instance this is attached to")2. From your laptop (the IR/investigator view). Same reads as admin or the
ir-readonly role — plus the policy simulator, which answers "allowed or denied?"
for specific actions without making a real call:
aws iam simulate-principal-policy \
--policy-source-arn arn:aws:iam::<acct>:role/ir-lab-prod-app \
--action-names s3:DeleteObject ssm:GetParameter iam:CreateUser \
--query 'EvaluationResults[].{action:EvalActionName,decision:EvalDecision}'
# → s3:DeleteObject allowed, ssm:GetParameter allowed, iam:CreateUser implicitDeny3. From the Terraform source (ground truth). 01-iam-foundation/iam.tf, the
prod_app_overbroad policy document. Or Console → IAM → Roles → ir-lab-prod-app.
The findinga web app needs maybe one parameter and one S3 prefix. This role has
s3:*+iam:Get*/List*on*. So one compromised box reads any bucket (including the secret-laden Terraform state), every secret in SSM, and the whole IAM layout. Over-privileged workload role = blast radius.
STS: the service that mints temporary credentials
Every "temporary credential" you've seen in this lab — the ones IMDS handed the box, and the ones you'll get in Exercise C — comes from one service: STS, the AWS Security Token Service. Its whole job is to issue short-lived credentials so nobody has to carry permanent keys.
Plain Englisha permanent IAM user key (like the prod-ci-service one in
Exercise B) never expires — that's the risk. STS instead vends a time-limited
credential set: you prove who you are, STS hands back keys that work for a few hours
and then die on their own.
A temporary credential is always three values, not one
AccessKeyId ASIA... ← note ASIA (temporary) vs AKIA (permanent user key)
SecretAccessKey ...
SessionToken ... ← the extra piece that marks it temporary; you MUST send all three
Expiration 2026-... ← after this, the creds are dead. No revocation needed.That's why in Exercise C you export all three (AWS_SESSION_TOKEN included) — leave
the session token out and AWS treats them as a malformed permanent key and rejects
them. And the ASIA prefix vs AKIA is how you tell at a glance, in CloudTrail or a
leaked credential, whether you're looking at a temporary STS cred or a permanent user
key.
The main STS operations
| Operation | You give it | You get back | Used by |
|---|---|---|---|
AssumeRole | a role ARN you're allowed to assume | temp creds for that role | Exercise C; IMDS (under the hood) |
AssumeRoleWithWebIdentity | an OIDC/JWT token | temp creds | GitHub Actions, EKS IRSA |
AssumeRoleWithSAML | a SAML assertion | temp creds | enterprise SSO logins |
GetCallerIdentity | nothing | who you are | the aws sts get-caller-identity you keep running |
AssumeRole flow (what Exercise C does, and what IMDS does for you)
caller (you / the EC2 instance) STS IAM
│ │ │
│ AssumeRole(role-arn, session-name) │ │
├─────────────────────────────────────────────────►│ │
│ │ 1. may the caller │
│ │ assume this role?│
│ │ - role TRUST policy │
│ │ allows the caller?├──┐
│ │ - caller has │ │ checks
│ │ sts:AssumeRole? │◄─┘
│ │ │
│ { AccessKeyId(ASIA), SecretAccessKey, │ 2. yes → mint creds │
│ SessionToken, Expiration } │ for the role │
│◄─────────────────────────────────────────────────┤ │
│ │ │
│ now call AWS as assumed-role/<role>/<session-name> │
▼
every later API call carries the 3 values; the ARN shows assumed-role/.../<session>Two gates must both pass: the role's trust policy must name you as an allowed
principal (for the IR roles, that's :root of this account — i.e. anyone in the
account; for prod-app it's the EC2 service), and your own identity needs
sts:AssumeRole permission. Miss either and you get AccessDenied.
Tie-back to IMDSwhen the box fetched creds from
169.254.169.254, AWS ran anAssumeRolefor the instance's role on your behalf — which is exactly whyget-caller-identityon the box showedassumed-role/ir-lab-prod-app/i-017ca7...(role + session-name = instance-id), not a user. Theassumed-roleprefix is STS's fingerprint. And GitHub OIDC is the same picture withAssumeRoleWithWebIdentityinstead of the metadata service.
SSRF: how the creds actually get stolen in the real world
In this lab you ssh'd onto the box to play attacker. In a real breach you usually
don't have SSH — you have a web app with a bug. The most common bug that turns
into "attacker holds the instance role's credentials" is SSRF (Server-Side Request
Forgery).
Plain EnglishSSRF is when you trick a server into making an HTTP request for you, to a destination you couldn't reach directly. The server has network access you don't — so you borrow its position on the network by feeding it a URL.
Why it's devastating on EC2the metadata service at 169.254.169.254 is
reachable only from the instance itself (see the section above). You can't curl it
from your laptop. But if you can make the app on the instance curl it for you and
hand you the response, you get the role's temporary credentials — no shell required.
Concrete example
Say the victim app has an innocent-looking feature — a URL preview / "fetch my avatar from a URL" / webhook tester. It takes a URL and the server fetches it:
POST /preview
Content-Type: application/json
{ "url": "https://example.com/me.png" } ← intended use: server fetches an imageThe server-side code does roughly:
@app.post("/preview")
def preview():
url = request.json["url"]
return requests.get(url).text # ← fetches WHATEVER url you give it. The bug.The attacker simply points url at the metadata service instead of an image:
POST /preview
{ "url": "http://169.254.169.254/latest/meta-data/iam/security-credentials/ir-lab-prod-app" } attacker (laptop) victim EC2 (the app) IMDS (local)
│ POST /preview │ │
│ url=http://169.254.169.254/... │ │
├───────────────────────────────────►│ the SERVER makes the call │
│ ├─────────────────────────────►│
│ │◄── AccessKeyId, Secret, │
│◄────────────────────────────────────┤ Token (JSON) │
│ response body = the role's TEMPORARY CREDENTIALS │The app dutifully fetches the metadata URL — because it'll fetch any URL — and
returns the JSON credentials in the HTTP response. The attacker exports those three
values and is now the ir-lab-prod-app role, exactly as in Exercise A, without ever
having a shell on the box. This is essentially the 2019 Capital One breach: an SSRF
through a misconfigured WAF reached IMDS, stole an over-privileged role, and read
~100M records out of S3.
Why IMDSv2 blunts this (tie-back to the previous section)
The naive SSRF above is a single GET with no custom headers — which is all most SSRF primitives can do. IMDSv2 requires a PUT to get a session token first, then a header carrying that token on the GET. A plain "fetch this URL" feature can't do the PUT or set the header, so the classic one-shot SSRF fails. Combined with the hop-limit-1 default (the response can't be proxied off-box), IMDSv2 closes the common path. (Defense in depth still matters: SSRF can also hit other internal services — databases, admin panels — so you also validate/allowlist outbound URLs and block the link-local range at the app.)
The lesson chainSSRF (app bug) → reach IMDS → steal the role's creds → over-privileged role → account-wide blast radius. Each link is a control you could have added: input validation, IMDSv2, least-privilege role, network egress limits.
Exercises
WHERE you run a command is the whole point. Two machines, two prompts. Watch which one you're at — the same
awscommand gives a different lesson depending on the identity it runs as.
Prompt you see Machine Identity you are [ec2-user@ip-172-31-x-x ~]$the victim EC2 (after ssh)the prod-approle (the attacker's stolen identity)➜ ... 01-iam-foundation/ your shellyour laptop admin(or whatever role you've assumed)Trap: if you run the recon commands on your laptop they "work" — because
admincan do everything. That proves nothing. The finding only counts when the command runs on the victim, as the limitedprod-approle.
A. Recon as the compromised workload (attacker view) — RUN ON THE VICTIM
# [laptop, in 02-billable/] SSH to the victim. After this your prompt becomes
# [ec2-user@ip-172-31-x-x ~]$ — EVERYTHING below runs there, NOT on your laptop.
eval "$(terraform output -raw ssh_command)"
# [victim] Who is this box? Ask the metadata service (IMDSv2: token first).
TOKEN=$(curl -sX PUT http://169.254.169.254/latest/api/token -H "X-aws-ec2-metadata-token-ttl-seconds: 300")
curl -s -H "X-aws-ec2-metadata-token: $TOKEN" http://169.254.169.254/latest/meta-data/iam/security-credentials/
# → expected: ir-lab-prod-app (the role name attached to this instance)
aws sts get-caller-identity
# → expected: "Arn": "arn:aws:sts::<acct>:assumed-role/ir-lab-prod-app/i-0123..."
# NOT user/admin — if you see admin, you're on your laptop, not the box.
# [victim] The over-privileged role reaches data it should never touch:
aws ssm get-parameter --name /ir-lab/prod/db-password --with-decryption \
--query Parameter.Value --output text
# → expected: L4b-Pl@ceholder-NotReal (the "production secret" — read by a workload role)
aws iam list-users
# → expected: a JSON list incl. admin + ir-lab-prod-ci-service (full account enumeration)
# [victim] Blast-radius kicker — the role's s3:* can list (and read) the STATE bucket:
aws s3 ls
# → expected: your buckets incl. ir-lab-tfstate-<acct>. Then prove the damage:
aws s3 cp s3://ir-lab-tfstate-<acct>/ir-lab/01-iam-foundation.tfstate - | head
# → expected: state JSON containing the service-account secret key + RDS password in plaintext.
exit # back to your laptop for Exercise BLessonan over-privileged workload role turns one popped box into account-wide read access — including the Terraform state bucket, where every other secret lives in plaintext. Note how little stands in your way.
B. Find the forgotten service account (persistence) — RUN ON YOUR LAPTOP
# [laptop] inspect the static-key service account — keys are a 01 output
cd 01-iam-foundation
aws iam list-access-keys --user-name ir-lab-prod-ci-service
# → expected: one AccessKeyMetadata entry, "Status": "Active"
aws iam get-access-key-last-used --access-key-id "$(terraform output -raw prod_ci_access_key_id)"
# → expected: LastUsedDate (or "N/A" if never used) — how you'd spot a dormant key
# [laptop] use the stolen keys (simulate the attacker who harvested them from git/CI/disk):
AWS_ACCESS_KEY_ID=$(terraform output -raw prod_ci_access_key_id) \
AWS_SECRET_ACCESS_KEY=$(terraform output -raw prod_ci_secret_access_key) \
aws sts get-caller-identity
# → expected: "Arn": "arn:aws:iam::<acct>:user/ir-lab-prod-ci-service" (NOT admin)Lessonstatic keys = durable persistence. They survive instance termination and
password resets; the only kill is deactivating/deleting the key (aws iam update-access-key --status Inactive ...).
C. Respond — investigate read-only, then acquire & isolate — RUN ON YOUR LAPTOP
# [laptop] role ARNs are 01 outputs. First DROP any stolen creds from Exercise B:
unset AWS_ACCESS_KEY_ID AWS_SECRET_ACCESS_KEY AWS_SESSION_TOKEN
cd 01-iam-foundation
# 1) [laptop] investigate WITHOUT the power to change anything. assume-role returns
# JSON; you must EXPORT the three fields or the next aws call ignores them.
creds=$(aws sts assume-role --role-arn "$(terraform output -raw ir_readonly_role_arn)" \
--role-session-name investigate --query Credentials --output json)
export AWS_ACCESS_KEY_ID=$(echo "$creds" | jq -r .AccessKeyId)
export AWS_SECRET_ACCESS_KEY=$(echo "$creds" | jq -r .SecretAccessKey)
export AWS_SESSION_TOKEN=$(echo "$creds" | jq -r .SessionToken)
aws sts get-caller-identity # → assumed-role/ir-lab-ir-readonly
aws ec2 describe-instances --query 'Reservations[].Instances[].InstanceId' # → works (read)
aws ec2 stop-instances --instance-ids <victim-id>
# → expected: AccessDenied. Read-only can LOOK, never TOUCH. That's the point.
# 2) [laptop] switch to the forensic role for the isolation flow. Drop readonly first:
unset AWS_ACCESS_KEY_ID AWS_SECRET_ACCESS_KEY AWS_SESSION_TOKEN
creds=$(aws sts assume-role --role-arn "$(terraform output -raw ir_forensic_role_arn)" \
--role-session-name acquire --query Credentials --output json)
export AWS_ACCESS_KEY_ID=$(echo "$creds" | jq -r .AccessKeyId)
export AWS_SECRET_ACCESS_KEY=$(echo "$creds" | jq -r .SecretAccessKey)
export AWS_SESSION_TOKEN=$(echo "$creds" | jq -r .SessionToken)
aws sts get-caller-identity # → assumed-role/ir-lab-ir-forensic
# [laptop] grab the pre-filled isolation commands and run them as the forensic role:
cd ../02-billable && terraform output isolation_cheatsheet
# → swaps the victim's instance profile to ir-lab-quarantine (deny-all),
# moves it to the quarantine SG, and lists its volume IDs to snapshot.Lessonseparation of duties — you investigate with read-only, and only the
scoped forensic role can touch the instance. Keep the Exercise-A SSH session open
on the victim, and the moment you run the cheatsheet's profile swap, re-run
aws s3 ls there — it starts returning AccessDenied. Containment, live. (The SSH
shell itself survives: it's role-independent, which is why this lab uses SSH, not
Session Manager.)
D. Snapshot → mount read-only
Create a snapshot of the victim's volume, make a volume from it, attach it to a
forensic instance, and mount -o ro,noexec,nosuid,nodev,norecovery. (Full steps in
the repo's forensics/digital-forensics.md.)
Cleanup reminder
terraform destroy each module in reverse order (02 → 01 → 00); the state bucket
goes last. Then rm -f 02-billable/ir-lab-key.pem. Confirm in the console + Budgets.