Security Notes
Hands-on Labs

Scenario: Identities, the "Default Service Account" Trap, and the Exercises

This is the story the lab tells. Read it once, then run the exercises.

13 min read 7 sections

First: "default service accounts" — AWS vs GCP vs Kubernetes

Your instinct ("there are default service accounts that, if you don't lock them down, an attacker can just use") is correct — but it's mostly a GCP / Kubernetes thing, not AWS. Worth getting straight, because the mental model differs per cloud:

  • GCP really does this: every project gets a default Compute Engine service account that is granted the broad Editor role and is attached to VMs by default. Compromise a VM that's using it → you inherit Editor on the whole project. That's the "default SA you forgot to lock down" you're picturing.
  • Kubernetes does it too: every namespace has a default ServiceAccount, and its token is auto-mounted into pods unless you opt out — so a popped pod often has an API token sitting at /var/run/secrets/... for free.
  • AWS has no such auto-attached "default service account." An EC2 instance has no identity at all unless you attach an instance profile. So in AWS the equivalent risks show up differently, and this lab emulates the two that matter:
    Over-privileged roles attached to compute

    the prod-app role here can read every secret and enumerate IAM. That's the AWS version of "the VM's SA can do too much."

    Long-lived IAM user access keys that were created once and never rotated

    the prod-ci-service user here. Static keys are the credential attackers prize: they don't expire, they're often over-privileged, and nobody notices them. This is the closest AWS analog to "a default credential left active."

Memory hook

One-linerGCP/K8s hand you an over-privileged identity by default; in AWS you hand it to yourself by attaching a fat role or leaving static keys lying around.


The identities this lab creates

Production (what an attacker lands in) — risky on purpose

IdentityTypeWhy it's risky (the lesson)
…-prod-appEC2 role (instance profile)Over-privileged workload role: s3:*, ssm:GetParameter, iam:List*/Get*. A compromised instance inherits all of it.
…-prod-ci-serviceIAM user + static access keyLong-lived "service account" keys with PowerUserAccess — the forgotten, never-rotated credential.
/ir-lab/prod/db-passwordSSM Parameter Store (SecureString)The juicy target the over-privileged role can read — proves "blast radius." (SSM instead of Secrets Manager so Section 1 stays free.)

Incident response (what you respond WITH)

IdentityTypePurpose
…-ir-readonlyrole (ReadOnlyAccess)Investigator — read everything, change nothing.
…-ir-forensicrole (scoped)Acquisition — snapshot, copy, isolate (modify SG), swap instance profile, SSM. Least-privilege for the job.
…-break-glassrole (AdministratorAccess)Emergency admin — normally unused; in real life you alarm on any use.
…-quarantineinstance profile (deny-all)The isolation identity you swap onto a compromised instance.

All three IR roles are assumable by principals in your account: aws sts assume-role --role-arn <arn> --role-session-name ir.

Where to run terraform output: identity outputs (role ARNs, service-account keys, prod_param_name) come from 01-iam-foundation/; the instance/Lambda/RDS outputs (ssh_command, etc.) come from 02-billable/. cd into the right module first. Exercises A–C need Section 2 applied (you need a running victim).


How an EC2 instance gets its identity (and how to audit it)

The single most important AWS concept in this lab: a process on an EC2 instance never holds AWS keys on disk — it borrows a role through the metadata service.

EC2 instance  i-017ca7bf...
   └─ attached:  Instance Profile  "ir-lab-prod-app"   (a thin wrapper/plug)
         └─ contains exactly ONE:  IAM Role  "ir-lab-prod-app"   (the identity)
               └─ policies on the role = what it's ALLOWED to do:
                     • inline policy "overbroad"  →  s3:* , ssm:GetParameter , iam:Get*/List* , ...
  • The instance profile is the plug; the role is the identity; the policies are the permissions. An instance profile holds exactly one role.
  • When any SDK/CLI on the box makes an AWS call, it fetches temporary credentials for that role from the Instance Metadata Service (IMDS) at 169.254.169.254. Those creds auto-rotate every few hours. Nothing is stored on disk — which is why, on a popped box, you can't "find the keys": the box's power is exactly the attached role's policies, delivered live.

Auditing "what can this box do?" — three ways

1. From the box itself (attacker self-enumeration). This role has iam:Get*/List*, so it can read its own permissions — a gift to an attacker:

bash
ROLE=ir-lab-prod-app
aws iam list-role-policies --role-name $ROLE
#   → "PolicyNames": ["overbroad"]
aws iam get-role-policy --role-name $ROLE --policy-name overbroad
#   → the statement: Action [s3:*, ssm:GetParameter*, iam:Get*/List*, ec2:Describe*], Resource "*"
aws iam list-attached-role-policies --role-name $ROLE
#   → []   (no AWS-managed policies here; the whole grant is the inline "overbroad")
aws iam get-instance-profile --instance-profile-name ir-lab-prod-app
#   → shows the profile → role link, and the role's trust policy (who may assume it:
#     Principal Service ec2.amazonaws.com — i.e. "any EC2 instance this is attached to")

2. From your laptop (the IR/investigator view). Same reads as admin or the ir-readonly role — plus the policy simulator, which answers "allowed or denied?" for specific actions without making a real call:

bash
aws iam simulate-principal-policy \
  --policy-source-arn arn:aws:iam::<acct>:role/ir-lab-prod-app \
  --action-names s3:DeleteObject ssm:GetParameter iam:CreateUser \
  --query 'EvaluationResults[].{action:EvalActionName,decision:EvalDecision}'
#   → s3:DeleteObject allowed, ssm:GetParameter allowed, iam:CreateUser implicitDeny

3. From the Terraform source (ground truth). 01-iam-foundation/iam.tf, the prod_app_overbroad policy document. Or Console → IAM → Roles → ir-lab-prod-app.

The findinga web app needs maybe one parameter and one S3 prefix. This role has s3:* + iam:Get*/List* on *. So one compromised box reads any bucket (including the secret-laden Terraform state), every secret in SSM, and the whole IAM layout. Over-privileged workload role = blast radius.


STS: the service that mints temporary credentials

Every "temporary credential" you've seen in this lab — the ones IMDS handed the box, and the ones you'll get in Exercise C — comes from one service: STS, the AWS Security Token Service. Its whole job is to issue short-lived credentials so nobody has to carry permanent keys.

Plain Englisha permanent IAM user key (like the prod-ci-service one in Exercise B) never expires — that's the risk. STS instead vends a time-limited credential set: you prove who you are, STS hands back keys that work for a few hours and then die on their own.

A temporary credential is always three values, not one

AccessKeyId      ASIA...     ← note ASIA (temporary) vs AKIA (permanent user key)
SecretAccessKey  ...
SessionToken     ...         ← the extra piece that marks it temporary; you MUST send all three
Expiration       2026-...    ← after this, the creds are dead. No revocation needed.

That's why in Exercise C you export all three (AWS_SESSION_TOKEN included) — leave the session token out and AWS treats them as a malformed permanent key and rejects them. And the ASIA prefix vs AKIA is how you tell at a glance, in CloudTrail or a leaked credential, whether you're looking at a temporary STS cred or a permanent user key.

The main STS operations

OperationYou give itYou get backUsed by
AssumeRolea role ARN you're allowed to assumetemp creds for that roleExercise C; IMDS (under the hood)
AssumeRoleWithWebIdentityan OIDC/JWT tokentemp credsGitHub Actions, EKS IRSA
AssumeRoleWithSAMLa SAML assertiontemp credsenterprise SSO logins
GetCallerIdentitynothingwho you arethe aws sts get-caller-identity you keep running

AssumeRole flow (what Exercise C does, and what IMDS does for you)

  caller (you / the EC2 instance)                         STS                    IAM
        │                                                  │                      │
        │  AssumeRole(role-arn, session-name)              │                      │
        ├─────────────────────────────────────────────────►│                     │
        │                                                  │  1. may the caller   │
        │                                                  │     assume this role?│
        │                                                  │  - role TRUST policy │
        │                                                  │    allows the caller?├──┐
        │                                                  │  - caller has        │  │ checks
        │                                                  │    sts:AssumeRole?   │◄─┘
        │                                                  │                      │
        │   { AccessKeyId(ASIA), SecretAccessKey,          │  2. yes → mint creds │
        │     SessionToken, Expiration }                   │     for the role     │
        │◄─────────────────────────────────────────────────┤                     │
        │                                                  │                      │
        │  now call AWS as  assumed-role/<role>/<session-name>                    │
        ▼
   every later API call carries the 3 values; the ARN shows assumed-role/.../<session>

Two gates must both pass: the role's trust policy must name you as an allowed principal (for the IR roles, that's :root of this account — i.e. anyone in the account; for prod-app it's the EC2 service), and your own identity needs sts:AssumeRole permission. Miss either and you get AccessDenied.

Tie-back to IMDSwhen the box fetched creds from 169.254.169.254, AWS ran an AssumeRole for the instance's role on your behalf — which is exactly why get-caller-identity on the box showed assumed-role/ir-lab-prod-app/i-017ca7... (role + session-name = instance-id), not a user. The assumed-role prefix is STS's fingerprint. And GitHub OIDC is the same picture with AssumeRoleWithWebIdentity instead of the metadata service.


SSRF: how the creds actually get stolen in the real world

In this lab you ssh'd onto the box to play attacker. In a real breach you usually don't have SSH — you have a web app with a bug. The most common bug that turns into "attacker holds the instance role's credentials" is SSRF (Server-Side Request Forgery).

Plain EnglishSSRF is when you trick a server into making an HTTP request for you, to a destination you couldn't reach directly. The server has network access you don't — so you borrow its position on the network by feeding it a URL.

Why it's devastating on EC2the metadata service at 169.254.169.254 is reachable only from the instance itself (see the section above). You can't curl it from your laptop. But if you can make the app on the instance curl it for you and hand you the response, you get the role's temporary credentials — no shell required.

Concrete example

Say the victim app has an innocent-looking feature — a URL preview / "fetch my avatar from a URL" / webhook tester. It takes a URL and the server fetches it:

POST /preview
Content-Type: application/json

{ "url": "https://example.com/me.png" }     ← intended use: server fetches an image

The server-side code does roughly:

python
@app.post("/preview")
def preview():
    url = request.json["url"]
    return requests.get(url).text          # ← fetches WHATEVER url you give it. The bug.

The attacker simply points url at the metadata service instead of an image:

POST /preview
{ "url": "http://169.254.169.254/latest/meta-data/iam/security-credentials/ir-lab-prod-app" }
        attacker (laptop)                 victim EC2 (the app)              IMDS (local)
            │   POST /preview                    │                              │
            │   url=http://169.254.169.254/...   │                              │
            ├───────────────────────────────────►│  the SERVER makes the call   │
            │                                     ├─────────────────────────────►│
            │                                     │◄── AccessKeyId, Secret,      │
            │◄────────────────────────────────────┤    Token (JSON)             │
            │   response body = the role's TEMPORARY CREDENTIALS                 │

The app dutifully fetches the metadata URL — because it'll fetch any URL — and returns the JSON credentials in the HTTP response. The attacker exports those three values and is now the ir-lab-prod-app role, exactly as in Exercise A, without ever having a shell on the box. This is essentially the 2019 Capital One breach: an SSRF through a misconfigured WAF reached IMDS, stole an over-privileged role, and read ~100M records out of S3.

Why IMDSv2 blunts this (tie-back to the previous section)

The naive SSRF above is a single GET with no custom headers — which is all most SSRF primitives can do. IMDSv2 requires a PUT to get a session token first, then a header carrying that token on the GET. A plain "fetch this URL" feature can't do the PUT or set the header, so the classic one-shot SSRF fails. Combined with the hop-limit-1 default (the response can't be proxied off-box), IMDSv2 closes the common path. (Defense in depth still matters: SSRF can also hit other internal services — databases, admin panels — so you also validate/allowlist outbound URLs and block the link-local range at the app.)

The lesson chainSSRF (app bug) → reach IMDS → steal the role's creds → over-privileged role → account-wide blast radius. Each link is a control you could have added: input validation, IMDSv2, least-privilege role, network egress limits.


Exercises

WHERE you run a command is the whole point. Two machines, two prompts. Watch which one you're at — the same aws command gives a different lesson depending on the identity it runs as.

Prompt you seeMachineIdentity you are
[ec2-user@ip-172-31-x-x ~]$the victim EC2 (after ssh)the prod-app role (the attacker's stolen identity)
➜ ... 01-iam-foundation / your shellyour laptopadmin (or whatever role you've assumed)

Trap: if you run the recon commands on your laptop they "work" — because admin can do everything. That proves nothing. The finding only counts when the command runs on the victim, as the limited prod-app role.

A. Recon as the compromised workload (attacker view) — RUN ON THE VICTIM

bash
# [laptop, in 02-billable/]  SSH to the victim. After this your prompt becomes
# [ec2-user@ip-172-31-x-x ~]$ — EVERYTHING below runs there, NOT on your laptop.
eval "$(terraform output -raw ssh_command)"

# [victim]  Who is this box? Ask the metadata service (IMDSv2: token first).
TOKEN=$(curl -sX PUT http://169.254.169.254/latest/api/token -H "X-aws-ec2-metadata-token-ttl-seconds: 300")
curl -s -H "X-aws-ec2-metadata-token: $TOKEN" http://169.254.169.254/latest/meta-data/iam/security-credentials/
#   → expected:  ir-lab-prod-app           (the role name attached to this instance)

aws sts get-caller-identity
#   → expected:  "Arn": "arn:aws:sts::<acct>:assumed-role/ir-lab-prod-app/i-0123..."
#                NOT user/admin — if you see admin, you're on your laptop, not the box.

# [victim]  The over-privileged role reaches data it should never touch:
aws ssm get-parameter --name /ir-lab/prod/db-password --with-decryption \
  --query Parameter.Value --output text
#   → expected:  L4b-Pl@ceholder-NotReal   (the "production secret" — read by a workload role)

aws iam list-users
#   → expected:  a JSON list incl. admin + ir-lab-prod-ci-service  (full account enumeration)

# [victim]  Blast-radius kicker — the role's s3:* can list (and read) the STATE bucket:
aws s3 ls
#   → expected:  your buckets incl. ir-lab-tfstate-<acct>. Then prove the damage:
aws s3 cp s3://ir-lab-tfstate-<acct>/ir-lab/01-iam-foundation.tfstate - | head
#   → expected:  state JSON containing the service-account secret key + RDS password in plaintext.

exit   # back to your laptop for Exercise B

Lessonan over-privileged workload role turns one popped box into account-wide read access — including the Terraform state bucket, where every other secret lives in plaintext. Note how little stands in your way.

B. Find the forgotten service account (persistence) — RUN ON YOUR LAPTOP

bash
# [laptop]  inspect the static-key service account — keys are a 01 output
cd 01-iam-foundation
aws iam list-access-keys --user-name ir-lab-prod-ci-service
#   → expected:  one AccessKeyMetadata entry, "Status": "Active"
aws iam get-access-key-last-used --access-key-id "$(terraform output -raw prod_ci_access_key_id)"
#   → expected:  LastUsedDate (or "N/A" if never used) — how you'd spot a dormant key

# [laptop]  use the stolen keys (simulate the attacker who harvested them from git/CI/disk):
AWS_ACCESS_KEY_ID=$(terraform output -raw prod_ci_access_key_id) \
AWS_SECRET_ACCESS_KEY=$(terraform output -raw prod_ci_secret_access_key) \
aws sts get-caller-identity
#   → expected:  "Arn": "arn:aws:iam::<acct>:user/ir-lab-prod-ci-service"  (NOT admin)

Lessonstatic keys = durable persistence. They survive instance termination and password resets; the only kill is deactivating/deleting the key (aws iam update-access-key --status Inactive ...).

C. Respond — investigate read-only, then acquire & isolate — RUN ON YOUR LAPTOP

bash
# [laptop]  role ARNs are 01 outputs. First DROP any stolen creds from Exercise B:
unset AWS_ACCESS_KEY_ID AWS_SECRET_ACCESS_KEY AWS_SESSION_TOKEN
cd 01-iam-foundation

# 1) [laptop]  investigate WITHOUT the power to change anything. assume-role returns
#    JSON; you must EXPORT the three fields or the next aws call ignores them.
creds=$(aws sts assume-role --role-arn "$(terraform output -raw ir_readonly_role_arn)" \
  --role-session-name investigate --query Credentials --output json)
export AWS_ACCESS_KEY_ID=$(echo "$creds" | jq -r .AccessKeyId)
export AWS_SECRET_ACCESS_KEY=$(echo "$creds" | jq -r .SecretAccessKey)
export AWS_SESSION_TOKEN=$(echo "$creds" | jq -r .SessionToken)
aws sts get-caller-identity         # → assumed-role/ir-lab-ir-readonly
aws ec2 describe-instances --query 'Reservations[].Instances[].InstanceId'   # → works (read)
aws ec2 stop-instances --instance-ids <victim-id>
#   → expected:  AccessDenied. Read-only can LOOK, never TOUCH. That's the point.

# 2) [laptop]  switch to the forensic role for the isolation flow. Drop readonly first:
unset AWS_ACCESS_KEY_ID AWS_SECRET_ACCESS_KEY AWS_SESSION_TOKEN
creds=$(aws sts assume-role --role-arn "$(terraform output -raw ir_forensic_role_arn)" \
  --role-session-name acquire --query Credentials --output json)
export AWS_ACCESS_KEY_ID=$(echo "$creds" | jq -r .AccessKeyId)
export AWS_SECRET_ACCESS_KEY=$(echo "$creds" | jq -r .SecretAccessKey)
export AWS_SESSION_TOKEN=$(echo "$creds" | jq -r .SessionToken)
aws sts get-caller-identity         # → assumed-role/ir-lab-ir-forensic

# [laptop]  grab the pre-filled isolation commands and run them as the forensic role:
cd ../02-billable && terraform output isolation_cheatsheet
#   → swaps the victim's instance profile to ir-lab-quarantine (deny-all),
#     moves it to the quarantine SG, and lists its volume IDs to snapshot.

Lessonseparation of duties — you investigate with read-only, and only the scoped forensic role can touch the instance. Keep the Exercise-A SSH session open on the victim, and the moment you run the cheatsheet's profile swap, re-run aws s3 ls there — it starts returning AccessDenied. Containment, live. (The SSH shell itself survives: it's role-independent, which is why this lab uses SSH, not Session Manager.)

D. Snapshot → mount read-only

Create a snapshot of the victim's volume, make a volume from it, attach it to a forensic instance, and mount -o ro,noexec,nosuid,nodev,norecovery. (Full steps in the repo's forensics/digital-forensics.md.)


Cleanup reminder

terraform destroy each module in reverse order (02 → 01 → 00); the state bucket goes last. Then rm -f 02-billable/ir-lab-key.pem. Confirm in the console + Budgets.