Security Notes
Hands-on Labs

AWS EKS Security Scenario Lab (Terraform)

The capstone of the AWS labs. ir-scenario stole an EC2 role's credentials through IMDS; cloudtrail-detection made that theft visible. This lab tells the same story one layer up — inside Kubernetes — and adds the attacks that only exist there: pods assuming IAM roles (IRSA), in-cluster RBAC privilege escalation, and pod-level containment.

4 min read 6 sections verified 2026-06

Last verified2026-06

00-bootstrap/        one-time: creates the S3 bucket that holds this lab's REMOTE state  (~pennies)
(lab root)           the cluster itself:
  EKS control plane  (managed API server + etcd, audit logging ON)
     ├─ OIDC provider ─────────► IRSA: a pod assumes an IAM role via STS
     ├─ managed node group ────► 1× t3.small (IMDS intentionally pod-reachable)
     └─ VPC CNI w/ NetworkPolicy enforcement ON
  k8s/ manifests:    IRSA workload • over-permissive RBAC • NetworkPolicy • Pod Security
⚠️

Throwaway sandbox account only. This builds an intentionally over-privileged IRSA role, over-permissive RBAC, and a weakened metadata service. Tear it down.

This lab is split into Terraform (the AWS cluster) and k8s/ manifests (the in-cluster objects). They're kept separate on purpose: the Kubernetes provider can't authenticate to a cluster that doesn't exist yet, so mixing them in one apply is fragile. Build the cluster with Terraform, then kubectl apply the rest.


What this costs

Last verified2026-06 — verify against current AWS pricing before a long run.

ThingCost
EKS control plane$0.10 / hour ($2.40/day, ~$73/mo) — charged whether or not you run pods. The main cost.
1× t3.small node$0.021/hr ($0.50/day)
EBS (20 GB gp3) + control-plane logsPennies/day
NAT / data transfer$0 — default VPC, public subnets, no NAT gateway

≈ $3/day. Easily inside a $180 credit, but terraform destroy the same day — the control plane bills hourly even while idle.


Prerequisites

bash
brew install kubectl                       # plus the toolchain from ir-scenario's README
  • A throwaway sandbox account + the non-root ir-lab CLI profile (see ir-scenario README Step 0–1).
  • An S3 state bucket — this lab's 00-bootstrap/ creates a dedicated one (eks-lab-tfstate-<account-id>) in Step 1 below. (Or reuse ir-scenario's shared ir-lab-tfstate-<account-id> bucket — this module uses its own state key either way.)
  • (Recommended) the cloudtrail-detection lab applied, so the IRSA AssumeRoleWithWebIdentity calls and the stolen-credential use show up in CloudTrail and the EKS audit log.

Spin up

bash
cd cloud/aws/labs/eks-scenario
cp .envrc.example .envrc && direnv allow

# Step 1 — create the state bucket (once). Skip if reusing an existing bucket.
cd 00-bootstrap
terraform init
terraform apply                                 # creates eks-lab-tfstate-<account-id>
terraform output -raw state_bucket              # copy this name
cd ..

# Step 2 — the cluster (remote state in that bucket)
cp backend.hcl.example backend.hcl              # paste the bucket name from Step 1
terraform init -backend-config=backend.hcl
terraform apply                                 # ~10–15 min (control plane + node group)

Point kubectl at the new cluster, then apply the in-cluster objects:

bash
eval "$(terraform output -raw update_kubeconfig)"   # aws eks update-kubeconfig ...
kubectl get nodes                                   # 1 node, Ready

# IRSA workload — substitute the role ARN Terraform created into the ServiceAccount:
ROLE=$(terraform output -raw overbroad_irsa_role_arn)
sed "s#ROLE_ARN_PLACEHOLDER#$ROLE#" k8s/10-irsa-workload.yaml | kubectl apply -f -

kubectl apply -f k8s/20-rbac-privesc.yaml
kubectl apply -f k8s/30-network-policy.yaml   # quarantine policy — inert until a pod is labelled
# NOTE: do NOT apply 35-default-deny-egress.yaml here — it blocks egress to AWS and
# is applied deliberately in the hardening exercise (D). Same for 40-psa-restricted
# (one of its pods is meant to be rejected).

One small node is tightCoreDNS wants two replicas; on a single t3.small one may stay Pending. The lab still works (one CoreDNS is enough). If pods won't schedule, terraform apply -var 'node_instance_type=t3.medium'.


Use it

The full walkthrough — attacker view, in-cluster privesc, the RBAC→IRSA→cloud bridge, containment, and hardening, plus interview Q&A — is in SCENARIO.md. The four exercises:

#StoryYou practice
AIRSA + IMDS theftPod assumes an IAM role; steal node creds via IMDS; blast radius
BIn-cluster RBAC privescA "deployer" SA escalates to cloud compromise
CPod IR / containmentIsolate a pod with NetworkPolicy, cordon/drain, revoke IRSA
DAdmission control / hardeningPod Security Standards, fix IMDS, default-deny egress

Tear down

bash
kubectl delete -f k8s/ --ignore-not-found     # optional; destroy removes the cluster anyway
terraform destroy                              # ~10–15 min (the cluster)
cd 00-bootstrap && terraform destroy && cd ..  # the state bucket, LAST

terraform destroy deletes the node group, control plane, OIDC provider, IRSA role, and SSM parameter. Confirm the EKS console reads clean and check that the node EC2 instance is gone (an orphaned node would keep billing).


Notes & gotchas

The control plane bills hourly, idle or not.

This is the one lab where "forgot to destroy it" actually costs real money. Destroy same day.

NetworkPolicy is off by default on EKS.

The VPC CNI accepts policies but doesn't enforce them unless enableNetworkPolicy=true — which this lab sets via the vpc-cni addon. Exercise C depends on it.

IMDS is intentionally weak here

(http_tokens=optional, hop limit 2) so the pod-steals-node-role attack works. Real clusters set hop limit 1 (or block 169.254.169.254 from pods) and require IMDSv2 — Exercise D flips it.

Audit logging is on

(api, audit, authenticator → CloudWatch /aws/eks/<name>/cluster). That's the in-cluster equivalent of CloudTrail, and how the RBAC privesc in Exercise B is detectable.

IRSA is the secure pattern done wrong here.

The point isn't that IRSA is bad — it's that a too-broad IRSA role has the same account-wide blast radius as a popped EC2. Scope the trust policy to the exact ServiceAccount and the permissions to the one parameter the app needs.