AWS Incident Response Playbooks
IR Framework for AWS
PICERL applied to AWS
CloudTrail enabled all regions, GuardDuty active, VPC Flow Logs, Security Hub
GuardDuty finding, CloudTrail anomaly, SIEM alert
isolate IAM credentials, isolate EC2, revoke sessions
remove backdoors, rotate creds, patch vulnerability
restore from clean snapshot/AMI, re-deploy from IaC
post-mortem, update runbooks, new detection rules
Scenario: Compromised IAM Credentials
DetectionGuardDuty UnauthorizedAccess:IAMUser/ConsoleLoginSuccess.B or unusual API calls from unexpected IP
Immediate Containment
# 1. Identify the credential type
aws sts get-caller-identity # shows account, user, assumed-role ARN
# 2. List access keys for the user
aws iam list-access-keys --user-name <username>
# 3. Deactivate (not delete yet — preserve forensics)
aws iam update-access-key --user-name <username> --access-key-id <key-id> --status Inactive
# 4. Attach an explicit DENY policy to stop all actions
aws iam put-user-policy --user-name <username> --policy-name INCIDENT-LOCKOUT --policy-document '{
"Version": "2012-10-17",
"Statement": [{"Effect":"Deny","Action":"*","Resource":"*"}]
}'
# 5. Revoke active console sessions (requires IAM identity center if SSO)
aws iam delete-login-profile --user-name <username>Investigation
# What did the compromised credential do? (CloudTrail)
aws cloudtrail lookup-events \
--lookup-attributes AttributeKey=Username,AttributeValue=<username> \
--start-time 2026-01-01T00:00:00Z \
--end-time 2026-01-02T00:00:00Z
# Check for new IAM users/keys created (persistence)
aws iam list-users --query 'Users[?CreateDate>=`2026-01-01`]'
aws iam list-access-keys # check all users
# Check for new roles with trust policies
aws iam list-roles --query 'Roles[?CreateDate>=`2026-01-01`]'
# What S3 data was accessed?
aws s3api list-buckets # then check S3 access logs or CloudTrail data events
# Check for EC2 instances launched
aws ec2 describe-instances --filters "Name=launch-time,Values=2026-01-01*"Eradication
# Delete malicious access keys
aws iam delete-access-key --user-name <username> --access-key-id <key-id>
# Remove backdoor users/roles
aws iam delete-user --user-name <backdoor-user>
aws iam delete-role --role-name <backdoor-role>
# Remove malicious policies
aws iam list-user-policies --user-name <username>
aws iam delete-user-policy --user-name <username> --policy-name <malicious-policy>
# Rotate legitimate credentials
aws iam create-access-key --user-name <username>Scenario: EC2 Instance Compromise
DetectionGuardDuty CryptoCurrency:EC2/BitcoinTool.B!DNS, unusual outbound connections in VPC Flow Logs
How to think about it — what exactly are you checking?
An EC2 compromise has two blast radii, and missing the second is the classic mistake:
the OS, processes, persistence, what the attacker ran on the box.
this is usually the bigger prize. The instance has a role (via its instance profile), and its credentials are reachable at the metadata endpoint 169.254.169.254. If the attacker got code execution, assume they stole the role's credentials and used them against the AWS API from their own infrastructure — which means the incident is no longer contained to the host.
So your investigation runs on two tracks in parallel:
| Track | What you check | Where |
|---|---|---|
| Host | Running processes, new cron/systemd, modified binaries, reverse shells, dropped tools, bash history, auth logs, outbound C2 | EBS snapshot + memory capture |
| Cloud / IAM | What the instance's role did in CloudTrail — especially API calls from a source IP that isn't the instance (= stolen creds used elsewhere) | CloudTrail, filtered to the role's session |
Memory hookThe key check — "did the role credentials walk?" Pull CloudTrail for the instance's role and compare the source IP of each call against the instance's own IP. AWS records the instance ID in the session for IMDS-derived creds, but calls coming from an external IP using that role are the smoking gun that credentials were exfiltrated and are being used off-box. This single comparison decides whether you're handling one bad server or an account-wide intrusion.
Order of operationsisolate (quarantine SG) → preserve (snapshot + memory) → disable IMDS so no further creds can be pulled → investigate both tracks → if creds walked, treat as a compromised-credential incident too.
Immediate Containment
# 1. Isolate EC2 — attach quarantine security group (deny all)
aws ec2 create-security-group --group-name QUARANTINE --description "IR isolation"
# (no inbound or outbound rules = deny all)
aws ec2 modify-instance-attribute \
--instance-id <i-xxx> \
--groups <quarantine-sg-id>
# 2. Take EBS snapshot for forensics (before any changes)
VOLUME=$(aws ec2 describe-instances --instance-ids <i-xxx> \
--query 'Reservations[*].Instances[*].BlockDeviceMappings[0].Ebs.VolumeId' --output text)
aws ec2 create-snapshot --volume-id $VOLUME --description "IR forensics $(date -u +%Y%m%dT%H%M%SZ)"
# 3. Capture memory (requires SSM agent or pre-installed tool)
aws ssm send-command \
--instance-ids <i-xxx> \
--document-name "AWS-RunShellScript" \
--parameters 'commands=["avml /tmp/memory.lime && aws s3 cp /tmp/memory.lime s3://ir-bucket/"]'
# 4. Disable instance metadata (prevent further cred theft)
aws ec2 modify-instance-metadata-options \
--instance-id <i-xxx> \
--http-endpoint disabledInvestigation
# VPC Flow Logs — find C2 connections
aws logs filter-log-events \
--log-group-name /aws/vpc/flowlogs \
--filter-pattern "[version, accountId, interfaceId, srcAddr, dstAddr, srcPort, dstPort, protocol, packets, bytes, start, end, action, logStatus]"
# Check what IAM role the instance had and what it accessed
aws cloudtrail lookup-events \
--lookup-attributes AttributeKey=ResourceName,AttributeValue=<instance-id>
# Analyse EBS snapshot: attach to forensic instance, mount read-only
aws ec2 create-volume --snapshot-id <snap-id> --availability-zone us-east-1a
# Mount: sudo mount -o ro /dev/xvdf1 /mnt/forensicsRecovery
- Launch replacement from known-good AMI via IaC (Terraform/CloudFormation)
- Do not restart compromised instance
- Terminate after forensics complete
Scenario: Compromised Pod in EKS
DetectionGuardDuty EKS findings (Execution:Kubernetes/ExecInKubeSystemPod, PrivilegeEscalation:Kubernetes/...), Falco runtime alerts (shell in container, sensitive mount), unusual egress from a pod, a workload doing AWS API calls it never did before.
How to think about it — the three escape ladders
A compromised pod is dangerous because of what it can reach beyond itself. Investigate along three escalation ladders, because the attacker will try all of them:
LADDER 1 — Pod → AWS (credential theft)
Does the pod have AWS permissions (IRSA / Pod Identity)? → attacker assumes that role
Can the pod reach the NODE's metadata (169.254.169.254)? → steals the NODE's role
(far more powerful — the node role often has broad EKS/ECR/EC2 perms)
LADDER 2 — Pod → Kubernetes (cluster escalation)
What can the pod's SERVICE ACCOUNT do via the K8s API? → list secrets? create pods?
Is the service-account token mounted? (default: yes) → attacker uses it against kube-apiserver
Any RBAC that lets it read secrets / escalate / exec? → cluster-wide compromise
LADDER 3 — Pod → Node → other pods (container escape)
Is the container privileged / hostPID / hostNetwork? → trivial escape to the node
hostPath mount of / or the docker socket? → own the node, then all pods on it
Known runtime-escape CVE? → break out of the containerWhat exactly to check
| Question | Where to look |
|---|---|
| What is the pod and what does it run? | kubectl describe pod, the image, its command, env vars |
| Does it have AWS creds? | Is a service account annotated for IRSA / associated via Pod Identity? What's the role's policy? Check CloudTrail for that role. |
| Can it reach node metadata? | Is IMDS reachable from pods? (It shouldn't be — block it / enforce IMDSv2 hop-limit 1.) Did the node role make unusual API calls? |
| What can its service account do? | kubectl auth can-i --list --as=system:serviceaccount:<ns>:<sa> — enumerate its RBAC |
| Is it over-privileged at the container level? | securityContext: privileged, allowPrivilegeEscalation, hostPID/hostNetwork/hostPath, capabilities |
| What did it touch? | EKS control-plane audit logs (API calls the SA made), Falco/runtime logs, VPC Flow Logs for egress, CloudTrail for any AWS calls |
Containment
# 1. Isolate the pod with a deny-all NetworkPolicy (cut its network)
kubectl label pod <pod> quarantine=true -n <ns>
# apply a NetworkPolicy selecting quarantine=true that allows no ingress/egress
# 2. Do NOT delete the pod yet — you lose live evidence. Cordon the node so
# nothing reschedules, and capture first.
kubectl cordon <node>
# 3. Capture evidence: running processes, container filesystem, memory if possible
kubectl exec <pod> -n <ns> -- ps aux # (only if safe; exec is itself logged/risky)
# better: snapshot the node's EBS volume and analyze offline
# 4. Revoke the pod's identity:
# - Remove the IRSA role mapping / Pod Identity association
# - If the K8s service-account token may be stolen, rotate it and review RBAC
# - If the NODE role may be compromised, treat the whole node as compromised
# 5. Once evidence is captured: delete the pod, and replace the node from a
# clean image (don't trust a node a container may have escaped to).Eradication & the key escalation question
The decisive question: did the attacker get off the pod?
delete pod, fix the vuln/image, rotate the pod's identity.
the node and its IAM role are compromised; cordon/drain/terminate the node, rotate the node role, and assume every pod that ran on it is suspect.
potentially cluster-wide; audit what the token did, rotate it, and review for created backdoors (new service accounts, role bindings, daemonsets).
Memory hook"a pod is a process, not a security boundary." The container shares the node's kernel, so the real questions in an EKS compromise are never "is the container bad" but "what can it reach": the node's IAM role, the Kubernetes API, or the kernel. Lock those three doors in advance — block pod access to IMDS, give each pod a least-privilege service account and IRSA role (never rely on the node role), and forbid privileged/hostPath pods via Pod Security admission — and a single bad pod stays a single bad pod. (See
../../kubernetes/incident-response.mdand../../detection/endpoint-detection.md.)
Scenario: S3 Data Exfiltration
DetectionMacie alert on PII access; CloudTrail GetObject from unusual IP; GuardDuty Exfiltration:S3/ObjectRead.Unusual
# 1. Identify what was accessed (CloudTrail data events must be enabled)
aws cloudtrail lookup-events \
--lookup-attributes AttributeKey=ResourceName,AttributeValue=<bucket-name>
# 2. Block public access immediately if not already
aws s3api put-public-access-block \
--bucket <bucket-name> \
--public-access-block-configuration "BlockPublicAcls=true,IgnorePublicAcls=true,BlockPublicPolicy=true,RestrictPublicBuckets=true"
# 3. Revoke the credential used for access (see IAM playbook above)
# 4. Preserve access logs
aws s3 cp s3://<logging-bucket>/ /tmp/s3-logs/ --recursive
# 5. Check bucket policy for unexpected principals
aws s3api get-bucket-policy --bucket <bucket-name>
# 6. Enable Object Lock on sensitive buckets going forward
aws s3api put-object-lock-configuration \
--bucket <bucket-name> \
--object-lock-configuration '{"ObjectLockEnabled":"Enabled"}'Scenario: Lambda Function Compromise
DetectionUnusual invocations, unexpected external calls, IAM calls from Lambda execution role
# 1. Throttle Lambda to stop execution
aws lambda put-function-concurrency \
--function-name <function-name> \
--reserved-concurrent-executions 0
# 2. Get current code for forensics
aws lambda get-function --function-name <function-name> --query 'Code.Location'
# Download and analyse
# 3. Check CloudTrail for what the execution role did
aws cloudtrail lookup-events \
--lookup-attributes AttributeKey=Username,AttributeValue=<execution-role-arn>
# 4. Rotate/detach execution role
aws lambda update-function-configuration \
--function-name <function-name> \
--role arn:aws:iam::123456789012:role/minimal-read-only-role
# 5. Redeploy from source control after reviewGuardDuty Finding → Response Mapping
| Finding | Immediate Action |
|---|---|
UnauthorizedAccess:IAMUser/ConsoleLoginSuccess.B | Lock user, investigate trail |
CryptoCurrency:EC2/BitcoinTool.B!DNS | Isolate EC2, snapshot |
Recon:EC2/PortProbeUnprotectedPort | Tighten SG, check if exploit followed |
Exfiltration:S3/ObjectRead.Unusual | Block public access, revoke creds |
PrivilegeEscalation:IAMUser/AnomalousBehavior | Deny-all policy on user, audit trail |
Persistence:IAMUser/NetworkPermissions | Audit new policies/roles, revoke |
Backdoor:EC2/C&CActivity.B!DNS | Isolate, snapshot, investigate |
Useful IR Queries (CloudTrail Insights)
# All root account activity (should be rare)
aws cloudtrail lookup-events \
--lookup-attributes AttributeKey=Username,AttributeValue=root
# Console logins from unusual sources
aws cloudtrail lookup-events \
--lookup-attributes AttributeKey=EventName,AttributeValue=ConsoleLogin
# IAM changes in last 24h
aws cloudtrail lookup-events \
--lookup-attributes AttributeKey=EventSource,AttributeValue=iam.amazonaws.com \
--start-time $(date -u -d '24 hours ago' +%Y-%m-%dT%H:%M:%SZ)
# New security group rules (potential firewall weakening)
aws cloudtrail lookup-events \
--lookup-attributes AttributeKey=EventName,AttributeValue=AuthorizeSecurityGroupIngressInterview Questions: AWS IR
CryptoCurrency:EC2/BitcoinTool.B!DNS. Walk me through containment.First isolate the instance by swapping its security groups for a quarantine group with no rules, cutting the C2 and mining traffic while keeping the box alive for forensics. Then preserve evidence before changing anything: snapshot the EBS volumes and, if I can, capture memory via SSM. Critically, I disable the instance metadata endpoint so no further role credentials can be pulled. Then I investigate two tracks in parallel — the host (processes, persistence, how they got in) from the snapshot, and the cloud side, pulling CloudTrail for the instance's IAM role and comparing source IPs to see whether the role's credentials were stolen and used off-box. If they were, it becomes an account-wide credential incident too. Recovery is to rebuild from a known-good AMI via IaC and terminate the compromised instance after forensics — never reboot and reuse it.
I avoid stopping it because shutting down loses volatile memory and can trigger attacker dead-man's switches. Instead I isolate it at the network layer with a quarantine security group so it can't do harm but stays running. I capture memory live — via the SSM agent running a tool like AVML and shipping the image to an S3 evidence bucket — because RAM holds the most ephemeral evidence: injected code, keys, and C2 state. Then I snapshot the EBS volumes for disk forensics, which I later attach read-only to a separate forensic instance. I also disable IMDS to stop further credential theft. Everything is hashed and logged for chain of custody. The instance keeps running, isolated, until analysis is complete, then it's terminated rather than reused.
One: deactivate the key — set it inactive, which instantly stops it but is reversible and preserves it as evidence, rather than deleting it outright. Two: scope its activity in CloudTrail filtered to that access key ID — every call, source IP, and timestamp, looking for recon-then-escalation-then-creation patterns. Three: hunt for persistence the attacker created — new IAM users, access keys, roles, or trust-policy changes — because that's what survives a simple key rotation. Four: contain any active abuse, like terminating crypto-mining EC2 they launched, and roll any secrets the key could read. Five: eradicate and harden — delete the key and the backdoor identities, and replace the IAM user with SSO or a role so there's no long-lived key to leak again, plus enable push protection and detections. The theme is that rotating the key alone is never enough.
The primary source is CloudTrail S3 data events, which log object-level GetObject calls — but they must be enabled in advance, so step zero is ensuring they're on. With them, I look for anomalous reads: large volumes of GetObject, access from unusual source IPs or principals, or a principal reading objects it never touched before. S3 server access logs are a secondary source. GuardDuty's S3 protection and Macie help — GuardDuty flags unusual object-read patterns, and Macie identifies which buckets hold sensitive data so I can prioritize. I'd also check the bucket policy and ACLs for unexpected principals or public access that enabled the exfil. The honest caveat I'd raise: if data events weren't enabled beforehand, object-level reads may simply not be recorded, which is itself a preparation gap to fix.
Deactivating sets the key to inactive — it immediately stops working but still exists, so it's fully reversible and, importantly, preserved as evidence and as a pivot for investigation. Deleting removes it permanently. In incident response you almost always deactivate first, because it achieves instant containment without destroying forensic value: you can still tie the key ID to its CloudTrail history and confirm scope. You delete only after scoping is complete, as part of eradication. Deleting prematurely is a mistake — it's irreversible, and you lose the ability to cleanly correlate the key to its activity.
PassRole is the permission to hand a role to a service, and the escalation is that someone who can pass any role plus run compute can launch an instance or Lambda with an admin role and borrow its permissions. The core fix is to scope PassRole tightly: in the identity policy, restrict the Resource to only the specific role ARNs that principal legitimately needs to pass, never a wildcard, and ideally add a condition on iam:PassedToService so a role can only be passed to the intended service. More broadly I treat any IAM-write permission — PassRole, AttachPolicy, CreateAccessKey, UpdateAssumeRolePolicy — as admin-equivalent, since they're paths to admin, and I'd run a tool like PMapper to map who can reach admin and close those edges. Permission boundaries also cap what self-created roles can do.