Security Notes
Cloud

AWS Incident Response Playbooks

13 min read 9 sections 6 model answers

IR Framework for AWS

PICERL applied to AWS

Preparation

CloudTrail enabled all regions, GuardDuty active, VPC Flow Logs, Security Hub

Identification

GuardDuty finding, CloudTrail anomaly, SIEM alert

Containment

isolate IAM credentials, isolate EC2, revoke sessions

Eradication

remove backdoors, rotate creds, patch vulnerability

Recovery

restore from clean snapshot/AMI, re-deploy from IaC

Lessons Learned

post-mortem, update runbooks, new detection rules


Scenario: Compromised IAM Credentials

DetectionGuardDuty UnauthorizedAccess:IAMUser/ConsoleLoginSuccess.B or unusual API calls from unexpected IP

Immediate Containment

bash
# 1. Identify the credential type
aws sts get-caller-identity   # shows account, user, assumed-role ARN

# 2. List access keys for the user
aws iam list-access-keys --user-name <username>

# 3. Deactivate (not delete yet — preserve forensics)
aws iam update-access-key --user-name <username> --access-key-id <key-id> --status Inactive

# 4. Attach an explicit DENY policy to stop all actions
aws iam put-user-policy --user-name <username> --policy-name INCIDENT-LOCKOUT --policy-document '{
  "Version": "2012-10-17",
  "Statement": [{"Effect":"Deny","Action":"*","Resource":"*"}]
}'

# 5. Revoke active console sessions (requires IAM identity center if SSO)
aws iam delete-login-profile --user-name <username>

Investigation

bash
# What did the compromised credential do? (CloudTrail)
aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=Username,AttributeValue=<username> \
  --start-time 2026-01-01T00:00:00Z \
  --end-time 2026-01-02T00:00:00Z

# Check for new IAM users/keys created (persistence)
aws iam list-users --query 'Users[?CreateDate>=`2026-01-01`]'
aws iam list-access-keys   # check all users

# Check for new roles with trust policies
aws iam list-roles --query 'Roles[?CreateDate>=`2026-01-01`]'

# What S3 data was accessed?
aws s3api list-buckets  # then check S3 access logs or CloudTrail data events

# Check for EC2 instances launched
aws ec2 describe-instances --filters "Name=launch-time,Values=2026-01-01*"

Eradication

bash
# Delete malicious access keys
aws iam delete-access-key --user-name <username> --access-key-id <key-id>

# Remove backdoor users/roles
aws iam delete-user --user-name <backdoor-user>
aws iam delete-role --role-name <backdoor-role>

# Remove malicious policies
aws iam list-user-policies --user-name <username>
aws iam delete-user-policy --user-name <username> --policy-name <malicious-policy>

# Rotate legitimate credentials
aws iam create-access-key --user-name <username>

Scenario: EC2 Instance Compromise

DetectionGuardDuty CryptoCurrency:EC2/BitcoinTool.B!DNS, unusual outbound connections in VPC Flow Logs

How to think about it — what exactly are you checking?

An EC2 compromise has two blast radii, and missing the second is the classic mistake:

The host itself

the OS, processes, persistence, what the attacker ran on the box.

The instance's IAM role

this is usually the bigger prize. The instance has a role (via its instance profile), and its credentials are reachable at the metadata endpoint 169.254.169.254. If the attacker got code execution, assume they stole the role's credentials and used them against the AWS API from their own infrastructure — which means the incident is no longer contained to the host.

So your investigation runs on two tracks in parallel:

TrackWhat you checkWhere
HostRunning processes, new cron/systemd, modified binaries, reverse shells, dropped tools, bash history, auth logs, outbound C2EBS snapshot + memory capture
Cloud / IAMWhat the instance's role did in CloudTrail — especially API calls from a source IP that isn't the instance (= stolen creds used elsewhere)CloudTrail, filtered to the role's session
Memory hook

The key check — "did the role credentials walk?" Pull CloudTrail for the instance's role and compare the source IP of each call against the instance's own IP. AWS records the instance ID in the session for IMDS-derived creds, but calls coming from an external IP using that role are the smoking gun that credentials were exfiltrated and are being used off-box. This single comparison decides whether you're handling one bad server or an account-wide intrusion.

Order of operationsisolate (quarantine SG) → preserve (snapshot + memory) → disable IMDS so no further creds can be pulled → investigate both tracks → if creds walked, treat as a compromised-credential incident too.

Immediate Containment

bash
# 1. Isolate EC2 — attach quarantine security group (deny all)
aws ec2 create-security-group --group-name QUARANTINE --description "IR isolation"
# (no inbound or outbound rules = deny all)

aws ec2 modify-instance-attribute \
  --instance-id <i-xxx> \
  --groups <quarantine-sg-id>

# 2. Take EBS snapshot for forensics (before any changes)
VOLUME=$(aws ec2 describe-instances --instance-ids <i-xxx> \
  --query 'Reservations[*].Instances[*].BlockDeviceMappings[0].Ebs.VolumeId' --output text)
aws ec2 create-snapshot --volume-id $VOLUME --description "IR forensics $(date -u +%Y%m%dT%H%M%SZ)"

# 3. Capture memory (requires SSM agent or pre-installed tool)
aws ssm send-command \
  --instance-ids <i-xxx> \
  --document-name "AWS-RunShellScript" \
  --parameters 'commands=["avml /tmp/memory.lime && aws s3 cp /tmp/memory.lime s3://ir-bucket/"]'

# 4. Disable instance metadata (prevent further cred theft)
aws ec2 modify-instance-metadata-options \
  --instance-id <i-xxx> \
  --http-endpoint disabled

Investigation

bash
# VPC Flow Logs — find C2 connections
aws logs filter-log-events \
  --log-group-name /aws/vpc/flowlogs \
  --filter-pattern "[version, accountId, interfaceId, srcAddr, dstAddr, srcPort, dstPort, protocol, packets, bytes, start, end, action, logStatus]"

# Check what IAM role the instance had and what it accessed
aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=ResourceName,AttributeValue=<instance-id>

# Analyse EBS snapshot: attach to forensic instance, mount read-only
aws ec2 create-volume --snapshot-id <snap-id> --availability-zone us-east-1a
# Mount: sudo mount -o ro /dev/xvdf1 /mnt/forensics

Recovery

  • Launch replacement from known-good AMI via IaC (Terraform/CloudFormation)
  • Do not restart compromised instance
  • Terminate after forensics complete

Scenario: Compromised Pod in EKS

DetectionGuardDuty EKS findings (Execution:Kubernetes/ExecInKubeSystemPod, PrivilegeEscalation:Kubernetes/...), Falco runtime alerts (shell in container, sensitive mount), unusual egress from a pod, a workload doing AWS API calls it never did before.

How to think about it — the three escape ladders

A compromised pod is dangerous because of what it can reach beyond itself. Investigate along three escalation ladders, because the attacker will try all of them:

LADDER 1 — Pod → AWS (credential theft)
  Does the pod have AWS permissions (IRSA / Pod Identity)?  → attacker assumes that role
  Can the pod reach the NODE's metadata (169.254.169.254)?  → steals the NODE's role
     (far more powerful — the node role often has broad EKS/ECR/EC2 perms)

LADDER 2 — Pod → Kubernetes (cluster escalation)
  What can the pod's SERVICE ACCOUNT do via the K8s API?    → list secrets? create pods?
  Is the service-account token mounted? (default: yes)      → attacker uses it against kube-apiserver
  Any RBAC that lets it read secrets / escalate / exec?     → cluster-wide compromise

LADDER 3 — Pod → Node → other pods (container escape)
  Is the container privileged / hostPID / hostNetwork?      → trivial escape to the node
  hostPath mount of / or the docker socket?                 → own the node, then all pods on it
  Known runtime-escape CVE?                                 → break out of the container

What exactly to check

QuestionWhere to look
What is the pod and what does it run?kubectl describe pod, the image, its command, env vars
Does it have AWS creds?Is a service account annotated for IRSA / associated via Pod Identity? What's the role's policy? Check CloudTrail for that role.
Can it reach node metadata?Is IMDS reachable from pods? (It shouldn't be — block it / enforce IMDSv2 hop-limit 1.) Did the node role make unusual API calls?
What can its service account do?kubectl auth can-i --list --as=system:serviceaccount:<ns>:<sa> — enumerate its RBAC
Is it over-privileged at the container level?securityContext: privileged, allowPrivilegeEscalation, hostPID/hostNetwork/hostPath, capabilities
What did it touch?EKS control-plane audit logs (API calls the SA made), Falco/runtime logs, VPC Flow Logs for egress, CloudTrail for any AWS calls

Containment

bash
# 1. Isolate the pod with a deny-all NetworkPolicy (cut its network)
kubectl label pod <pod> quarantine=true -n <ns>
# apply a NetworkPolicy selecting quarantine=true that allows no ingress/egress

# 2. Do NOT delete the pod yet — you lose live evidence. Cordon the node so
#    nothing reschedules, and capture first.
kubectl cordon <node>

# 3. Capture evidence: running processes, container filesystem, memory if possible
kubectl exec <pod> -n <ns> -- ps aux            # (only if safe; exec is itself logged/risky)
#   better: snapshot the node's EBS volume and analyze offline

# 4. Revoke the pod's identity:
#    - Remove the IRSA role mapping / Pod Identity association
#    - If the K8s service-account token may be stolen, rotate it and review RBAC
#    - If the NODE role may be compromised, treat the whole node as compromised

# 5. Once evidence is captured: delete the pod, and replace the node from a
#    clean image (don't trust a node a container may have escaped to).

Eradication & the key escalation question

The decisive question: did the attacker get off the pod?

Stayed in the pod

delete pod, fix the vuln/image, rotate the pod's identity.

Reached the node (metadata or escape)

the node and its IAM role are compromised; cordon/drain/terminate the node, rotate the node role, and assume every pod that ran on it is suspect.

Reached the K8s API with a powerful service account

potentially cluster-wide; audit what the token did, rotate it, and review for created backdoors (new service accounts, role bindings, daemonsets).

Memory hook

"a pod is a process, not a security boundary." The container shares the node's kernel, so the real questions in an EKS compromise are never "is the container bad" but "what can it reach": the node's IAM role, the Kubernetes API, or the kernel. Lock those three doors in advance — block pod access to IMDS, give each pod a least-privilege service account and IRSA role (never rely on the node role), and forbid privileged/hostPath pods via Pod Security admission — and a single bad pod stays a single bad pod. (See ../../kubernetes/incident-response.md and ../../detection/endpoint-detection.md.)


Scenario: S3 Data Exfiltration

DetectionMacie alert on PII access; CloudTrail GetObject from unusual IP; GuardDuty Exfiltration:S3/ObjectRead.Unusual

bash
# 1. Identify what was accessed (CloudTrail data events must be enabled)
aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=ResourceName,AttributeValue=<bucket-name>

# 2. Block public access immediately if not already
aws s3api put-public-access-block \
  --bucket <bucket-name> \
  --public-access-block-configuration "BlockPublicAcls=true,IgnorePublicAcls=true,BlockPublicPolicy=true,RestrictPublicBuckets=true"

# 3. Revoke the credential used for access (see IAM playbook above)

# 4. Preserve access logs
aws s3 cp s3://<logging-bucket>/ /tmp/s3-logs/ --recursive

# 5. Check bucket policy for unexpected principals
aws s3api get-bucket-policy --bucket <bucket-name>

# 6. Enable Object Lock on sensitive buckets going forward
aws s3api put-object-lock-configuration \
  --bucket <bucket-name> \
  --object-lock-configuration '{"ObjectLockEnabled":"Enabled"}'

Scenario: Lambda Function Compromise

DetectionUnusual invocations, unexpected external calls, IAM calls from Lambda execution role

bash
# 1. Throttle Lambda to stop execution
aws lambda put-function-concurrency \
  --function-name <function-name> \
  --reserved-concurrent-executions 0

# 2. Get current code for forensics
aws lambda get-function --function-name <function-name> --query 'Code.Location'
# Download and analyse

# 3. Check CloudTrail for what the execution role did
aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=Username,AttributeValue=<execution-role-arn>

# 4. Rotate/detach execution role
aws lambda update-function-configuration \
  --function-name <function-name> \
  --role arn:aws:iam::123456789012:role/minimal-read-only-role

# 5. Redeploy from source control after review

GuardDuty Finding → Response Mapping

FindingImmediate Action
UnauthorizedAccess:IAMUser/ConsoleLoginSuccess.BLock user, investigate trail
CryptoCurrency:EC2/BitcoinTool.B!DNSIsolate EC2, snapshot
Recon:EC2/PortProbeUnprotectedPortTighten SG, check if exploit followed
Exfiltration:S3/ObjectRead.UnusualBlock public access, revoke creds
PrivilegeEscalation:IAMUser/AnomalousBehaviorDeny-all policy on user, audit trail
Persistence:IAMUser/NetworkPermissionsAudit new policies/roles, revoke
Backdoor:EC2/C&CActivity.B!DNSIsolate, snapshot, investigate

Useful IR Queries (CloudTrail Insights)

bash
# All root account activity (should be rare)
aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=Username,AttributeValue=root

# Console logins from unusual sources
aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=EventName,AttributeValue=ConsoleLogin

# IAM changes in last 24h
aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=EventSource,AttributeValue=iam.amazonaws.com \
  --start-time $(date -u -d '24 hours ago' +%Y-%m-%dT%H:%M:%SZ)

# New security group rules (potential firewall weakening)
aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=EventName,AttributeValue=AuthorizeSecurityGroupIngress

Interview Questions: AWS IR

Q
A GuardDuty alert fires for CryptoCurrency:EC2/BitcoinTool.B!DNS. Walk me through containment.
Model answer

First isolate the instance by swapping its security groups for a quarantine group with no rules, cutting the C2 and mining traffic while keeping the box alive for forensics. Then preserve evidence before changing anything: snapshot the EBS volumes and, if I can, capture memory via SSM. Critically, I disable the instance metadata endpoint so no further role credentials can be pulled. Then I investigate two tracks in parallel — the host (processes, persistence, how they got in) from the snapshot, and the cloud side, pulling CloudTrail for the instance's IAM role and comparing source IPs to see whether the role's credentials were stolen and used off-box. If they were, it becomes an account-wide credential incident too. Recovery is to rebuild from a known-good AMI via IaC and terminate the compromised instance after forensics — never reboot and reuse it.

Q
How do you forensically preserve an EC2 instance without stopping it?
Model answer

I avoid stopping it because shutting down loses volatile memory and can trigger attacker dead-man's switches. Instead I isolate it at the network layer with a quarantine security group so it can't do harm but stays running. I capture memory live — via the SSM agent running a tool like AVML and shipping the image to an S3 evidence bucket — because RAM holds the most ephemeral evidence: injected code, keys, and C2 state. Then I snapshot the EBS volumes for disk forensics, which I later attach read-only to a separate forensic instance. I also disable IMDS to stop further credential theft. Everything is hashed and logged for chain of custody. The instance keeps running, isolated, until analysis is complete, then it's terminated rather than reused.

Q
An IAM access key leaked on GitHub. What are the first five actions you take?
Model answer

One: deactivate the key — set it inactive, which instantly stops it but is reversible and preserves it as evidence, rather than deleting it outright. Two: scope its activity in CloudTrail filtered to that access key ID — every call, source IP, and timestamp, looking for recon-then-escalation-then-creation patterns. Three: hunt for persistence the attacker created — new IAM users, access keys, roles, or trust-policy changes — because that's what survives a simple key rotation. Four: contain any active abuse, like terminating crypto-mining EC2 they launched, and roll any secrets the key could read. Five: eradicate and harden — delete the key and the backdoor identities, and replace the IAM user with SSO or a role so there's no long-lived key to leak again, plus enable push protection and detections. The theme is that rotating the key alone is never enough.

Q
How would you detect data exfiltration from S3 after the fact?
Model answer

The primary source is CloudTrail S3 data events, which log object-level GetObject calls — but they must be enabled in advance, so step zero is ensuring they're on. With them, I look for anomalous reads: large volumes of GetObject, access from unusual source IPs or principals, or a principal reading objects it never touched before. S3 server access logs are a secondary source. GuardDuty's S3 protection and Macie help — GuardDuty flags unusual object-read patterns, and Macie identifies which buckets hold sensitive data so I can prioritize. I'd also check the bucket policy and ACLs for unexpected principals or public access that enabled the exfil. The honest caveat I'd raise: if data events weren't enabled beforehand, object-level reads may simply not be recorded, which is itself a preparation gap to fix.

Q
What's the difference between deactivating and deleting an access key in an IR context?
Model answer

Deactivating sets the key to inactive — it immediately stops working but still exists, so it's fully reversible and, importantly, preserved as evidence and as a pivot for investigation. Deleting removes it permanently. In incident response you almost always deactivate first, because it achieves instant containment without destroying forensic value: you can still tie the key ID to its CloudTrail history and confirm scope. You delete only after scoping is complete, as part of eradication. Deleting prematurely is a mistake — it's irreversible, and you lose the ability to cleanly correlate the key to its activity.

Q
How do you prevent an attacker from escalating privileges via iam:PassRole?
Model answer

PassRole is the permission to hand a role to a service, and the escalation is that someone who can pass any role plus run compute can launch an instance or Lambda with an admin role and borrow its permissions. The core fix is to scope PassRole tightly: in the identity policy, restrict the Resource to only the specific role ARNs that principal legitimately needs to pass, never a wildcard, and ideally add a condition on iam:PassedToService so a role can only be passed to the intended service. More broadly I treat any IAM-write permission — PassRole, AttachPolicy, CreateAccessKey, UpdateAssumeRolePolicy — as admin-equivalent, since they're paths to admin, and I'd run a tool like PMapper to map who can reach admin and close those edges. Permission boundaries also cap what self-created roles can do.