Kubernetes Incident Response
IR Phases in a K8s Context
Detect → Triage → Contain → Investigate → Eradicate → Recover → Post-mortemDetection Sources
| Source | What It Catches |
|---|---|
| Falco alerts | Runtime anomalies (shell in container, suspicious syscalls) |
| K8s Audit Logs | API server actions (who exec'd, what secrets were accessed) |
| Network policy violations | Unexpected east-west or egress traffic |
| Cloud SIEM (CloudTrail, Chronicle) | Control plane API calls (kubectl from unexpected IP) |
| Image scanner alerts | CVE in running image, new critical CVE published |
| Node metrics | Unexpected CPU spike (crypto mining) |
Triage Checklist
# 1. Identify affected pod/namespace
kubectl get pods --all-namespaces | grep -v Running
# 2. Check recent events
kubectl get events --all-namespaces --sort-by=.lastTimestamp | tail -50
# 3. Check audit logs for suspicious API calls
# (location depends on cluster setup — check /var/log/kube-apiserver-audit.log or cloud logging)
# 4. Who exec'd into pods recently? (from audit log)
grep '"verb":"create".*"resource":"pods/exec"' /var/log/audit.log
# 5. What secrets were accessed?
grep '"resource":"secrets"' /var/log/audit.log | grep '"verb":"get"'
# 6. Unusual outbound connections
kubectl exec -it <pod> -- netstat -an # or ss -an
kubectl exec -it <pod> -- cat /proc/net/tcp
# 7. Unusual processes in pod
kubectl exec -it <pod> -- ps auxf
# 8. Check for modified files
kubectl exec -it <pod> -- find / -newer /tmp -type f 2>/dev/nullMemory hookin K8s IR, the audit log is your CloudTrail. The Kubernetes API server audit log is the single most important evidence source — it records every action against the cluster: who exec'd into which pod, who read which secret, who created a ClusterRoleBinding, and crucially the source IP. The catch (just like AWS CloudTrail data events): audit logging is often not configured to capture the useful detail by default, so "is the audit policy logging secret access and exec at Request level?" is both a preparation question and the first thing you check in an incident. The other reflex: don't
kubectl delete podfirst — the Deployment recreates it and you lose the live evidence; isolate with a deny-all NetworkPolicy and capture before you kill.
Scenario: Compromised Container (Shell Spawned)
DetectionFalco alert — "A shell was run in a container"
Immediate containment
# 1. Isolate pod — apply a "quarantine" network policy (deny all ingress/egress)
kubectl apply -f - <<EOF
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: quarantine-pod
namespace: <namespace>
spec:
podSelector:
matchLabels:
app: <compromised-app>
policyTypes: [Ingress, Egress]
ingress: []
egress: []
EOF
# 2. Capture memory/state before killing
kubectl exec -it <pod> -- ps auxf > /tmp/ps_snapshot.txt
kubectl exec -it <pod> -- netstat -an > /tmp/netstat_snapshot.txt
kubectl exec -it <pod> -- find / -newer /proc/1 -type f 2>/dev/null > /tmp/modified_files.txt
# 3. Capture container filesystem
kubectl cp <pod>:/tmp /tmp/forensics-$(date +%s)
# 4. Get container image ID and check against known good
kubectl get pod <pod> -o jsonpath='{.status.containerStatuses[*].imageID}'
# 5. Kill the pod (Deployment will recreate from known-good image)
kubectl delete pod <pod>Investigation
- Review audit logs for API calls made by the pod's ServiceAccount
- Check if SA token was used from an external IP
- Review network flow logs for C2 connections
- Determine how attacker got shell (exploit? misconfigured exec permission?)
Scenario: Secrets Exfiltration via SA Token
DetectionCloud audit log shows SA token used from unexpected IP/region; many get secret API calls
# Find what SA is associated with the pod
kubectl get pod <pod> -o jsonpath='{.spec.serviceAccountName}'
# Check SA permissions
kubectl auth can-i --list --as=system:serviceaccount:<ns>:<sa>
# Check if token is mounted and has been read
kubectl exec -it <pod> -- cat /var/run/secrets/kubernetes.io/serviceaccount/token
# Rotate SA token (delete old secret, K8s auto-creates new one)
kubectl delete secret <sa-token-secret>
# Disable automounting on the SA going forward
kubectl patch serviceaccount <sa> -p '{"automountServiceAccountToken": false}'Scenario: Crypto Mining (High CPU DaemonSet)
DetectionNode CPU 100%, kubectl top nodes; Falco alert on known miner process names
# Identify the pod consuming resources
kubectl top pods --all-namespaces --sort-by=cpu
# Check process
kubectl exec -it <pod> -- ps auxf | grep -i 'xmrig\|monero\|miner\|stratum'
# Check network connections (mining pool ports 3333, 4444, 14444)
kubectl exec -it <pod> -- ss -tun | grep -E '3333|4444|14444'
# Containment: quarantine, capture, deleteRoot cause analysis
- Was the image compromised? (supply chain)
- Was there a vulnerability exploited (RCE → container compromise)?
- Is
kubectl execaccessible from the internet?
Scenario: Cluster-Admin Privilege Escalation
DetectionAudit log shows creation of ClusterRoleBinding granting cluster-admin
# Check for new/unexpected ClusterRoleBindings
kubectl get clusterrolebindings -o wide | grep cluster-admin
# Check audit log for who created it
grep '"resource":"clusterrolebindings".*"verb":"create"' audit.log
# Who has cluster-admin?
kubectl get clusterrolebinding cluster-admin -o jsonpath='{.subjects}'
# Immediate revocation
kubectl delete clusterrolebinding <malicious-binding>
# Check what actions the escalated account took in the window
grep '"user":{"username":"<attacker>"}' audit.logScenario: Container Escape to Node
DetectionAlert on --privileged container; node-level processes accessed from container; Falco rule triggered
# On the node (via SSH or node debug pod)
# Check for unexpected processes
ps auxf | grep -v '\[' | sort -k3 -rn | head -20
# Check /proc for suspicious mounts
cat /proc/mounts | grep -E 'overlay|tmpfs'
# Look for crypto miners, reverse shells
netstat -antp | grep -E '3333|4444|1337|4545'
# Check for new cron entries, modified /etc/
ls -la /etc/cron*
find /etc -newer /etc/hosts -type f
# Containment: cordon node (stop scheduling), drain, investigate
kubectl cordon <node>
kubectl drain <node> --ignore-daemonsets --delete-emptydir-dataRecoveryTerminate compromised node; replace with fresh from AMI/image.
Audit Log Analysis
Enable K8s audit logging (API server flag):
--audit-log-path=/var/log/kube-apiserver-audit.log
--audit-policy-file=/etc/kubernetes/audit-policy.yaml# audit-policy.yaml — log sensitive operations
apiVersion: audit.k8s.io/v1
kind: Policy
rules:
- level: RequestResponse
resources:
- group: ""
resources: ["secrets", "configmaps"]
- level: Request
verbs: ["create", "update", "patch", "delete"]
- level: Metadata
resources:
- group: ""
resources: ["pods/exec", "pods/portforward"]Key audit log fields
{
"verb": "create",
"user": {"username": "system:serviceaccount:default:compromised-sa"},
"sourceIPs": ["198.51.100.42"],
"objectRef": {"resource": "secrets", "name": "db-creds"},
"responseStatus": {"code": 200}
}Post-Mortem Template
- Timeline — When detected? When did compromise start (check audit log timestamps)?
- Root cause — How did attacker get in? Misconfiguration, vulnerability, supply chain?
- Blast radius — What namespaces, secrets, data were accessed?
- Containment actions taken
- Remediation — What was fixed? New policies applied?
- Detection gap — Why wasn't this caught earlier? What new alert would catch it next time?
- Action items — Owner + deadline for each
Interview Questions: K8s IR
Production containers run one defined process and shouldn't have interactive shells, so I treat it as a likely compromise. I scope first without tipping off: identify the pod, namespace, image, node, and pull what the shell did from Falco and the API audit log. Then I think in escalation ladders — did it reach the pod's cloud credentials or node metadata, did it use the service account token against the API server, and is the pod privileged enough to escape to the node. For containment I apply a deny-all NetworkPolicy to isolate the pod and cordon the node, rather than immediately deleting the pod, so I preserve evidence — and I capture process list, connections, and the filesystem first. Then I revoke the pod's identity, and if it may have reached the node or API I treat those as compromised too. Eradicate by redeploying from a clean image and fixing the entry vector.
The key is to capture before the Deployment recreates the pod and erases evidence. First I isolate it with a deny-all ingress/egress NetworkPolicy so it can't do more harm or phone home while I work. Then I snapshot the volatile state — the process tree, network connections, and recently modified files — via kubectl exec, and copy out the relevant filesystem with kubectl cp, recording the exact image ID for comparison against known-good. If I need deeper forensics, I snapshot the underlying node's disk, since the container's writable layer lives there. Only after capture do I delete the pod. Throughout I avoid rebooting or deleting anything that holds evidence, and I note that exec itself is logged and slightly alters the container, so I document what I ran.
That pattern means the token was exfiltrated and is being replayed from outside the cluster. I find which service account and pod it belongs to and enumerate its permissions with kubectl auth can-i --list, because that tells me the blast radius — can it read secrets, create pods, exec. I pull the audit log filtered to that service account to see exactly what it did and from which IPs. For containment I rotate the token by deleting the token secret so the stolen one stops working, set automountServiceAccountToken false going forward, and tighten the SA's RBAC to least privilege. Then I hunt for what it touched and any persistence it created — new bindings, service accounts, or workloads — and rotate any secrets it accessed. Root cause is usually an app compromise that read the mounted token, so I fix that too.
Detection signals include Falco rules for host namespace access or sensitive mounts, an alert on a privileged or hostPath pod, and node-level anomalies like unexpected processes or outbound connections that correlate with a container. Response: once I suspect the node is compromised, I treat the whole node as hostile, not just the pod. I cordon it to stop new scheduling, capture node-level evidence — processes, connections, cron, modified system files, and ideally a disk snapshot — then drain it. Because a container escape means the attacker had root on the node, I assume every pod that ran there and the node's IAM role are compromised, so I rotate the node role and rebuild the node from a clean image rather than trusting it. Recovery is replace, not repair.
The clearest is creation or modification of RBAC that grants broad power — a new ClusterRoleBinding to cluster-admin, or RoleBindings handing a service account elevated verbs. Also: use of the impersonate verb, create on pods with privileged or hostPath specs, create on pods/exec into sensitive pods, and bulk get on secrets. I'd look at the verb, the resource, the user or service account, and the source IP together — for instance a service account that normally only reads its own namespace suddenly creating cluster-wide bindings from an external IP. The audit log's RequestResponse level on RBAC and secrets is what makes these visible, which is why configuring the audit policy to capture them in advance matters.
I separate network containment from destruction. I can contain fast without losing evidence by applying a deny-all NetworkPolicy to isolate the pod and cordoning the node — that stops C2, exfiltration, and lateral movement immediately while the pod stays alive for capture. What destroys evidence is deleting the pod, which triggers recreation, so I do that only after snapshotting processes, connections, filesystem, and ideally the node disk. The exception is active harm — if it's exfiltrating data or spreading right now, containment wins and I accept evidence loss. And I scope before I broadly remediate, because revoking one thing prematurely can tip off the attacker before I've mapped the blast radius. So the rule is isolate-then-capture-then-eradicate, with active-harm as the override.