Security Notes
Kubernetes

Kubernetes Incident Response

9 min read 11 sections 6 model answers

IR Phases in a K8s Context

Detect → Triage → Contain → Investigate → Eradicate → Recover → Post-mortem

Detection Sources

SourceWhat It Catches
Falco alertsRuntime anomalies (shell in container, suspicious syscalls)
K8s Audit LogsAPI server actions (who exec'd, what secrets were accessed)
Network policy violationsUnexpected east-west or egress traffic
Cloud SIEM (CloudTrail, Chronicle)Control plane API calls (kubectl from unexpected IP)
Image scanner alertsCVE in running image, new critical CVE published
Node metricsUnexpected CPU spike (crypto mining)

Triage Checklist

bash
# 1. Identify affected pod/namespace
kubectl get pods --all-namespaces | grep -v Running

# 2. Check recent events
kubectl get events --all-namespaces --sort-by=.lastTimestamp | tail -50

# 3. Check audit logs for suspicious API calls
# (location depends on cluster setup — check /var/log/kube-apiserver-audit.log or cloud logging)

# 4. Who exec'd into pods recently? (from audit log)
grep '"verb":"create".*"resource":"pods/exec"' /var/log/audit.log

# 5. What secrets were accessed?
grep '"resource":"secrets"' /var/log/audit.log | grep '"verb":"get"'

# 6. Unusual outbound connections
kubectl exec -it <pod> -- netstat -an     # or ss -an
kubectl exec -it <pod> -- cat /proc/net/tcp

# 7. Unusual processes in pod
kubectl exec -it <pod> -- ps auxf

# 8. Check for modified files
kubectl exec -it <pod> -- find / -newer /tmp -type f 2>/dev/null

Memory hook

in K8s IR, the audit log is your CloudTrail. The Kubernetes API server audit log is the single most important evidence source — it records every action against the cluster: who exec'd into which pod, who read which secret, who created a ClusterRoleBinding, and crucially the source IP. The catch (just like AWS CloudTrail data events): audit logging is often not configured to capture the useful detail by default, so "is the audit policy logging secret access and exec at Request level?" is both a preparation question and the first thing you check in an incident. The other reflex: don't kubectl delete pod first — the Deployment recreates it and you lose the live evidence; isolate with a deny-all NetworkPolicy and capture before you kill.

Scenario: Compromised Container (Shell Spawned)

DetectionFalco alert — "A shell was run in a container"

Immediate containment

bash
# 1. Isolate pod — apply a "quarantine" network policy (deny all ingress/egress)
kubectl apply -f - <<EOF
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: quarantine-pod
  namespace: <namespace>
spec:
  podSelector:
    matchLabels:
      app: <compromised-app>
  policyTypes: [Ingress, Egress]
  ingress: []
  egress: []
EOF

# 2. Capture memory/state before killing
kubectl exec -it <pod> -- ps auxf > /tmp/ps_snapshot.txt
kubectl exec -it <pod> -- netstat -an > /tmp/netstat_snapshot.txt
kubectl exec -it <pod> -- find / -newer /proc/1 -type f 2>/dev/null > /tmp/modified_files.txt

# 3. Capture container filesystem
kubectl cp <pod>:/tmp /tmp/forensics-$(date +%s)

# 4. Get container image ID and check against known good
kubectl get pod <pod> -o jsonpath='{.status.containerStatuses[*].imageID}'

# 5. Kill the pod (Deployment will recreate from known-good image)
kubectl delete pod <pod>

Investigation

  • Review audit logs for API calls made by the pod's ServiceAccount
  • Check if SA token was used from an external IP
  • Review network flow logs for C2 connections
  • Determine how attacker got shell (exploit? misconfigured exec permission?)

Scenario: Secrets Exfiltration via SA Token

DetectionCloud audit log shows SA token used from unexpected IP/region; many get secret API calls

bash
# Find what SA is associated with the pod
kubectl get pod <pod> -o jsonpath='{.spec.serviceAccountName}'

# Check SA permissions
kubectl auth can-i --list --as=system:serviceaccount:<ns>:<sa>

# Check if token is mounted and has been read
kubectl exec -it <pod> -- cat /var/run/secrets/kubernetes.io/serviceaccount/token

# Rotate SA token (delete old secret, K8s auto-creates new one)
kubectl delete secret <sa-token-secret>

# Disable automounting on the SA going forward
kubectl patch serviceaccount <sa> -p '{"automountServiceAccountToken": false}'

Scenario: Crypto Mining (High CPU DaemonSet)

DetectionNode CPU 100%, kubectl top nodes; Falco alert on known miner process names

bash
# Identify the pod consuming resources
kubectl top pods --all-namespaces --sort-by=cpu

# Check process
kubectl exec -it <pod> -- ps auxf | grep -i 'xmrig\|monero\|miner\|stratum'

# Check network connections (mining pool ports 3333, 4444, 14444)
kubectl exec -it <pod> -- ss -tun | grep -E '3333|4444|14444'

# Containment: quarantine, capture, delete

Root cause analysis

  • Was the image compromised? (supply chain)
  • Was there a vulnerability exploited (RCE → container compromise)?
  • Is kubectl exec accessible from the internet?

Scenario: Cluster-Admin Privilege Escalation

DetectionAudit log shows creation of ClusterRoleBinding granting cluster-admin

bash
# Check for new/unexpected ClusterRoleBindings
kubectl get clusterrolebindings -o wide | grep cluster-admin

# Check audit log for who created it
grep '"resource":"clusterrolebindings".*"verb":"create"' audit.log

# Who has cluster-admin?
kubectl get clusterrolebinding cluster-admin -o jsonpath='{.subjects}'

# Immediate revocation
kubectl delete clusterrolebinding <malicious-binding>

# Check what actions the escalated account took in the window
grep '"user":{"username":"<attacker>"}' audit.log

Scenario: Container Escape to Node

DetectionAlert on --privileged container; node-level processes accessed from container; Falco rule triggered

bash
# On the node (via SSH or node debug pod)
# Check for unexpected processes
ps auxf | grep -v '\[' | sort -k3 -rn | head -20

# Check /proc for suspicious mounts
cat /proc/mounts | grep -E 'overlay|tmpfs'

# Look for crypto miners, reverse shells
netstat -antp | grep -E '3333|4444|1337|4545'

# Check for new cron entries, modified /etc/
ls -la /etc/cron*
find /etc -newer /etc/hosts -type f

# Containment: cordon node (stop scheduling), drain, investigate
kubectl cordon <node>
kubectl drain <node> --ignore-daemonsets --delete-emptydir-data

RecoveryTerminate compromised node; replace with fresh from AMI/image.


Audit Log Analysis

Enable K8s audit logging (API server flag):

yaml
--audit-log-path=/var/log/kube-apiserver-audit.log
--audit-policy-file=/etc/kubernetes/audit-policy.yaml
yaml
# audit-policy.yaml — log sensitive operations
apiVersion: audit.k8s.io/v1
kind: Policy
rules:
  - level: RequestResponse
    resources:
      - group: ""
        resources: ["secrets", "configmaps"]
  - level: Request
    verbs: ["create", "update", "patch", "delete"]
  - level: Metadata
    resources:
      - group: ""
        resources: ["pods/exec", "pods/portforward"]

Key audit log fields

json
{
  "verb": "create",
  "user": {"username": "system:serviceaccount:default:compromised-sa"},
  "sourceIPs": ["198.51.100.42"],
  "objectRef": {"resource": "secrets", "name": "db-creds"},
  "responseStatus": {"code": 200}
}

Post-Mortem Template

  1. Timeline — When detected? When did compromise start (check audit log timestamps)?
  2. Root cause — How did attacker get in? Misconfiguration, vulnerability, supply chain?
  3. Blast radius — What namespaces, secrets, data were accessed?
  4. Containment actions taken
  5. Remediation — What was fixed? New policies applied?
  6. Detection gap — Why wasn't this caught earlier? What new alert would catch it next time?
  7. Action items — Owner + deadline for each

Interview Questions: K8s IR

Q
Falco fires "shell in prod container" at 3am. Walk me through your response.
Model answer

Production containers run one defined process and shouldn't have interactive shells, so I treat it as a likely compromise. I scope first without tipping off: identify the pod, namespace, image, node, and pull what the shell did from Falco and the API audit log. Then I think in escalation ladders — did it reach the pod's cloud credentials or node metadata, did it use the service account token against the API server, and is the pod privileged enough to escape to the node. For containment I apply a deny-all NetworkPolicy to isolate the pod and cordon the node, rather than immediately deleting the pod, so I preserve evidence — and I capture process list, connections, and the filesystem first. Then I revoke the pod's identity, and if it may have reached the node or API I treat those as compromised too. Eradicate by redeploying from a clean image and fixing the entry vector.

Q
How do you forensically capture state from a pod before killing it?
Model answer

The key is to capture before the Deployment recreates the pod and erases evidence. First I isolate it with a deny-all ingress/egress NetworkPolicy so it can't do more harm or phone home while I work. Then I snapshot the volatile state — the process tree, network connections, and recently modified files — via kubectl exec, and copy out the relevant filesystem with kubectl cp, recording the exact image ID for comparison against known-good. If I need deeper forensics, I snapshot the underlying node's disk, since the container's writable layer lives there. Only after capture do I delete the pod. Throughout I avoid rebooting or deleting anything that holds evidence, and I note that exec itself is logged and slightly alters the container, so I document what I ran.

Q
A service account token was used from an unexpected IP. How do you contain and investigate?
Model answer

That pattern means the token was exfiltrated and is being replayed from outside the cluster. I find which service account and pod it belongs to and enumerate its permissions with kubectl auth can-i --list, because that tells me the blast radius — can it read secrets, create pods, exec. I pull the audit log filtered to that service account to see exactly what it did and from which IPs. For containment I rotate the token by deleting the token secret so the stolen one stops working, set automountServiceAccountToken false going forward, and tighten the SA's RBAC to least privilege. Then I hunt for what it touched and any persistence it created — new bindings, service accounts, or workloads — and rotate any secrets it accessed. Root cause is usually an app compromise that read the mounted token, so I fix that too.

Q
How do you detect and respond to a container escape to the node?
Model answer

Detection signals include Falco rules for host namespace access or sensitive mounts, an alert on a privileged or hostPath pod, and node-level anomalies like unexpected processes or outbound connections that correlate with a container. Response: once I suspect the node is compromised, I treat the whole node as hostile, not just the pod. I cordon it to stop new scheduling, capture node-level evidence — processes, connections, cron, modified system files, and ideally a disk snapshot — then drain it. Because a container escape means the attacker had root on the node, I assume every pod that ran there and the node's IAM role are compromised, so I rotate the node role and rebuild the node from a clean image rather than trusting it. Recovery is replace, not repair.

Q
What audit log entries indicate a privilege escalation attempt?
Model answer

The clearest is creation or modification of RBAC that grants broad power — a new ClusterRoleBinding to cluster-admin, or RoleBindings handing a service account elevated verbs. Also: use of the impersonate verb, create on pods with privileged or hostPath specs, create on pods/exec into sensitive pods, and bulk get on secrets. I'd look at the verb, the resource, the user or service account, and the source IP together — for instance a service account that normally only reads its own namespace suddenly creating cluster-wide bindings from an external IP. The audit log's RequestResponse level on RBAC and secrets is what makes these visible, which is why configuring the audit policy to capture them in advance matters.

Q
How do you balance "contain quickly" versus "preserve evidence" in K8s IR?
Model answer

I separate network containment from destruction. I can contain fast without losing evidence by applying a deny-all NetworkPolicy to isolate the pod and cordoning the node — that stops C2, exfiltration, and lateral movement immediately while the pod stays alive for capture. What destroys evidence is deleting the pod, which triggers recreation, so I do that only after snapshotting processes, connections, filesystem, and ideally the node disk. The exception is active harm — if it's exfiltrating data or spreading right now, containment wins and I accept evidence loss. And I scope before I broadly remediate, because revoking one thing prematurely can tip off the attacker before I've mapped the blast radius. So the rule is isolate-then-capture-then-eradicate, with active-harm as the override.