GCP Incident Response Playbooks
Detection Sources
| Source | What It Catches |
|---|---|
| SCC Event Threat Detection | Real-time threats from audit logs |
| SCC Container Threat Detection | GKE runtime anomalies |
| Cloud Audit Logs (Admin Activity) | All write operations, IAM changes |
| Cloud Audit Logs (Data Access) | Secret reads, storage access (opt-in) |
| VPC Flow Logs | Network traffic, unusual connections |
| Chronicle SIEM | Correlation across all log sources |
| Cloud Monitoring alerts | Resource anomalies (CPU spike = mining) |
Scenario: Compromised Service Account Key
DetectionSCC CREDENTIAL_ACCESS_GET_SECRET_VALUE from unusual IP; Chronicle alert on SA used from non-corp location
Memory hookthe SA key is GCP's leaked-AWS-key, and "disable then delete" is the same drill. A
...iam.gserviceaccount.comJSON key is the GCP equivalent of a leakedAKIAaccess key — long-lived, found-by-bots-in-seconds, and good until you act. The response rhymes with AWS: disable the key first (reversible, preserves evidence) → scope what it did in Cloud Audit Logs → hunt for the persistence it created (new SA keys, new SAs, new IAM bindings) → then delete + rotate. Two GCP specifics: (1) use an IAM Deny policy for instant org-wide lockout of the principal — it's GCP's "explicit deny always wins"; (2) the #1 missed step is the same as AWS — rotating the one key while leaving the attacker's newly-minted SA keys and bindings in place. The deeper fix is Org PolicydisableServiceAccountKeyCreation+ Workload Identity so there's no key to leak next time.
Containment
# 1. Identify and disable the compromised key
gcloud iam service-accounts keys list \
--iam-account=<sa>@<project>.iam.gserviceaccount.com
gcloud iam service-accounts keys disable <key-id> \
--iam-account=<sa>@<project>.iam.gserviceaccount.com
# 2. Deny all actions from this SA immediately (org policy won't help here — use IAM deny)
# Add a deny policy binding (IAM Deny, GA in 2022+)
gcloud iam deny-policies create incident-lockout \
--attachment-point="cloudresourcemanager.googleapis.com/projects/<project>" \
--policy-file=deny-sa-all.json
# deny-sa-all.json: deny all actions for the SA principal
# 3. Alternatively: remove all roles from the SA
gcloud projects get-iam-policy <project> --format=json > policy.json
# Edit policy.json to remove SA bindings
gcloud projects set-iam-policy <project> policy.jsonInvestigation
# What did the SA do in the last 24h?
gcloud logging read \
'protoPayload.authenticationInfo.principalEmail="<sa>@<project>.iam.gserviceaccount.com"' \
--freshness=24h \
--format=json | jq '[.[] | {time: .timestamp, method: .protoPayload.methodName, resource: .resource.labels}]'
# Check for new SA keys created (persistence)
gcloud logging read \
'protoPayload.methodName="google.iam.admin.v1.CreateServiceAccountKey"' \
--freshness=7d
# Check for new IAM bindings added
gcloud logging read \
'protoPayload.methodName="SetIamPolicy"' \
--freshness=24h
# What data was accessed in GCS?
gcloud logging read \
'protoPayload.authenticationInfo.principalEmail="<sa>@..." AND protoPayload.serviceName="storage.googleapis.com"' \
--freshness=24hEradication
# Delete compromised key
gcloud iam service-accounts keys delete <key-id> \
--iam-account=<sa>@<project>.iam.gserviceaccount.com
# Delete any backdoor SAs created during the compromise
gcloud iam service-accounts list --filter="email:<suspicious>"
gcloud iam service-accounts delete <backdoor-sa>
# Rotate secrets accessed
gcloud secrets versions add <secret-name> --data-file=./new-secret.txt
gcloud secrets versions disable <old-version> --secret=<secret-name>Scenario: Crypto Mining on GCE
DetectionSCC CRYPTO_MINING_EXECUTION (VM Threat Detection); Cloud Monitoring CPU alert; VPC Flow Logs showing connections to mining pool ports
# 1. Identify the instance
gcloud compute instances list --filter="status:RUNNING" --format="table(name,zone,machineType,status)"
# 2. Snapshot disk for forensics (before stopping)
DISK=$(gcloud compute instances describe <instance> --zone=<zone> \
--format="value(disks[0].source)")
gcloud compute disks snapshot $DISK \
--snapshot-names="ir-forensics-$(date +%Y%m%d)" \
--zone=<zone>
# 3. Isolate instance — remove from load balancer, apply deny-all firewall rule
gcloud compute firewall-rules create ir-quarantine \
--direction=INGRESS \
--action=DENY \
--rules=all \
--target-tags=quarantine \
--priority=0 # highest priority
gcloud compute instances add-tags <instance> \
--tags=quarantine --zone=<zone>
# Block all egress too
gcloud compute firewall-rules create ir-quarantine-egress \
--direction=EGRESS \
--action=DENY \
--rules=all \
--target-tags=quarantine \
--priority=0
# 4. Capture runtime state via SSH (or OS Login)
gcloud compute ssh <instance> --zone=<zone> -- "ps auxf > /tmp/ps.txt && netstat -an > /tmp/netstat.txt"
gcloud compute scp <instance>:/tmp/ps.txt ./forensics/ --zone=<zone>
# 5. Stop instance, attach disk to forensic instance
gcloud compute instances stop <instance> --zone=<zone>Root Cause Investigation
- Review VPC Flow Logs for initial access (unusual SSH, exploit traffic)
- Check metadata server access logs: was the default SA used?
- Review GCE startup scripts, custom images
- Check if instance was exposed on 0.0.0.0/0 (SCC
OPEN_FIREWALLfinding)
Scenario: GCS Bucket Data Exfiltration
DetectionSCC PUBLIC_BUCKET_ACL; unusual storage.objects.list + storage.objects.get from external IP; Macie-equivalent data classification alert
# 1. Immediately remove public access
gsutil iam ch -d allUsers gs://<bucket-name>
gsutil iam ch -d allAuthenticatedUsers gs://<bucket-name>
# Enable uniform bucket-level access (disables legacy ACLs)
gsutil uniformbucketlevelaccess set on gs://<bucket-name>
# 2. Investigate who accessed what
gcloud logging read \
'protoPayload.resourceName:"projects/_/buckets/<bucket-name>"
AND protoPayload.methodName=("storage.objects.get" OR "storage.objects.list")' \
--freshness=7d \
--format=json | jq '[.[] | {time: .timestamp, caller: .protoPayload.authenticationInfo.principalEmail, ip: .httpRequest.remoteIp, object: .protoPayload.resourceName}]'
# 3. Check bucket ACL and IAM history
gsutil iam get gs://<bucket-name>
gcloud logging read 'protoPayload.methodName="storage.setIamPermissions"' \
--freshness=30d
# 4. Enable data access audit logging for storage (to catch future reads)
gcloud projects get-iam-policy <project> \
--flatten="auditConfigs()" \
--filter="auditConfigs.service=storage.googleapis.com"Scenario: GKE Cluster Compromise
See kubernetes/incident-response.md for pod-level playbooks.
GKE-specific actions
# Check GKE audit logs
gcloud logging read \
'resource.type="k8s_cluster"
AND protoPayload.methodName=("io.k8s.core.v1.pods.exec" OR "io.k8s.rbac.v1.clusterrolebindings.create")' \
--freshness=24h
# Check node pool versions
gcloud container clusters describe <cluster> --zone=<zone> \
--format="value(nodePools[].version)"
# Enable Workload Identity if not enabled
gcloud container clusters update <cluster> \
--workload-pool=<project>.svc.id.goog
# Enable Binary Authorization
gcloud container clusters update <cluster> \
--binauthz-evaluation-mode=PROJECT_SINGLETON_POLICY_ENFORCESCC Finding → Response Mapping
| Finding | Immediate Action |
|---|---|
PUBLIC_BUCKET_ACL | Remove allUsers/allAuthenticatedUsers ACL |
OPEN_FIREWALL | Restrict firewall rule to known CIDR |
CRYPTO_MINING_EXECUTION | Isolate GCE instance, snapshot, investigate |
PERSISTENCE_IAM_ANOMALOUS_GRANT | Audit and revoke anomalous IAM binding |
CREDENTIAL_ACCESS_GET_SECRET_VALUE | Rotate secret, investigate SA |
INITIAL_ACCESS_LOG4J_BAD_USER_AGENT | Patch, isolate, check for post-exploit |
EXFILTRATION_BIGQUERY_DATA_EXFILTRATION | Revoke SA, check VPC SC gaps |
CONTAINER_RUNTIME_UNEXPECTED_CHILD_SHELL | Quarantine pod, investigate |
Useful Audit Queries
# Root/super-admin activity
gcloud logging read \
'protoPayload.authenticationInfo.principalEmail:("@yourdomain.com")
AND protoPayload.authorizationInfo.granted=true
AND resource.type="project"' \
--freshness=24h
# Service accounts creating other service accounts (persistence)
gcloud logging read \
'protoPayload.methodName="google.iam.admin.v1.CreateServiceAccount"' \
--freshness=7d
# Firewall rule changes
gcloud logging read \
'protoPayload.methodName:("compute.firewalls.insert" OR "compute.firewalls.patch")' \
--freshness=7d
# Secret access from outside expected IPs
gcloud logging read \
'protoPayload.serviceName="secretmanager.googleapis.com"
AND protoPayload.methodName="google.cloud.secretmanager.v1.SecretManagerService.AccessSecretVersion"' \
--freshness=24hChronicle SIEM Investigation
# YARA-L: detect anomalous SA key creation
rule sa_key_persistence {
meta:
severity = "HIGH"
events:
$e.metadata.event_type = "USER_RESOURCE_CREATION"
$e.target.resource.type = "SERVICE_ACCOUNT_KEY"
not $e.principal.ip in %corp_ip_ranges
condition:
$e
}
# YARA-L: detect public bucket grant
rule public_bucket_grant {
events:
$e.metadata.event_type = "USER_RESOURCE_UPDATE_CONTENT"
$e.target.resource.type = "GCS_BUCKET"
$e.target.resource.attribute.labels["binding_member"] = "allUsers"
condition:
$e
}Interview Questions: GCP IR
Isolate, preserve, investigate. First I isolate the VM by applying a highest-priority deny-all firewall rule on both ingress and egress via a quarantine tag, cutting the mining traffic and any C2 while keeping the box alive. Before stopping it I snapshot the disk for forensics and capture live state — process list and connections — over SSH. Then I investigate two tracks like an EC2 compromise: the host, for how they got in and what they ran, and the cloud side, checking whether the VM's service account token was used — especially if it's the over-privileged default SA — by querying Cloud Audit Logs for that principal. I look at VPC Flow Logs for the initial access vector and whether the instance was exposed to 0.0.0.0/0. Recovery is rebuild from a clean image and terminate the compromised VM; if the SA was used off-box, it also becomes a credential incident.
Several layers, fast first. If a key is involved, disable it immediately — reversible and preserves it as evidence. For instant, decisive lockout of the principal regardless of its grants, I attach an IAM Deny policy denying all actions for that service account, which overrides any allow, GCP's explicit-deny-wins. I can also strip its role bindings from the project IAM policy. Then I scope what it did in Cloud Audit Logs and, crucially, hunt for persistence it created — new service account keys, new service accounts, or new IAM bindings — because just disabling the one key leaves backdoors. Finally I delete the compromised key, remove backdoor identities, rotate any secrets it accessed, and prevent recurrence with Org Policy disabling key creation plus Workload Identity.
First contain by removing the allUsers and allAuthenticatedUsers bindings and enabling uniform bucket-level access. Then scope: I query Cloud Audit Logs for object access on that bucket — storage.objects.get and list — over the exposure window, extracting the caller, source IP, and which objects were read, to distinguish external readers from normal internal traffic. The catch is that object-level reads are Data Access logs, which are opt-in, so if they weren't enabled I may not have read visibility, in which case I fall back to what I do have and treat unknown exposure conservatively. I also pull the IAM-change history to see who made it public and when, to fix root cause. If sensitive data was confirmed read by external parties, it escalates to a data-breach response with notification considerations, and I close the loop with Org Policy to block public buckets and an alert for next time.
VPC Service Controls puts a perimeter around the storage resources so that API requests crossing the boundary are denied even when the caller has valid IAM permission. In the public-bucket or stolen-credential case, the attacker copying objects out to an external or personal GCP project would be blocked at the perimeter, because the destination is outside it — IAM said yes, but the perimeter says the data can't leave. It specifically defends the gap that IAM can't: a legitimately-authorized identity, or a leaked credential, exfiltrating data. You'd sanction real exceptions with Access Levels for corp IPs or trusted service accounts. So it converts "anyone with read access can pull the data anywhere" into "the data physically cannot leave the trust boundary."
Chronicle, now Google SecOps, is a cloud-native SIEM built on Google's infrastructure, so it's designed for petabyte-scale ingestion with flat-rate pricing rather than the per-GB licensing that makes traditional SIEMs costly at volume. It normalizes everything into a Unified Data Model so detections work across sources, uses YARA-L as its detection language, and natively ingests GCP telemetry — Cloud Audit Logs, VPC Flow, GKE audit — plus third-party data. A standout capability is retroactive analysis: because it retains large volumes affordably, you can apply a brand-new detection rule against a year of historical data to find past activity, which traditional SIEMs struggle to do. It's the data-lake-style approach applied to security, which is exactly the modern detection architecture a large GCP shop favors.
Whether it's legitimate or an escalation, by examining the binding in context: who was granted what role on what resource, and who granted it. The red flags are a high-privilege or impersonation-enabling role — Owner, securityAdmin, or serviceAccountTokenCreator — granted to an unexpected principal, especially an external account or a freshly created service account, and a grantor who is itself suspicious or operating from an unusual IP. I'd pull the SetIamPolicy event, identify the actor and their source, and check what else that actor did around the same time — did they create a service account or key first, suggesting a privilege-escalation chain. If it looks malicious I scope the actor's full activity before reverting, so I don't tip them off prematurely, then revoke the binding and any persistence. The instinct is treat an out-of-band IAM grant as potential escalation until proven a legitimate change.