Kubernetes Security
RBAC (Role-Based Access Control)
Objects
| Object | Scope | Description |
|---|---|---|
| Role | Namespace | Grants permissions within a namespace |
| ClusterRole | Cluster | Grants cluster-wide or namespace permissions |
| RoleBinding | Namespace | Binds a Role/ClusterRole to subjects in a namespace |
| ClusterRoleBinding | Cluster | Binds a ClusterRole to subjects cluster-wide |
Subjects
User— external human user (authenticated via OIDC, certs, etc.)Group— group of usersServiceAccount— pod identity within the cluster
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: pod-reader
namespace: default
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "watch", "list"]
---
kind: RoleBinding
metadata:
name: read-pods
namespace: default
subjects:
- kind: ServiceAccount
name: my-app-sa
namespace: default
roleRef:
kind: Role
name: pod-reader
apiGroup: rbac.authorization.k8s.ioDangerous RBAC Permissions
| Permission | Risk |
|---|---|
* on * (wildcard) | Full cluster control |
create pods | Can run privileged containers |
create/update rolebindings | Privilege escalation — self-assign roles |
exec on pods | Lateral movement into running workloads |
get secrets | Access all secrets in namespace |
impersonate | Act as any user or service account |
patch/update deployments | Inject malicious containers |
Audit command
kubectl auth can-i --list --as=system:serviceaccount:default:my-saMemory hookRole/ClusterRole = the verb list; Binding = who gets it. A Role/ClusterRole is just a set of permissions ("get/list pods") with no one attached — it's a job description, not a hiring. A RoleBinding/ClusterRoleBinding is what actually grants it to a user, group, or service account. The Role/ClusterRole split is namespace vs cluster-wide scope; the binding split is the same. The interview trap: a ClusterRole used in a RoleBinding grants its permissions only within that one namespace — so you reuse one ClusterRole across many namespaces via per-namespace RoleBindings. Mnemonic: Role = what, Binding = who, Cluster prefix = where (everywhere).
Pod Security Standards (PSS)
Replaced PodSecurityPolicy (PSP) in Kubernetes 1.25. Applied via namespace labels.
| Profile | Description |
|---|---|
privileged | Unrestricted (avoid for production) |
baseline | Prevents known privilege escalations |
restricted | Hardened; requires non-root, no capabilities, seccomp |
# Apply restricted policy to namespace
apiVersion: v1
kind: Namespace
metadata:
name: prod
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/audit: restricted
pod-security.kubernetes.io/warn: restrictedSecure Pod Spec
securityContext:
runAsNonRoot: true
runAsUser: 10001
readOnlyRootFilesystem: true
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
seccompProfile:
type: RuntimeDefaultNetwork Policies
By default, all pods can communicate with all pods. Network Policies restrict this.
# Deny all ingress to 'payments' namespace, then selectively allow
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: deny-all-ingress
namespace: payments
spec:
podSelector: {} # applies to all pods
policyTypes: [Ingress]
ingress: [] # empty = deny all
---
# Allow only from 'api' namespace
kind: NetworkPolicy
metadata:
name: allow-from-api
namespace: payments
spec:
podSelector:
matchLabels:
app: payment-svc
policyTypes: [Ingress]
ingress:
- from:
- namespaceSelector:
matchLabels:
name: api
ports:
- port: 8443NoteNetwork Policies require a supporting CNI plugin (Calico, Cilium, Weave). Flannel does not enforce them.
Admission Controllers
Webhook-based plugins that intercept API server requests after authentication/authorization but before persisting to etcd.
| Controller | Purpose |
|---|---|
MutatingAdmissionWebhook | Modify resources (inject sidecars, add labels) |
ValidatingAdmissionWebhook | Reject non-compliant resources |
PodSecurity | Enforce Pod Security Standards |
ImagePolicyWebhook | Allow/deny images based on policy |
NodeRestriction | Limits kubelet permissions |
OPA / Gatekeeper
Open Policy Agent provides policy-as-code for Kubernetes:
# Deny privileged containers
package k8s.deny_privileged
violation[{"msg": msg}] {
container := input.review.object.spec.containers[_]
container.securityContext.privileged == true
msg := sprintf("Privileged container forbidden: %v", [container.name])
}Kyverno
Kubernetes-native policy engine (no Rego needed):
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: disallow-privileged
spec:
validationFailureAction: Enforce
rules:
- name: check-privileged
match:
resources:
kinds: [Pod]
validate:
message: "Privileged mode is not allowed."
pattern:
spec:
containers:
- =(securityContext):
=(privileged): falseSecrets Management
| Approach | Security Level | Notes |
|---|---|---|
| Native K8s Secrets (default) | Low | base64 only; etcd unencrypted |
| etcd encryption at rest | Medium | AES-GCM; protects etcd backups | | Sealed Secrets (Bitnami) | Medium | Encrypted at rest in git; decrypted by controller | | External Secrets Operator | High | Syncs from AWS SM, GCP SM, Vault | | Vault Agent Injector | High | Injects dynamic secrets into pod env at runtime | | CSI Secret Store Driver | High | Mounts secrets as volume from external provider |
Memory hooka K8s Secret is "base64, not encrypted." The #1 Kubernetes security misconception: a
Secretis not secret by default. It's just base64-encoded (trivially decoded) and stored in etcd in plaintext unless you explicitly enable encryption-at-rest. Anyone who can read the Secret object, read etcd, or read an etcd backup has the cleartext. The fixes, in increasing strength: (1) enable etcd encryption-at-rest, (2) restrict RBACget secretstightly, (3) push secrets to an external manager (Vault, AWS Secrets Manager via External Secrets Operator / CSI driver) so the real secret never lives in etcd at all. The one-liner: "K8s Secrets are obfuscated, not encrypted — turn on etcd encryption and prefer an external secrets store."
Enable etcd encryption
# /etc/kubernetes/encryption-config.yaml
apiVersion: apiserver.config.k8s.io/v1
kind: EncryptionConfiguration
resources:
- resources: [secrets]
providers:
- aescbc:
keys:
- name: key1
secret: <base64-encoded-32-byte-key>
- identity: {}Runtime Security with Falco
Falco detects anomalous behaviour in running containers using kernel syscall inspection.
Example rules
- rule: Shell in container
desc: A shell was spawned in a container
condition: container.id != host and proc.name = bash
output: "Shell in container (user=%user.name container=%container.name)"
priority: WARNING
- rule: Write to /etc in container
condition: container and open_write and fd.directory startswith /etc
output: "Write to /etc in container: %container.name %proc.cmdline"
priority: ERRORWhat Falco detects
- Shell spawned in production container
- Unexpected outbound connections from container
- Sensitive file reads (
/etc/shadow,/etc/kubernetes/admin.conf) - Privilege escalation attempts (setuid, capability changes)
- Crypto mining process names
Container Escape Paths (Interview Must-Know)
Memory hook"a pod is a process, not a VM." Every escape below exploits the fact that containers share the host kernel — they're isolated processes, not separate machines. So the escape routes are the bits of isolation you removed:
privileged(drop the walls), a mounted docker/containerd socket (ask the host to run a container for you), hostPID/hostNetwork/hostPath (share a host namespace or the host filesystem), a dangerous capability likeCAP_SYS_ADMIN, or a kernel/runtime CVE (one shared kernel). Defenders flip every row: non-root, drop all caps, no privileged, no host namespaces, no socket mounts — enforced by Pod Securityrestricted+ admission policy.
| Technique | Condition | Notes |
|---|---|---|
--privileged | Container runs privileged | Mount host /dev/sda → chroot into host |
docker.sock mounted | /var/run/docker.sock in container | Create new privileged container → escape |
hostPID: true | Pod has host PID namespace | See host processes, inject into them |
hostNetwork: true | Pod has host network namespace | Reach host services, bypass network policies |
CAP_SYS_ADMIN | Container has this capability | Many kernel primitives available |
| Cgroups release_agent | Writable /sys/fs/cgroup | Classic CVE-2022-0492 container escape |
| runc CVE-2019-5736 | Runnable PoC overwrites runc binary | Requires container exec |
Supply Chain Security in K8s
Binary Authorization (GKE) / Image Policy
- Only signed images from trusted registries can run
- Sigstore/cosign attestation enforced at admission
SBOM (Software Bill of Materials)
syftgenerates SBOM from imagegrypescans SBOM for CVEs- Attach SBOM as OCI artifact:
cosign attest --type cyclonedx
Key Security Audit Checks
# Find pods running as root
kubectl get pods --all-namespaces -o jsonpath='{range .items[*]}{.metadata.namespace}{"\t"}{.metadata.name}{"\t"}{.spec.containers[*].securityContext.runAsUser}{"\n"}{end}'
# Find privileged pods
kubectl get pods --all-namespaces -o json | jq '.items[] | select(.spec.containers[].securityContext.privileged==true) | .metadata.name'
# Find SA tokens automounted
kubectl get pods --all-namespaces -o json | jq '.items[] | select(.spec.automountServiceAccountToken!=false) | .metadata.name'
# Check exposed services
kubectl get svc --all-namespaces | grep LoadBalancer
# RBAC who can exec
kubectl get clusterrolebindings -o json | jq '.items[] | select(.roleRef.name=="cluster-admin") | .subjects'Interview Questions: K8s Security
Both define a set of permissions — verbs on resources — but differ in scope. A Role is namespaced: it only grants access to resources within its namespace. A ClusterRole is cluster-wide and can grant access to cluster-scoped resources like nodes, or to namespaced resources across all namespaces. The subtlety is the binding: a ClusterRole referenced by a RoleBinding applies only within that binding's namespace, which is how you define one reusable permission set — say "secret-reader" — and grant it per-namespace. You use a Role for app-team permissions scoped to their namespace, and a ClusterRole for platform-level access or for reusable permission templates bound per namespace.
By default every pod gets its service account's token mounted into the filesystem, and that token can call the Kubernetes API with whatever RBAC the service account has. If the pod is compromised — say via an app RCE — the attacker immediately has that token and can use it against the API server to escalate: list secrets, create pods, or exec into others, depending on the SA's permissions. So an over-permissioned service account plus automounting turns a single app compromise into cluster escalation. The mitigations are to set automountServiceAccountToken false on pods that don't need API access, give each workload its own least-privilege service account, and use short-lived projected tokens. It's the Kubernetes equivalent of stealing the pod's identity.
A privileged container runs with all Linux capabilities and access to host devices, effectively dropping the isolation between container and host. With that, an attacker inside the container can see the host's block devices under /dev, mount the host root filesystem, and chroot into it — now operating as root on the node. They could also load kernel modules or write to host paths to establish persistence. The root reason it works is that a container is just an isolated process on the shared host kernel, and privileged removes the restrictions that kept it contained. The defense is never running privileged in production, enforced by Pod Security Standards restricted and an admission policy that rejects privileged pods.
Defense in depth. At the workload level, set the securityContext with runAsNonRoot true and a non-zero runAsUser, drop all capabilities, readOnlyRootFilesystem, and allowPrivilegeEscalation false. But individual specs get forgotten, so I enforce it at the cluster level: apply the Pod Security Standards "restricted" profile to production namespaces via the built-in Pod Security admission, which rejects pods that run as root or request privileges. For richer rules I'd add a policy engine — Kyverno or OPA/Gatekeeper — to validate and even mutate specs at admission. And I'd bake non-root USER into the images themselves. The principle is to make root the rejected exception, enforced by admission control, not left to each developer.
Both are policy engines that run as admission controllers to validate or mutate Kubernetes resources, enforcing policy-as-code. The main difference is the language and ergonomics. OPA/Gatekeeper uses Rego, a powerful general-purpose policy language that's more expressive but has a learning curve, and OPA can be used beyond Kubernetes. Kyverno is Kubernetes-native — policies are written as YAML CRDs that look like Kubernetes resources, so there's no new language to learn, and it does validation, mutation, and generation. In practice teams pick Kyverno for approachability and K8s-only use, and Gatekeeper/Rego when they want maximum expressiveness or a policy language shared across systems.
I treat it as a potential active compromise, because production containers shouldn't have interactive shells — they run one defined process. First I scope without tipping off: identify the pod, image, and node, and pull what the shell did from Falco and the audit logs. Then I think in the three escalation ladders — did it try to reach the pod's cloud credentials or the node metadata, did it use the service account token against the API server, and is the pod privileged enough to escape to the node. For containment I isolate the pod with a deny-all NetworkPolicy and cordon the node rather than immediately deleting the pod, so I preserve evidence; capture what I can; then revoke the pod's identity and, if it may have reached the node or API, treat those as compromised too. Afterward, eradicate by redeploying from a clean image and fixing how the shell got there. It's the EKS pod IR playbook.
Binary Authorization is an admission-time control that only allows container images meeting a policy to run — typically images that are signed and attested by trusted parties, from approved registries. At deploy time the admission controller verifies cryptographic signatures and attestations (via Sigstore/cosign) before the pod is scheduled, so an unsigned image, one from an untrusted source, or one that didn't pass required checks like a vulnerability scan is rejected. This closes a major supply-chain gap: it stops a tampered or rogue image — say one an attacker pushed to the registry — from silently running in the cluster, and it enforces that everything in production came through your trusted build and signing pipeline. It's the runtime enforcement of the "verify provenance" half of supply-chain security.
A Kubernetes Secret is only base64-encoded, not encrypted, and by default it's stored in etcd in plaintext — so anyone who can read the Secret object via the API, read etcd directly, or get an etcd backup has the cleartext. Three fixes, increasing in strength: first, enable etcd encryption-at-rest so the data is encrypted in the datastore and in backups; second, lock down RBAC so very few principals can get secrets, since broad "get secrets" is itself a major risk; and third, the strongest, keep secrets out of etcd entirely by sourcing them from an external manager like Vault or AWS Secrets Manager through the External Secrets Operator or the CSI Secrets Store driver, so the real secret is injected at runtime and never persisted by Kubernetes. The headline is that K8s Secrets are obfuscated, not encrypted.