Kubernetes Fundamentals
Architecture Overview
┌──────────────── Control Plane ─────────────────┐
│ kube-apiserver ←─── only entry point to state │
│ etcd ←─── all cluster state stored │
│ kube-scheduler ←─── assigns pods to nodes │
│ controller-mgr ←─── reconciliation loops │
│ cloud-controller←─── cloud-provider integration│
└─────────────────────────────────────────────────┘
↕ kubelet + kube-proxy
┌──────── Worker Node ────────────────────────────┐
│ kubelet ←─── node agent, talks to API │
│ kube-proxy ←─── iptables/ipvs rules │
│ container runtime (containerd / CRI-O) │
│ Pods │
└─────────────────────────────────────────────────┘etcd — distributed key-value store; contains all secrets, configs, and resource state. If compromised = full cluster compromise.
kube-apiserver — all kubectl commands hit the API server. Authentication → Authorization (RBAC) → Admission Control → etcd.
Memory hookeverything is "desired state in etcd, reconciled by controllers." The one idea that makes Kubernetes click: you don't run things, you declare what you want ("3 replicas of this image") and store that in etcd via the API server. Controllers are loops that constantly compare desired state (etcd) to actual state (the nodes) and nudge reality toward the declaration. Kill a pod and the Deployment controller notices the gap and makes a new one. So K8s is a declarative reconciliation engine, not an imperative runner — which is why GitOps works (git = desired state), why
kubectl delete poddoesn't really "fix" anything (it gets recreated), and why the API server + etcd are the crown jewels.
The Big Picture: How Kubernetes Is Actually Used End-to-End
Architecture diagrams hide the part people actually live in: the workflow from a developer's code to a running, monitored, secured workload. Here's the whole pipeline.
┌─ DEVELOPER ─┐ ┌──── CI (build & check) ────┐ ┌─ REGISTRY ─┐ ┌── GitOps ──┐ ┌──── CLUSTER ────┐
│ writes code │ │ build image (Dockerfile) │ │ store the │ │ ArgoCD/Flux│ │ kube-apiserver │
│ + a Dockerfile│ │ scan image (Trivy/Grype) │──▶│ signed │ │ watches the│──▶│ admission checks│
│ + k8s manifests│ │ run tests, SAST, IaC scan │ │ image │ │ git repo & │ │ schedules pods │
│ → git push │──▶│ sign image (cosign) │ │ (ECR/GCR/ │ │ syncs to │ │ runtime (Falco) │
└──────────────┘ │ push manifest change to git│ │ GHCR) │ │ the cluster│ │ + service mesh │
│ └────────────────────────────┘ └────────────┘ └────────────┘ └─────────────────┘
│ │
└──────────────── observability: metrics (Prometheus), logs, traces, alerts ◀──────────┘The two deployment models: "push" vs "pull" (GitOps)
| Model | How | Tools | Why it matters |
|---|---|---|---|
| Push (imperative/CI-push) | CI runs kubectl apply / helm upgrade against the cluster | GitHub Actions + kubectl/Helm | Simpler, but CI needs prod cluster credentials (a juicy secret), and the cluster can drift from git |
| Pull (GitOps) | An in-cluster agent watches a git repo and syncs the cluster to match it | ArgoCD, Flux | Git is the single source of truth; cluster auto-corrects drift; CI never holds cluster creds; every change is a reviewed, audited git commit |
Memory hookGitOps = "git is the desired state; an in-cluster robot makes reality match." Because K8s is already a reconciliation engine, GitOps just extends the loop out to git: you
git pusha manifest change, and ArgoCD/Flux (running inside the cluster) pulls it and applies it. Huge security wins: the cluster credentials never leave the cluster (CI can't be the breach path to prod), every deploy is a signed/reviewed git commit (audit + rollback =git revert), and config drift becomes detectable — if the live cluster differs from git, that's either an unauthorized change or an incident. This is how most modern shops deploy, and "do you use GitOps?" is a common interview question.
What lives in the GitHub repo
A typical app repo (or a separate "infra"/"deploy" repo) contains:
my-service/
├── src/ # application code
├── Dockerfile # how to build the container image
├── .github/workflows/ci.yml # build, scan, sign, push, bump manifest
└── deploy/
├── base/ # Kubernetes manifests (Deployment, Service, etc.)
│ ├── deployment.yaml
│ └── service.yaml
└── overlays/ # per-environment patches (dev/staging/prod)
├── staging/
└── prod/ # ArgoCD watches this path- Manifests are YAML describing the desired objects. Raw YAML doesn't scale across environments, so people template it:Helm
package manager; templated charts with values per environment (
values-prod.yaml). Think "npm for K8s."Kustomizepatch/overlay model (no templating language); built into
kubectl. Abase/plus environmentoverlays/.
Core Workload Objects
| Object | Description |
|---|---|
| Pod | Smallest deployable unit; one or more containers sharing network + storage |
| Deployment | Manages ReplicaSets; declarative rollouts and rollbacks |
| StatefulSet | Pods with stable identity and persistent storage (databases) |
| DaemonSet | One pod per node (logging agents, security tools) |
| Job / CronJob | Run-to-completion workloads |
| Namespace | Logical isolation; scopes RBAC, network policies, quotas |
Networking
Pod-to-Pod
- Every Pod gets a unique cluster IP
- All pods can reach all other pods by default (flat network)
- Network Policies restrict this
Service Types
| Type | Scope |
|---|---|
| ClusterIP | Internal only (default) |
| NodePort | Exposed on each node's IP:port |
| LoadBalancer | External; provisions cloud LB |
| ExternalName | DNS alias to external service |
DNS
<service>.<namespace>.svc.cluster.local<pod-ip>.<namespace>.pod.cluster.local
CNI Plugins
| Plugin | Notes |
|---|---|
| Calico | Network policy enforcement, BGP routing |
| Cilium | eBPF-based; L7 policy, observability |
| Flannel | Simple overlay; no network policy |
| Weave | Simple; supports encryption |
Container Runtimes & Sandboxing (gVisor, Kata, Firecracker)
When a pod runs, something actually starts the container. That stack has layers, and the security-critical question is how strong the isolation is between the container and the host kernel.
kubelet
│ speaks CRI (Container Runtime Interface)
▼
containerd or CRI-O ← the "high-level" runtime (manages images, lifecycle)
│ calls a "low-level" runtime
▼
runc (default) │ gVisor (runsc) │ Kata Containers │ Firecracker (via Kata)
───────────────────────────────────────────────────────────────────────────────
shares host │ user-space │ real lightweight │ microVM (~125ms boot,
kernel directly │ kernel intercepts│ VM per pod │ minimal device model)
(fast, weakest │ syscalls (strong │ (own kernel, │ (AWS Lambda/Fargate
isolation) │ isolation, some │ strong isolation,│ use this under the hood)
│ perf cost) │ more overhead) │| Runtime | Isolation model | Trade-off | When to use |
|---|---|---|---|
| runc | Namespaces + cgroups, shared host kernel | Fastest; a kernel exploit = host escape | Default; trusted workloads |
gVisor (runsc, Google) | A user-space kernel intercepts syscalls so the container never talks directly to the host kernel | ~Reduced attack surface; some syscall-heavy perf cost & compat gaps | Multi-tenant / running untrusted code |
| Kata Containers | Each pod runs in a real lightweight VM with its own kernel | Hardware-enforced isolation; higher memory/boot overhead | Strong tenancy boundaries needed |
| Firecracker (AWS) | Minimal microVM (tiny device model, fast boot) | VM isolation at near-container speed; used via Kata | Serverless / Fargate / function isolation |
Memory hookthese exist because "a container is a process, not a VM." Default
runccontainers share the host kernel, so one kernel-level container-escape bug compromises every workload on the node. gVisor, Kata, and Firecracker each put a barrier back between the container and the host kernel — gVisor by emulating the kernel in user space (intercepting syscalls), Kata/Firecracker by wrapping the pod in a genuine lightweight VM. The spectrum is isolation vs speed: runc is fastest/weakest, microVMs are strongest/heaviest. You reach for them when running untrusted or multi-tenant code — e.g., a CI system or a platform that runs customers' arbitrary containers. Interviewers love "how would you safely run untrusted containers?" → answer: gVisor or Kata/Firecracker, not plain runc.
Service Mesh & Envoy
As soon as you have many services talking to each other, you need consistent encryption, retries, traffic routing, and observability between them — without baking it into every app. That's a service mesh, and its workhorse is Envoy.
Envoy is a high-performance L7 proxy. In a mesh it runs as a sidecar — a second container injected into every pod — so all traffic in and out of the app goes through its local Envoy:
┌──────── Pod A ────────┐ ┌──────── Pod B ────────┐
│ app ⇄ Envoy sidecar│ ──mTLS──│ Envoy sidecar ⇄ app │
└───────────────────────┘ └───────────────────────┘
(app talks plaintext to its local Envoy; the Envoys
encrypt + authenticate the wire between pods)
Control plane (Istio's istiod / Linkerd) configures all the Envoys:
who can talk to whom, mTLS certs, routing rules, retries, telemetry| You get | Without changing app code |
|---|---|
| Automatic mTLS | Every service-to-service call is encrypted + mutually authenticated — "zero trust" inside the cluster |
| Identity-based authz | "service A may call service B" enforced at the proxy (workload identity, not IP) |
| Traffic management | Canary/blue-green, retries, timeouts, circuit breaking |
| Observability | Golden metrics, traces, and access logs for every call, for free |
the heavyweight, feature-rich mesh (Envoy sidecars + istiod control plane).
lighter, simpler, security-focused (its own micro-proxy).
meshes are emerging to cut the per-pod sidecar overhead.
Memory hooka service mesh is "TLS + authz + telemetry for east-west traffic, moved out of the app into a sidecar." The app keeps speaking plain HTTP to
localhost; the injected Envoy proxy transparently handles mTLS encryption, who-can-call-whom authorization, retries, and metrics for it. Security value: you get encryption-in-transit and identity-based access control between every microservice without trusting developers to implement it, plus per-call visibility that's gold for detection. The cost is operational complexity and a sidecar in every pod. Mnemonic: mesh = mutual-TLS + metrics + routing, as infrastructure.
Storage & Persisting Data
Containers are ephemeral — when a pod dies, its writable layer is gone. So "where does the data live?" is a real design question.
Ephemeral (dies with the pod) Persistent (survives pod restart/reschedule)
────────────────────────────── ─────────────────────────────────────────────
• container writable layer • PersistentVolume (PV) ← real storage (EBS, GCE PD, NFS)
• emptyDir volume • PersistentVolumeClaim (PVC) ← a pod's request for a PV
• ConfigMap / Secret (read-only) • StorageClass ← dynamically provisions PVs on demand
• bound to a StatefulSet for stable identity (databases)| Object | Purpose |
|---|---|
| PersistentVolume (PV) | Cluster-level storage resource (backed by EBS, GCE PD, NFS, etc.) |
| PersistentVolumeClaim (PVC) | Pod's request for storage ("I need 20Gi") — binds to a PV |
| StorageClass | Dynamic provisioner — auto-creates a PV from the cloud when a PVC asks |
| CSI driver | Container Storage Interface — the plugin that lets K8s talk to any storage backend |
| StatefulSet | Gives pods stable identity + their own PVC — how you run databases in K8s |
| ConfigMap | Non-sensitive config data |
| Secret | Base64-encoded sensitive data (not encrypted by default in etcd) |
Memory hook"PVC is the request, PV is the storage, StorageClass is the vending machine." A pod doesn't grab disk directly; it files a claim (PVC: "give me 20Gi, fast SSD"), and either an admin pre-made a matching PV or a StorageClass dynamically provisions one from the cloud (an EBS volume on EKS, a PD on GKE) via a CSI driver. Stateful things (databases) use a StatefulSet so each replica keeps a stable name and its own persistent volume across restarts. The reality-check most people learn: for serious databases, many teams use the cloud's managed DB (RDS/Cloud SQL) instead of running stateful DBs in K8s at all — K8s shines for stateless services. Security angle: a PV often outlives pods, so it can hold sensitive data and snapshots that need encryption and access control of their own.
Scheduling
# Node affinity
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: topology.kubernetes.io/zone
operator: In
values: [us-east-1a]
# Resource requests and limits
resources:
requests:
cpu: "250m"
memory: "64Mi"
limits:
cpu: "500m"
memory: "128Mi"prevent pods from scheduling on nodes unless they tolerate the taint
minimum available pods during voluntary disruption
ConfigMaps and Secrets
# Secret (base64, not encrypted at rest unless KMS envelope encryption enabled)
apiVersion: v1
kind: Secret
metadata:
name: db-creds
type: Opaque
data:
password: cGFzc3dvcmQ= # base64("password") — NOT secure without etcd encryption
# Consume as env var
env:
- name: DB_PASSWORD
valueFrom:
secretKeyRef:
name: db-creds
key: passwordSecure secret storageenable etcd encryption at rest (AES-GCM KMS provider) or use External Secrets Operator + cloud secret manager.
Probes
livenessProbe: # restart container if fails
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
readinessProbe: # remove from Service endpoints if fails
httpGet:
path: /ready
port: 8080
startupProbe: # delay other probes until app is started
httpGet:
path: /started
port: 8080
failureThreshold: 30
periodSeconds: 10How Developers Authenticate to the Cluster
A crucial real-world point: Kubernetes has no built-in user database. There are no "user accounts" in etcd for humans. The API server authenticates requests by trusting an external mechanism, then authorizes with RBAC. So "how do people log in" = "what external identity does the API server trust?"
kubectl ──(credential in kubeconfig)──▶ kube-apiserver
│ 1. AUTHENTICATE (who are you?)
│ • client cert (rare for humans)
│ • OIDC token (Okta/Entra/Google) ← modern
│ • cloud IAM (EKS/GKE) ← managed clusters
│ 2. AUTHORIZE (RBAC: what can you do?)
│ 3. ADMISSION (policy checks)
▼
etcdThe common human-auth options
| Method | How it works | Reality |
|---|---|---|
| Client certificates | kubeconfig holds a cert the API server trusts | Hard to revoke (no CRL), no groups — fine for break-glass, bad for teams |
| OIDC (Okta, Entra ID, Google, Dex) | API server is configured to trust an OIDC provider; users log in via SSO and get a short-lived ID token; the token's groups claim maps to RBAC | The standard way to give humans cluster access — SSO + MFA + central deprovisioning |
| Cloud IAM (EKS / GKE) | The cloud's IAM is the identity layer (see below) | Default on managed clusters; ties cluster access to AWS/GCP IAM you already run |
Wiring Okta (or any IdP) to a cluster — the shape of it
1. Register the cluster as an OIDC app in Okta (issuer URL, client ID)
2. Point the kube-apiserver at Okta:
--oidc-issuer-url=https://company.okta.com
--oidc-client-id=kubernetes
--oidc-username-claim=email
--oidc-groups-claim=groups ← Okta groups → K8s RBAC groups
3. Bind Okta groups to RBAC:
a RoleBinding gives group "okta:platform-admins" the cluster-admin ClusterRole
4. Developers run `kubectl` via a helper (kubelogin/Dex) that does the Okta SSO
flow, gets a short-lived OIDC token, and kubectl presents it.Memory hookK8s authenticates by trust delegation, not its own user list. There's no
kubectl create user. The API server says "I trust Okta (or AWS IAM, or this CA) to tell me who you are," then RBAC decides what that identity can do. So adding a team is really "map your IdP groups to RBAC RoleBindings." Why it matters for security: you get SSO, MFA, and instant offboarding for free (disable the person in Okta → cluster access gone), and the tokens are short-lived so there's no standing kubeconfig secret to steal. The classic anti-pattern is sharing one long-lived admin kubeconfig — no attribution, no revocation.
Service accounts vs human users
OIDC / cloud IAM (above).
(in-cluster): ServiceAccount tokens (projected, short-lived).
(out to AWS/GCP): IRSA / Pod Identity (EKS) or Workload Identity (GKE) — see security.md.
Local vs EKS vs GKE — What's Actually Different
The big distinction is who runs the control plane (API server, etcd, scheduler). That determines what you secure and what's the cloud's problem.
| Local (kind/minikube/k3s) | EKS (AWS) | GKE (Google) | |
|---|---|---|---|
| Control plane | You run it (or it's a single binary) | AWS manages it (you never touch etcd/apiserver hosts) | Google manages it (most "Kubernetes-native" experience) |
| You secure | Everything | Nodes, RBAC, network, workloads | Nodes, RBAC, workloads (more is automated) |
| Cluster auth | Certs / local | AWS IAM → maps to RBAC (EKS access entries / aws-auth) | GCP IAM → maps to RBAC natively |
| Pod→cloud identity | n/a | IRSA / EKS Pod Identity | Workload Identity |
| Node upgrades/patching | You | You (or managed node groups / Fargate) | Largely automated (auto-upgrade, Autopilot = nodeless) |
| Image scanning | DIY (Trivy) | ECR scanning (Inspector) | Artifact Registry scanning + Binary Authorization |
| Built-in threat detection | None | GuardDuty EKS Protection | GKE Security Posture + SCC |
| Use for | Dev, CI, learning, edge (k3s) | AWS-centric orgs | Teams wanting the most managed K8s (Autopilot) |
Memory hookmanaged clusters move the control plane (and its risk) to the cloud; you still own the workloads. On EKS/GKE you never see etcd or the API server hosts — the cloud patches and secures them, which removes a huge class of "is my control plane hardened?" worries (and the etcd-exposure incidents that plagued self-managed clusters). What stays yours on every platform: RBAC, network policy, pod security, secrets, and the images you run. The other big difference is identity glue — EKS and GKE bolt the cloud's IAM onto cluster auth (IAM → RBAC) and onto pod-to-cloud access (IRSA / Workload Identity), so you manage one identity system instead of two. GKE Autopilot goes furthest — you don't manage nodes at all.
Monitoring, Image Scanning & "Is Something Wrong?"
Three different questions, three different toolsets.
1. Is the cluster healthy / behaving normally? (observability)
Metrics → Prometheus (scrapes pods/nodes) + Grafana (dashboards) + Alertmanager
Logs → Fluent Bit/Fluentd → Loki / Elasticsearch / cloud logging
Traces → OpenTelemetry → Jaeger/Tempo (request flow across services)
"Golden signals": latency, traffic, errors, saturation2. Is the image safe? (image scanning — shift left + at runtime)
Trivy, Grype, Docker Scout, ECR/Artifact Registry scanning — scan for known CVEs in OS packages and app deps before deploy, fail the build on criticals.
sign images (cosign) and enforce signatures at admission (Kyverno / Binary Authorization) so only trusted images run.
(syft) so you can answer "are we affected by CVE-X?" in minutes, not days.
3. Is the configuration wrong / drifting? (posture + admission + runtime)
(prevent): Pod Security Standards + Kyverno/OPA reject privileged pods, root containers, :latest tags, missing limits.
(detect): kube-bench (CIS benchmark), kubescape, Trivy's misconfig scan, cloud posture (GKE Security Posture, GuardDuty).
(GitOps): ArgoCD flags any live object that differs from git — unauthorized change or incident.
(detect active threats): Falco / eBPF tools watch syscalls for shell-in-container, sensitive mounts, unexpected egress.
Memory hookscan the image before it runs, enforce policy as it's admitted, watch behavior while it runs. Three time-points: build/registry (image CVE scanning + signing — "shift left"), admission (policy gate — reject bad configs and unsigned images before they're scheduled), and runtime (Falco/eBPF — catch what got through). A mature setup has all three, because each catches what the others miss: scanning misses logic/config flaws, admission misses zero-days, runtime is the last line. "How do you know if something's wrong with an image or config?" → name those three layers.
Key kubectl Commands
# Context / cluster
kubectl config get-contexts
kubectl config use-context <ctx>
# Inspect
kubectl get pods -n <ns> -o wide
kubectl describe pod <pod> -n <ns>
kubectl logs <pod> -n <ns> --previous
kubectl exec -it <pod> -n <ns> -- /bin/sh
# RBAC
kubectl auth can-i create pods --as=system:serviceaccount:default:my-sa
kubectl get rolebindings,clusterrolebindings --all-namespaces | grep <subject>
# Audit
kubectl get events --sort-by=.lastTimestamp -n <ns>Reality Check: Do You Actually Dump Memory in Container IR?
You asked the right question — and the honest answer is rarely, in most teams. Classic host DFIR puts memory acquisition near the top (order of volatility). Container/cloud-native IR usually works differently, for practical reasons:
Why memory dumping is uncommon in container IR:
A pod is one process group on a shared kernel; capturing its RAM cleanly is fiddly (you'd typically dump the node's memory, not the pod's), and the pod may be gone before you act.
the API audit log (who did what), runtime telemetry (Falco/eBPF — the syscalls, processes, connections), container/node logs, the image (you can pull and analyze the exact bytes), and cloud control-plane logs. These answer "what happened" without RAM.
The container filesystem layer, the image digest, and the manifest tell you most of what ran.
Since you rebuild from a known-good image anyway, deep live-forensics on a doomed pod has lower ROI.
When teams do reach for memory:
- Fileless / in-memory malware that leaves little on disk (a crypto-miner injected into a process, an in-memory implant) — here RAM is where the evidence lives.
- Advanced/targeted intrusions where you need to recover injected code, keys, or decrypted configs.
- Malware reverse-engineering, not routine SOC triage.
- In practice it's usually a node-level capture (snapshot the node's disk and/or memory) or grabbing
/proc/<pid>/of the suspect process, done by a specialized DFIR function — not the everyday IR responder.
Memory hookThe honest interview answer"In container IR I prioritize the API audit log, runtime/eBPF telemetry, logs, and the image — they answer 'what happened' faster and more reliably than pod memory, and the workload is ephemeral and gets rebuilt anyway. I'd only do memory forensics for fileless/in-memory malware or a targeted intrusion, and even then it's typically a node-level capture by a DFIR specialist. So no — routine pod memory dumping is not standard practice, and saying otherwise would be cargo-culting host forensics into a cloud-native world." Saying this — that you match the technique to the environment — signals real experience.
Interview Questions: K8s Fundamentals
kubectl apply -f deployment.yaml.kubectl sends the manifest to the kube-apiserver, which runs it through three gates: authentication (who are you — via your kubeconfig cert, OIDC token, or cloud IAM), authorization (does RBAC permit this verb on this resource), and admission control (mutating then validating webhooks and Pod Security — inject defaults, reject policy violations like privileged pods). If it passes, the desired state is written to etcd. From there controllers take over: the Deployment controller sees a new desired state, creates a ReplicaSet, which creates Pod objects; the scheduler assigns each pod to a node based on resources, affinity, and taints; and the kubelet on that node tells the container runtime to pull the image and start the containers, reporting status back. The key insight is that apply just records desired state — controllers reconcile reality to match it asynchronously.
A Deployment manages stateless, interchangeable pods — they get random names, any replica is as good as any other, and they scale and roll out freely. A StatefulSet is for workloads that need stable identity and persistent state, like databases: each pod gets a stable ordinal name, its own persistent volume that follows it across restarts and reschedules, and ordered, controlled rollout and scaling. So you use a Deployment for web services and APIs, and a StatefulSet when each replica is distinct and must keep its data and identity. The practical caveat I'd add is that many teams avoid running serious databases in Kubernetes at all and use a managed cloud database, reserving K8s for the stateless tier.
Kubernetes has no built-in user database — the API server delegates authentication to something it trusts, then authorizes with RBAC. For humans the modern approach is OIDC: you configure the API server to trust an identity provider like Okta with its issuer URL and client ID, and map a claim like email to the username and the groups claim to RBAC groups. Then you bind those IdP groups to roles — for example, the Okta group platform-admins gets the cluster-admin ClusterRole via a RoleBinding. Developers run kubectl through a helper that performs the Okta SSO login, receives a short-lived OIDC token, and presents it. The wins are SSO with MFA, central deprovisioning — disable the user in Okta and cluster access is gone — and short-lived tokens with no standing kubeconfig secret to steal. On EKS and GKE the cloud's IAM plays this role instead.
The biggest difference is who runs the control plane. Locally — kind, minikube, k3s — you run everything yourself and secure all of it, which is great for dev and CI but you own etcd and the API server. On EKS and GKE the cloud manages the control plane, so you never touch etcd or the apiserver hosts and the provider patches and secures them, removing a whole class of control-plane hardening concerns. What stays yours everywhere is workloads, RBAC, network policy, pod security, and images. The other big difference is identity glue: EKS and GKE wire the cloud's IAM into cluster auth — IAM identities mapped to RBAC — and into pod-to-cloud access via IRSA or Pod Identity on EKS and Workload Identity on GKE. GKE leans most managed, with Autopilot removing node management entirely. So managed clusters shift the control plane and its risk to the cloud while you keep owning the workloads.
I think in three time-points. Before it runs — in CI and the registry — I scan images for known CVEs with something like Trivy or Grype, fail builds on criticals, generate an SBOM so I can answer impact questions fast, and sign images with cosign. As it's admitted, I enforce policy at admission with Pod Security Standards plus Kyverno or OPA, rejecting privileged or root containers, latest tags, and missing limits, and requiring signed images via Binary Authorization. While it runs, I detect at runtime with Falco or eBPF tooling watching for shells in containers, sensitive mounts, and unexpected egress, and I use posture scanners like kube-bench for CIS compliance. And with GitOps, ArgoCD flags configuration drift — anything live that differs from git is an unauthorized change or an incident. Each layer catches what the others miss: scanning misses config and logic flaws, admission misses zero-days, runtime is the last line.
A service mesh handles the cross-cutting concerns of service-to-service traffic — encryption, identity, routing, observability — without putting that logic in each app. It works by injecting an Envoy proxy as a sidecar into every pod, so all of a pod's traffic flows through its local Envoy, and a control plane like Istio's istiod configures all the Envoys. The security value is significant: automatic mutual TLS encrypts and authenticates every service-to-service call, giving you zero trust inside the cluster without developers implementing it; authorization policies enforce which workload identities may call which services at the proxy rather than relying on network IPs; and you get per-call telemetry that's excellent for detection. The cost is operational complexity and a sidecar in every pod, which is why lighter meshes like Linkerd and sidecar-less approaches exist.
The problem is that default runc containers share the host kernel, so a single kernel-level escape compromises every workload on the node — unacceptable when you're running untrusted code, like a CI platform or something executing customers' arbitrary containers. So I'd add a stronger isolation boundary using a sandboxed runtime: gVisor, which runs a user-space kernel that intercepts the container's syscalls so it never talks directly to the host kernel, or Kata Containers and Firecracker, which wrap each pod in a real lightweight VM with its own kernel — that's what AWS uses under Fargate and Lambda. The trade-off is performance and some compatibility cost, so I'd reserve it for the untrusted tier. Alongside that I'd enforce non-root, drop capabilities, seccomp, no privileged pods, strict network policy, and per-tenant namespaces with resource quotas — defense in depth, with the sandboxed runtime as the kernel-isolation backstop.
A Secret is only base64-encoded, not encrypted, and it's stored in etcd in plaintext unless you turn on encryption-at-rest — so anyone who can read the Secret via the API, read etcd, or get an etcd backup has the cleartext. Three fixes in increasing strength: enable etcd encryption-at-rest, ideally with a KMS provider so the encryption key isn't on the host; restrict RBAC so very few principals can get secrets, since broad get-secrets is itself a major risk; and strongest, keep secrets out of etcd entirely by sourcing them from an external manager like Vault or a cloud secret manager via the External Secrets Operator or the CSI Secrets Store driver, so the real secret is injected at runtime and never persisted by Kubernetes. The headline is that Secrets are obfuscated, not encrypted.
etcd is the cluster's single source of truth — it stores all object state, configuration, and Secrets. That makes it the crown jewel: anyone who can read etcd can read every Secret in the cluster, and anyone who can write to it can manipulate any resource, effectively owning the cluster while bypassing the API server's authentication, authorization, and admission controls entirely. So protecting it is foundational: encrypt it at rest ideally with KMS, require mutual TLS for all etcd access, tightly restrict network access so only the API server talks to it, and secure its backups, because an unencrypted etcd snapshot in a bucket is a full credential dump. On managed clusters like EKS and GKE the cloud runs and secures etcd for you, which is one of the main security benefits of going managed.