Security Notes
Start here

Concepts Over Tools — A Security Engineer's Foundation

19 min read 14 sections

Why this file existsChasing new tools is a trap. Terraform vs Pulumi, Falco vs Wazuh, k8s Audit vs eBPF — the specific tool matters less than understanding the concept the tool implements. Tools come and go; the underlying principles remain constant. This file maps every major security domain to its foundational concept, so when you encounter a new tool, you can immediately ask: "What concept is this implementing?" and reason about its trade-offs from first principles.


The Core Problem With Tool-First Thinking

When you learn a tool without understanding the concept it embodies, you can only use the tool — you cannot reason about it, adapt when it fails, or evaluate its replacements. When you understand the concept, you can:

  • Use any tool in that category after a 20-minute read of the docs
  • Spot what the tool can't do (its gaps)
  • Explain to a non-technical audience what problem is being solved
  • Debug it when it behaves unexpectedly
  • Interview confidently even if you've never used the specific tool the interviewer asks about

"I've used Falco in production" is one data point. "I understand eBPF-based syscall tracing and why it beats auditd for real-time detection" tells you whether I can use the next five tools like it."


Identity and Access Management

The Concept: Least Privilege and the Principle of Need-to-Know

Every security model for access control derives from two questions:

Identity

Who (or what) is making this request?

Authorization

Of the things this identity is allowed to do, is this specific action one of them?

The concept that unifies all IAM is least privilege — grant only the minimum permissions required to accomplish the specific task, for the shortest time necessary, to the most specific resource possible. This is not a configuration setting; it's a design philosophy.

Tool-first thinking:    "I need to learn Okta, then AWS IAM, then GCP IAM, then K8s RBAC"
Concept-first thinking: "All of these are implementations of the same model:
                         identity (principal) → role (collection of permissions)
                         → resource (what the permissions apply to).
                         Learn the model once, map new tools onto it."

The model that underlies every IAM tool:

Principal → (authenticated as) → Identity
Identity  → (assigned to)      → Role / Policy
Role      → (grants)           → Permission (action + resource + condition)
Permission → (is evaluated against) → Request (who, what action, which resource, when, from where)

Every tool — AWS IAM policies, GCP IAM bindings, Kubernetes RBAC, Okta, Vault ACL policies, GitHub Actions permissions — is a different syntax for expressing this same model.

What to ask about any IAM tool:

  • How does it establish identity? (passwords, certificates, OIDC tokens, service account keys)
  • How granular are its permissions? (per-resource vs per-service vs per-account)
  • Does it support conditions? (time-based, network-based, MFA-required)
  • How does it handle service-to-service authentication? (is there a machine identity mechanism?)
  • How do you audit who had what access, when?

Concept: Just-In-Time (JIT) access

Permanent high-privilege roles are a persistent risk surface — compromised credentials give an attacker permanent access. JIT access means permissions are granted only when needed, for a fixed window, and expire automatically. The concept is the same whether the tool is CyberArk, HashiCorp Boundary, AWS IAM Identity Center's time-limited sessions, or a custom Vault lease.


Network Security

The Concept: Defence in Depth Through Layered Filtering

Network security is not about any single firewall or WAF rule. It is about ensuring that traffic must pass through multiple independent filtering points, each operating on different data (IP vs hostname vs HTTP content vs TLS certificate), so that bypass of one layer does not mean bypass of all layers.

Concept: Every filtering decision is a question with two parts:
1. What information is available at this point in the stack to make the decision?
2. What is the default when the decision is ambiguous? (allow or deny)

IP firewall: sees source/dest IP and port. Default deny = zero-trust network.
DNS firewall: sees hostname. Catches C2 domains that change IPs constantly.
TLS inspection: sees certificate and SNI. Catches encrypted C2 over HTTPS.
WAF: sees HTTP headers and body. Catches SQLi/XSS/SSRF that IP rules miss.
Service mesh: sees service identity (mTLS cert). Enforces which services may talk.

The concept of stateful vs stateless filtering:

A stateless firewall (old iptables -j ACCEPT/DROP rules) evaluates each packet independently. It cannot tell whether a packet is part of an established connection or a new connection attempt. A stateful firewall (nftables ct state established,related accept) tracks connection state — it allows response packets automatically without explicit rules.

Why it matters: a stateless firewall that allows inbound port 80 also allows an attacker to craft response-looking packets from arbitrary sources. A stateful firewall only allows packets that are responses to connections the protected host initiated.

What to ask about any network security tool:

  • At which layer of the network stack does this operate? (L3/L4 vs L7)
  • What information does it have to make decisions?
  • Is it stateful or stateless?
  • Is it inline (blocks traffic) or passive (observes and alerts)?
  • What is the default: allow-everything-not-denied, or deny-everything-not-allowed?

Concept: Zero Trust Networking

Traditional perimeter security assumes everything inside the network boundary is trusted. Zero Trust treats every request as potentially hostile regardless of network location — every access request is authenticated, authorized, and encrypted, even for internal service-to-service calls. The concept doesn't require a specific tool; it requires:

  1. Identity verification on every request (not just at the perimeter)
  2. Least-privilege access to resources (not subnets)
  3. Assume breach — log everything, detect lateral movement

Tools (BeyondCorp, Zscaler, Cloudflare Access, Pomerium) are implementations. Understand the concept and you can evaluate any of them.


Endpoint Security

The Concept: Behavioural Detection vs Signature Detection

Signature detection (traditional AV): compare file hashes and byte patterns against a database of known malware. Fast, low false-positive rate, but blind to anything not in the database. A new malware variant with one changed byte has a different hash.

Behavioural detection (modern EDR): observe what programs do (which syscalls they make, which files they write, which network connections they open, which processes they spawn) and determine whether the behaviour matches attack patterns — regardless of what the binary looks like. A binary that injects code into another process is suspicious whether it's a known tool or a novel one.

Signature:  "This exact sequence of bytes was seen in Mimikatz"
Behavioural: "A process opened lsass.exe with PROCESS_VM_READ access" → suspicious regardless of what the process is

Signature catches: known malware, known C2 IPs/domains
Behavioural catches: LOLBins (legitimate tools used maliciously), fileless attacks, novel malware
Neither catches perfectly: attackers customise signatures, and legitimate tools do many suspicious things

The concept of an IOC vs an IOA:

IOC (Indicator of Compromise)

evidence that a compromise already happened — a file hash, an IP address, a registry key. Reactive. CrowdStrike's IOC-based detection is like matching signatures.

IOA (Indicator of Attack)

a behavioural pattern indicating an attack is in progress — a sequence of actions (encoded PowerShell → write to startup folder → network connection). Proactive. CrowdStrike's IOA engine fires on TTPs regardless of the specific binary.

What to ask about any EDR/AV tool:

  • Does it operate primarily on signatures or behaviour?
  • At what layer? (userspace hooks, kernel callbacks, eBPF, hardware-level)
  • What is its detection latency? (milliseconds vs minutes)
  • Can it block (inline) or only alert (passive)?
  • How does it handle living-off-the-land attacks (LOLBins)?

Cloud Security

The Concept: The Shared Responsibility Model and the Control Plane

Cloud providers do not give you a server — they give you a control plane API through which you request resources. This creates two distinct security domains:

  1. The control plane (IAM, APIs, configuration): who can call aws ec2 run-instances, who can modify bucket policies, who can create service accounts. This is entirely your responsibility and is the primary attack surface for cloud-native attacks.

  2. The data plane (what runs inside the resources): the OS, application code, network traffic between workloads. Partially shared responsibility.

Most cloud breaches are control plane attacks:
- Stolen IAM keys → call CreateUser → persist
- Misconfigured S3 bucket → public read → data exfiltration
- Overpermissioned Lambda → ssm:GetParameter → secrets exfiltration
- Service account key committed to git → lateral movement across projects

Not OS exploits. Not application vulnerabilities. IAM and configuration.

The concept of the Metadata Service (IMDS) as an attack primitive:

Every major cloud provider has an instance metadata service (AWS: 169.254.169.254, GCP: metadata.google.internal). It gives running instances their identity and temporary credentials. SSRF vulnerabilities in applications become cloud credential theft when attackers can reach the metadata service. IMDSv2 (AWS) requires a PUT request first (mitigates SSRF that can only do GET). Understanding this concept applies to any cloud provider.

What to ask about any cloud security tool (GuardDuty, Security Command Center, Defender for Cloud):

  • Is it analysing control plane events (API calls) or data plane events (network, process)?
  • What is its detection latency and coverage?
  • Can it detect misconfigurations proactively, or only active attacks?
  • How does it handle multi-account / multi-project environments?

Container and Kubernetes Security

The Concept: The Container Is Not a Security Boundary

Containers share the host kernel. A container escape (via a kernel vulnerability or misconfiguration) gives the attacker access to the host and potentially all other containers on it. Containers provide isolation (process, filesystem, network namespaces), not security boundaries.

Security boundaries (do isolate):   VM hypervisor, separate physical hosts, gVisor/Kata
Isolation primitives (don't isolate): namespaces, cgroups, seccomp, AppArmor

A privileged container (--privileged) has almost the same capabilities as the host.
A container with access to /var/run/docker.sock effectively has root on the host.
A container with hostPID:true can see all host processes.

The concept of Pod Security Standards (PSS):

Rather than configuring dozens of security settings per pod, PSS defines three predefined levels:

  • Privileged: no restrictions (for system daemonsets)
  • Baseline: minimum restrictions to prevent known escalations (no hostPID, no hostNet, no privileged containers)
  • Restricted: hardened (no root, seccomp required, capabilities dropped)

This is the concept — any admission controller (Kyverno, OPA/Gatekeeper, native PSA) is just a tool implementing the same enforcement.

The concept of workload identity in Kubernetes:

Applications running in pods need to authenticate to cloud APIs. The wrong way: store cloud credentials in a Secret and mount them. The right way: use workload identity — the pod's Kubernetes Service Account is federated to a cloud IAM identity via OIDC. The pod gets a short-lived OIDC token automatically; the cloud provider validates it and issues temporary credentials. No long-lived secrets anywhere.

Tools: AWS IRSA, GCP Workload Identity, Azure Managed Identity. All implement the same concept.

What to ask about any Kubernetes security tool:

  • Does it operate at admission (before resource creation) or at runtime?
  • Does it inspect configuration (static) or behaviour (dynamic)?
  • What is its false-positive rate in a busy cluster?
  • Does it integrate with the cloud provider's IAM for workload identity?

CI/CD and Supply Chain Security

The Concept: Every Step in the Pipeline Is a Trust Boundary

A CI/CD pipeline is a chain of automated steps, each with access to progressively more powerful credentials and systems. An attacker who compromises any step in the chain inherits all the trust of subsequent steps.

Developer laptop → git push → GitHub Actions runner → AWS credentials → production
                ↑ compromise any link = compromise everything after it

Common attack surfaces:
  - Developer laptop: git credentials, SSH keys
  - Source repository: branch protection, PR review requirements
  - CI runner: secrets injection, environment variables, runner isolation
  - Build artifacts: tampering between build and deploy
  - Package registries: publishing malicious updates
  - Production: deployment credentials, rollback capability

The concept of SLSA (Supply chain Levels for Software Artifacts):

SLSA is a framework for reasoning about the integrity of the software supply chain. It defines four levels of build provenance:

L1

build process is documented

L2

build is hosted on a build service and provenance is generated

L3

build is isolated, provenance is signed and non-falsifiable

L4

two-person review required, hermetic reproducible builds

Tools (Sigstore, cosign, SBOM generators) are mechanisms to achieve specific SLSA levels. The concept is: prove, cryptographically, that artifact X was built from source Y using build system Z, and that nobody tampered with it in transit.

The concept of OIDC-based short-lived credentials in CI:

Instead of storing a static AWS key in a CI secret (which persists, can be leaked, is hard to rotate), CI/CD systems federate their identity with the cloud provider via OIDC. GitHub Actions gets a short-lived OIDC token; AWS validates it and issues a temporary 15-minute IAM credential. If the token leaks, it expires. No long-lived secrets in CI. The concept applies to any CI platform and cloud provider combination.


Cryptography

The Concept: Cryptography Solves Three Problems, Not One

Beginners treat cryptography as monolithic ("encrypt the data"). Professionals separate three distinct problems that require different solutions:

1. Confidentiality:  prevent unauthorized parties from reading the data
   → Symmetric encryption (AES-256-GCM) for bulk data
   → Asymmetric encryption (RSA-OAEP, ECDH) for key exchange

2. Integrity:        detect if data was tampered with in transit or storage
   → HMAC (keyed hash), AEAD ciphers (AES-GCM authenticates AND encrypts)
   → Signatures (RSA-PSS, ECDSA, Ed25519)

3. Authenticity:     verify the identity of the sender or the origin of data
   → Digital signatures (private key signs → public key verifies)
   → Certificates (trusted CA binds identity to public key)

Using AES-CBC without a MAC gives confidentiality but not integrity — an attacker can flip bits in the ciphertext and the decryption produces corrupted plaintext silently. This is the padding oracle vulnerability class. AEAD modes (AES-GCM, ChaCha20-Poly1305) solve all three at once — always prefer AEAD.

The concept of key hierarchy

You never encrypt data directly with a master key. Instead:

Master Key (stored in HSM / never leaves secure hardware)
    └── encrypts Data Encryption Key (DEK)
              └── DEK encrypts the actual data

To decrypt: retrieve encrypted DEK → decrypt with master key → decrypt data with DEK
To rotate: generate new DEK → re-encrypt data → update encrypted DEK → don't need to touch master key

This is envelope encryption — used by AWS KMS, GCP Cloud KMS, HashiCorp Vault. The concept is the same regardless of tool.


Observability and Detection

The Concept: You Can't Detect What You Don't Log, and You Can't Log Everything

Detection engineering is fundamentally a signal-to-noise problem. The goal is not to log everything — it's to collect the minimum set of signals that allows you to detect the most impactful threats with acceptable latency and false-positive rates.

Too little logging:  attacker dwell time measured in months (median: 21 days as of 2023)
Too much logging:    alert fatigue → analysts stop investigating → zero effective detections
Right amount:        high-fidelity signals for high-impact TTPs, correlated across sources

The concept of detection layers (defense in depth for monitoring):

Network layer:    Who is talking to whom? (NetFlow, DNS logs, proxy logs)
→ Catches: beaconing, lateral movement, exfiltration, C2 communication

Host layer:       What did a process do? (auditd, EDR, eBPF, Sysmon)
→ Catches: process injection, credential access, persistence installation, privilege escalation

Application layer: What did the app do? (application logs, WAF, SIEM correlation)
→ Catches: authentication attacks, SQLi, API abuse, insider threats

Cloud/Identity layer: Who called what API? (CloudTrail, GCP Audit Logs, Azure Activity Log)
→ Catches: credential misuse, privilege escalation, data exfiltration via APIs

No single layer catches everything. Attackers who evade network detection (encrypted C2) are often visible at the host layer. Attackers who evade host detection (LOLBins) are visible at the network layer.

The concept of the detection pyramid (MITRE ATT&CK layers):

Most volatile (tools change often)    →  specific tool hashes, IP addresses, domains
     ↓                                   network/host artifacts
     ↓                                   TTPs (Tactics, Techniques, Procedures)
Most stable (hard for attacker to      →  adversary objectives (persistence, exfil, lateral movement)
change without fundamentally
changing their approach)

Write detections at the TTP level, not the IOC level. An IOC (C2 IP) changes with every infrastructure burn. A TTP (process injection via NtCreateThreadEx) requires the attacker to fundamentally retool.

What to ask about any SIEM/detection tool:

  • At which layer does it collect data?
  • What is its detection latency?
  • Does it detect on IOCs (signatures) or TTPs (behavioural)?
  • How does it handle log ingestion at scale without alert fatigue?
  • What does MITRE ATT&CK coverage look like?

Infrastructure as Code

The Concept: Infrastructure State Is Code, and Code Has a Review Process

IaC is not about which tool you use (Terraform, Pulumi, CloudFormation, Ansible). It's about applying software engineering practices to infrastructure:

Without IaC:
  - Configuration drift: prod is different from staging in unknown ways
  - No audit trail: who changed what, when, why?
  - Manual deployment: human error, inconsistency, undocumented snowflake servers
  - Rollback: impossible or manual

With IaC:
  - Infrastructure is version-controlled → git blame, PR review, audit trail
  - Declarative: state file or plan shows exactly what will change before it changes
  - Reproducible: identical infra in any environment from the same code
  - Drift detection: `terraform plan` shows deviation from desired state

The security concept of "plan before apply":

Every IaC tool has a dry-run mode (Terraform plan, Pulumi preview, CloudFormation change sets). The security discipline is: never apply without a reviewed plan. In CI/CD: PR → plan output in PR comment → human review → merge → apply. This is not a Terraform feature — it's a process that every IaC tool supports.

The concept of state drift and why it's a security concern:

Infrastructure state drift (manual changes made outside IaC) is not just an operational problem — it's a security problem. Attackers who achieve cloud console access commonly create resources (EC2 instances, IAM users, Lambda functions) that are not in any IaC state file. Drift detection detects them: run terraform plan and look for unexpected changes the tool wants to destroy.


Secrets Management

The Concept: Secrets Have a Lifecycle, Not Just a Location

The beginner's approach to secrets: "don't commit them to git." The professional approach: manage the entire lifecycle.

Secret lifecycle:
  1. Generation:   how is it created? (random, derived, issued by IdP?)
  2. Storage:      where does it live? (encrypted at rest, access controlled)
  3. Distribution: how does it reach consumers? (env var, file, API, OIDC token)
  4. Rotation:     how often is it changed? (scheduled, on-breach, at-will)
  5. Revocation:   how do you invalidate it immediately?
  6. Auditing:     who accessed it, when?

The concept of secret zero

Every secrets management system has a bootstrapping problem: to authenticate to the secrets manager (HashiCorp Vault, AWS Secrets Manager), you need a credential — but where does that credential come from? This is the "secret zero" problem. Solutions:

  • Cloud workload identity (IRSA, GCP Workload Identity) — the cloud provider's identity service, using OIDC, is the "zero secret"
  • HashiCorp Vault agent with AWS IAM auth — the instance's IAM role is the zero secret
  • Hardware-bound keys (TPM, HSM) — the zero secret is physically bound to hardware

Understanding this concept lets you evaluate any secrets management architecture.

Why environment variables are not a good secrets distribution mechanism:

  • printenv as any user → visible to all processes on the host
  • Appears in /proc/PID/environ → any process with read access sees it
  • Accidentally logged by frameworks, crash reporters, monitoring tools
  • No rotation mechanism — changing the value requires restarting the process
  • No audit trail — you can't know if/when the env var was read

Better: fetch the secret at runtime from a secrets manager API, use it, discard it from memory immediately.


Incident Response

The Concept: Preserve Evidence, Then Contain

The most common IR mistake is destroying evidence by acting too quickly. Reimaging a compromised host or blocking a C2 IP before collecting forensic evidence may remove your only chance to understand the full scope of the breach.

Wrong order:
  1. Reimage the infected host → evidence destroyed
  2. Block the C2 IP → attacker notices and burns other infrastructure
  3. Reset passwords → attacker's backdoor user still exists (you didn't find it yet)

Right order:
  1. Detect → assess scope (how many systems? which credentials? what data?)
  2. Collect evidence (memory dump, disk image, logs) while the host is still running
  3. Contain simultaneously across ALL affected systems (not incrementally)
  4. Eradicate (remove malware, revoke credentials, patch vulnerabilities)
  5. Recover (restore from clean backups, rebuild from IaC)
  6. Post-incident review

The concept of the order of volatility:

Digital evidence degrades in a specific order — most volatile first:

CPU registers, cache → RAM (processes, network connections, decrypted keys) → swap/hibernation
→ disk (files, logs) → remote logs (SIEM, cloud audit logs) → backups

Memory must be collected first because it contains: running process state, decrypted keys (disk encryption key in RAM), active network connections, injected shellcode (only exists in memory, not on disk), and recently used credentials.

The concept of an IR retainer

Most organisations shouldn't build a full IR team. The concept is a retainer — a pre-negotiated contract with a DFIR firm (Mandiant, CrowdStrike Services, Secureworks) that guarantees a defined response time and forensic capability when you call them. The retainer means you don't negotiate contracts during an active incident. Budget this as an insurance cost.


The Pipeline Mental Model

The most important concept for understanding how security fits into modern infrastructure is the end-to-end pipeline view. Every security control should be understood as a gate or monitor at a specific point in the flow:

Developer           Source           Build              Artifact           Deploy          Runtime
   │                Control          System             Registry              │               │
   ▼                   │                │                   │                 ▼               ▼
  IDE                 git              CI                Docker/OCI          K8s           Production
  linting           branch          SAST/DAST            signing          admission       EDR/SIEM
  pre-commit        protection      dependency          SBOM attach        control         runtime
  hooks             signed          scanning            vulnerability       policy          detection
                    commits         secret              scanning            enforce
                    PR review       detection           dist. to
                    CODEOWNERS      SLSA                private
                                   provenance          registry
                                   signing

The key insighta vulnerability found in the IDE (via linting or static analysis) costs essentially nothing to fix. The same vulnerability found in production costs orders of magnitude more — in remediation time, breach risk, and potential customer impact. Security controls shifted left (towards the developer) are cheaper per finding.

But "shift left" does not mean "only left": you still need runtime detection (EDR, SIEM) because:

  • Some vulnerabilities are not detectable statically (only appear at runtime)
  • The developer pipeline can itself be compromised (supply chain attack)
  • Configuration drift and misuse in production is invisible to pre-deploy tooling

Every security tool you encounter lives somewhere in this pipeline. When you learn a new tool, ask: where does it sit in the pipeline, what does it receive as input, and what does it produce as output? That question alone puts any tool in context regardless of what it's called.


Quick Reference: Concept → Tool Category Mapping

ConceptWhat it solvesTool examples (all interchangeable)
Least privilege + short-lived tokensCredential theft, lateral movementAWS IAM + OIDC, GCP Workload Identity, Vault
Stateful network filteringPerimeter control, default denynftables, iptables, AWS Security Groups, GCP VPC firewall
Service mesh + mTLSEast-west encryption, service identityIstio, Linkerd, Envoy, Consul Connect
Behavioural detection (host)Novel malware, LOLBins, fileless attacksCrowdStrike, SentinelOne, Tetragon, Falco
Syscall filtering (container)Kernel attack surface reductionseccomp, AppArmor, SELinux
Image signing and provenanceSupply chain integritycosign, Sigstore, Notary v2
SIEM / log correlationMulti-source threat detectionSplunk, Elastic SIEM, Chronicle, Datadog
SBOM + vulnerability scanningKnown CVE in dependenciesGrype, Trivy, Snyk, Dependabot
Secret lifecycle managementCredential exposure and rotationVault, AWS Secrets Manager, GCP Secret Manager
IaC drift detectionUnauthorized infrastructure changesTerraform plan, Pulumi preview, Driftctl
OIDC-federated CI credentialsStatic secret elimination in CIGitHub OIDC + AWS IRSA, GCP Workload Identity
AEAD encryptionConfidentiality + integrity in oneAES-256-GCM, ChaCha20-Poly1305
Certificate transparencyUnauthorized certificate issuancecrt.sh, Google CT log, certstream
Admission control (K8s)Enforce policy at deploy timeOPA/Gatekeeper, Kyverno, PSA
Runtime file integrityDetect filesystem tamperingAIDE, Tripwire, IMA, Falco file rules