Linux Security Fundamentals
Breadth layernotes-security-core-knowledge.md Also see: Linux Privilege Escalation · Linux Exploits · Linux Hardening
Secure Boot Chain
Secure Boot ensures every component in the boot path is cryptographically signed before execution. Each stage verifies the next; an unsigned or tampered binary halts the boot.
Memory hookSecure Boot is a "chain of trust": each link checks the next before letting go. Firmware verifies the bootloader, which verifies the kernel, which verifies the next stage — like a bucket brigade where everyone checks the badge of the person they hand to. Break any link (unsigned/tampered binary) and the brigade stops. The whole thing roots in keys baked into firmware (UEFI db = trusted, dbx = revoked/banned). Shim exists for a political reason worth knowing: hardware only trusts Microsoft's key, so distros ship a tiny Microsoft-signed Shim that then trusts the distro's key — letting Linux boot on Secure Boot hardware without Microsoft signing every kernel.
UEFI Firmware (ROM)
└─ verifies → Shim (signed by Microsoft CA — distro-provided)
└─ verifies → GRUB2 (signed by distro key embedded in Shim's MokList)
└─ verifies → vmlinuz kernel (signed by distro key)
└─ verifies → initramfs (optional, depends on distro)Key concepts
| Component | Role |
|---|---|
| UEFI db | Database of trusted signing certificates stored in firmware NVRAM |
| UEFI dbx | Revocation database — hashes/certs of banned bootloaders (e.g. BootHole-vulnerable GRUBs) |
| Shim | Thin first-stage bootloader signed by the Microsoft UEFI CA; allows distros to manage their own keys without enrolling them into the UEFI db |
| MokManager | Interactive tool to enroll a Machine Owner Key (MOK) — required for custom/unsigned kernels and out-of-tree modules |
| mokutil | CLI to manage MOK enrollment |
# Check Secure Boot status
mokutil --sb-state
# Check enrolled MOKs
mokutil --list-enrolled
# Enroll a new key (signs out-of-tree modules e.g. DKMS, VirtualBox)
openssl req -new -x509 -newkey rsa:2048 -keyout mok.key -out mok.crt -days 36500 -subj "/CN=MOK/"
openssl x509 -in mok.crt -out mok.der -outform DER
mokutil --import mok.der # prompts for one-time password; enroll on next reboot
sbsign --key mok.key --cert mok.crt --output vmlinuz-signed vmlinuz
# Sign a kernel module
/usr/src/linux-headers-$(uname -r)/scripts/sign-file sha256 mok.key mok.crt mymodule.koWhat Secure Boot does NOT protect against
- Attacks after the kernel hands off to userspace
- Signed but vulnerable kernels (dbx revocation list lags behind disclosure)
- Physical attacks (Evil Maid can reset NVRAM if firmware password is not set)
- DMA attacks on PCIe devices (mitigated by IOMMU — see below)
Measured Boot and TPM
Secure Boot = won't boot untrusted code.
Measured Boot = records exactly what booted, so you can prove it later.
Memory hookSecure Boot is a bouncer, Measured Boot is a security camera. The bouncer (Secure Boot) blocks anything unsigned at the door. The camera (Measured Boot) doesn't block anything — it records a tamper-proof hash of every stage into the TPM's PCRs so you can later prove what booted. They combine powerfully via sealing: you can seal a secret (like a LUKS disk-encryption key) to a specific set of PCR values, so the disk only auto-unlocks if the machine booted exactly the expected software — change the kernel or bootloader and the PCRs differ and the secret stays locked. Mnemonic: Secure Boot prevents, Measured Boot proves; sealing ties a secret to "the machine booted clean."
TPM Platform Configuration Registers (PCRs)
The TPM is a tamper-resistant chip that stores SHA-256 measurements. Each PCR is a one-way accumulator: PCR_new = SHA256(PCR_old || new_value).
| PCR | Measures |
|---|---|
| 0 | UEFI firmware code |
| 1 | UEFI firmware configuration |
| 2 | UEFI option ROMs |
| 4 | MBR / bootloader |
| 5 | GPT partition table |
| 7 | Secure Boot state (db, dbx, policy) |
| 8–9 | GRUB config, kernel command line |
| 10 | IMA measurement log |
| 11–16 | OS / application use (configurable) |
# Read all PCR values
tpm2_pcrread sha256
# Read specific PCR
tpm2_pcrread sha256:0,7,10
# Seal a secret to current PCR state (unlocks only if PCRs match at boot)
tpm2_createprimary -G ecc -c primary.ctx
tpm2_create -C primary.ctx -L "pcr:sha256:0,7" -u sealed.pub -r sealed.priv -i secret.txt
tpm2_load -C primary.ctx -u sealed.pub -r sealed.priv -c sealed.ctx
tpm2_unseal -c sealed.ctx -p pcr:sha256:0,7 # only succeeds if PCR 0 and 7 matchTPM-backed LUKS (systemd-cryptenroll)
The most practical use: disk decryption key is sealed to PCRs 0+7 (firmware + Secure Boot state). If someone tampers with the bootloader or disables Secure Boot, the PCRs change, the TPM refuses to unseal the key, and the disk stays encrypted.
# Enroll TPM2 as LUKS key slot (requires systemd >= 248)
systemd-cryptenroll --tpm2-device=auto --tpm2-pcrs=0+7 /dev/sda2
# Test: cryptsetup open should now work without passphrase if PCRs match
cryptsetup open /dev/sda2 rootdm-verity — Block Device Integrity
dm-verity builds a Merkle tree of all blocks in a partition. At mount time, the kernel verifies each block's hash as it is read. Any single-bit modification is detected and the read fails. Used by Android verified boot, ChromeOS, and immutable Linux distros.
Root hash (single 32-byte value)
└── Hash tree level N
└── Hash tree level N-1
└── ...
└── Data blocks (4 KiB each)# Create a verity device from a read-only image
veritysetup format rootfs.img rootfs.verity
# → outputs: Root hash: <64-char hex>
# Verify without mounting
veritysetup verify rootfs.img rootfs.verity <root-hash>
# Open (kernel will verify every read)
veritysetup open rootfs.img rootveri rootfs.verity <root-hash>
mount -o ro /dev/mapper/rootveri /mnt
# In a production image build, embed the root hash in the kernel cmdline:
# GRUB: linux /vmlinuz ... ro verity.roothash=<hash> verity.hashdevice=/dev/sda3The root hash itself must be protected — typically signed with a key and verified as part of the Secure Boot chain (stored in the kernel image or passed via UEFI variables).
Immutable Root Filesystem
An immutable root makes it impossible for an attacker who gains code execution to persist changes across reboots.
Approach 1 — tmpfs overlay (container-style)
/ → read-only (dm-verity backed)
/tmp → tmpfs (RAM, wiped on reboot)
/var → separate rw partition or overlayfs upper layer
/home → separate rw partition# Mount root read-only at boot (kernel cmdline)
ro
# In /etc/fstab — volatile runtime directories on tmpfs
tmpfs /tmp tmpfs defaults,noexec,nosuid,nodev,size=256m 0 0
tmpfs /var/tmp tmpfs defaults,noexec,nosuid,nodev,size=64m 0 0Approach 2 — OSTree / image-based OS
Distributions like Fedora CoreOS, RHEL CoreOS, and Fedora Silverblue use OSTree:
- The entire OS tree is versioned and content-addressed (like git)
/usris a read-only bind mount from the OSTree commit- Updates replace the commit atomically and can be rolled back
/etcand/varare writable overlays on top
# Check current deployment
rpm-ostree status
# Upgrade
rpm-ostree upgrade
# Rollback to previous version
rpm-ostree rollbackApproach 3 — chattr +i for individual files
# Make a file immutable (even root cannot write/delete until flag removed)
chattr +i /etc/passwd /etc/shadow /etc/sudoers
# Check immutable flag
lsattr /etc/passwd
# Remove flag (requires CAP_LINUX_IMMUTABLE)
chattr -i /etc/passwdUse chattr +i for files that should never change at runtime: /etc/resolv.conf when you want a static DNS configuration, critical cron files, or /etc/hosts.
Memory hookWhy this beats normal permissions — it stops root too. Standard file permissions don't restrain root; root can write anything. The immutable flag (
+i) is enforced by the filesystem layer, so even a root process gets "Operation not permitted" trying to modify or delete the file. Locking/etc/passwd,/etc/shadow, and/etc/sudoersdefeats the entire family of "append a UID-0 user" / "rewrite root's hash" attacks — an attacker who lands as root has to first notice the flag and runchattr -i(which needsCAP_LINUX_IMMUTABLE), and that extra step is something you can audit and alert on. Caveat: it doesn't stop bugs that write under the VFS layer — e.g. DirtyCOW wrote straight to the page cache — and a determined root can clear the flag, so treat it as a speed-bump and tripwire, not a wall.
Preventing Kernel Module Loads
Allowing arbitrary kernel module loading is equivalent to allowing arbitrary code execution in ring 0.
1. Signed modules only (build-time)
# Check kernel config
grep CONFIG_MODULE_SIG /boot/config-$(uname -r)
# CONFIG_MODULE_SIG=y → signing support compiled in
# CONFIG_MODULE_SIG_FORCE=y → unsigned modules are REJECTED
# CONFIG_MODULE_SIG_ALL=y → all in-tree modules signed at build
# CONFIG_MODULE_SIG_KEY="..." → path to signing key
# Check if a module is signed
modinfo mymodule.ko | grep sig2. Kernel lockdown mode
Lockdown is a Linux Security Module (LSM) that restricts even root from actions that could compromise kernel integrity.
Modes:
none → disabled
integrity → prevents modifications that could subvert integrity
(no /dev/mem writes, no kprobes, no module loading without signature)
confidentiality → integrity + prevents reading kernel memory
(no hibernation, no PCI BAR access)# Check current lockdown mode
cat /sys/kernel/security/lockdown
# Set at boot via kernel cmdline
lockdown=integrity
# or
lockdown=confidentiality
# When Secure Boot is active, some kernels automatically enter integrity lockdown3. Disable all module loading after boot
# One-way — no new modules can be loaded until reboot
# (already-loaded modules remain; cannot be undone without reboot)
echo 1 > /proc/sys/kernel/modules_disabled
# or via sysctl:
sysctl -w kernel.modules_disabled=14. Blacklist specific modules
# /etc/modprobe.d/blacklist.conf
blacklist usb_storage # prevent USB storage devices
blacklist firewire_core # prevent DMA via FireWire
blacklist thunderbolt # prevent DMA via Thunderbolt
# Stronger: replace the module command with /bin/true
# prevents loading even if requested explicitly
install usb_storage /bin/true
install cramfs /bin/true
install freevxfs /bin/true
install jffs2 /bin/true
install hfs /bin/true
install hfsplus /bin/true
install squashfs /bin/true
install udf /bin/true
# Apply without reboot
depmod -a5. IOMMU — prevent DMA attacks
Even a read-only kernel can be compromised if a PCIe device (NIC, GPU, Thunderbolt) can DMA into arbitrary physical memory.
# Enable IOMMU in GRUB (/etc/default/grub)
GRUB_CMDLINE_LINUX="intel_iommu=on iommu=pt"
# or for AMD:
GRUB_CMDLINE_LINUX="amd_iommu=on iommu=pt"
update-grub # Debian/Ubuntu
grub2-mkconfig -o /boot/grub2/grub.cfg # RHEL/CentOS
# Verify IOMMU is active
dmesg | grep -i iommunftables — Stateful Firewall
nftables replaces iptables, ip6tables, arptables, and ebtables with a single framework. It has a cleaner syntax, better performance via JIT compilation, and native set/map support.
Core concepts
Table → namespace (inet = IPv4+IPv6)
└── Chain → hook point (input, forward, output, prerouting, postrouting)
└── Rule → condition + verdict (accept, drop, reject, log, jump)
└── Set/Map → efficient lookup for IPs, ports, etc.Minimal hardened ruleset
# /etc/nftables.conf
table inet filter {
# Sets — update dynamically without flushing ruleset
set blocked_ips {
type ipv4_addr
flags dynamic, timeout
timeout 1h
}
chain input {
type filter hook input priority 0; policy drop; # default DENY
# Allow established/related connections (stateful)
ct state established,related accept
# Drop invalid packets
ct state invalid drop
# Allow loopback
iifname "lo" accept
# ICMP — allow ping but rate-limit to prevent flooding
icmp type echo-request limit rate 5/second accept
icmpv6 type { echo-request, nd-neighbor-solicit, nd-router-advert, nd-neighbor-advert } accept
# SSH — only from management subnet
ip saddr 10.0.1.0/24 tcp dport 22 accept
# Drop blocked IPs
ip saddr @blocked_ips drop
# Log and drop everything else
log prefix "nft-drop-input: " drop
}
chain forward {
type filter hook forward priority 0; policy drop;
}
chain output {
type filter hook output priority 0; policy accept;
# Block outbound to common C2 ports (optional egress filtering)
tcp dport { 4444, 4445, 5555, 8080 } drop
}
}
# Persistence
# systemctl enable nftables
# nft -f /etc/nftables.conf
# Reload live
nft -f /etc/nftables.conf
nft list rulesetDynamic brute-force blocking
table inet filter {
set brute_force {
type ipv4_addr
flags dynamic, timeout
timeout 5m # auto-expire after 5 minutes
}
chain input {
type filter hook input priority 0; policy drop;
ct state established,related accept
iifname "lo" accept
# Count new SSH connections; add offending IP to set after 5 attempts
tcp dport 22 ct state new \
add @brute_force { ip saddr limit rate 5/minute burst 5 packets } accept
ip saddr @brute_force drop
tcp dport 22 ct state new accept
}
}Useful nft commands
nft list ruleset # show full ruleset
nft list table inet filter # show specific table
nft list set inet filter blocked_ips # show set contents
nft add element inet filter blocked_ips { 1.2.3.4 } # add element
nft delete element inet filter blocked_ips { 1.2.3.4 } # remove element
nft monitor # watch rule events live
nft -c -f /etc/nftables.conf # dry-run / syntax checkMandatory Access Control
Memory hookDAC vs MAC: "the owner decides" vs "the policy decides." Normal Linux permissions are Discretionary Access Control — the owner of a file decides who can read it (and root can do anything), so a compromised root process is game over. Mandatory Access Control adds a second, system-wide policy layer that even root can't override: a process is confined to exactly what its policy allows, regardless of file ownership. So if your web server is popped, MAC can stop it from reading
/etc/shadoweven though it's running as root, because the policy says "the httpd type may not touch the shadow type." Mnemonic: DAC = discretion of the owner, MAC = mandate of the system; MAC is what contains a compromised root.SELinux vs AppArmor — the two implementations: SELinux labels every object (
user:role:type:level) and is extremely powerful but complex (RHEL/Fedora/Android). AppArmor confines by file path in human-readable profiles — simpler, less granular (Ubuntu/SUSE). Mnemonic: SELinux = labels (powerful, painful), AppArmor = paths (simple, readable).
SELinux
SELinux labels every process (subject) and file/socket/device (object) with a security context: user:role:type:level. Policy defines which type can access which type.
# Check mode
getenforce # Enforcing / Permissive / Disabled
sestatus
# Temporarily switch to permissive (log but don't deny)
setenforce 0
# Check label on a file or process
ls -Z /etc/passwd
ps -Z -p $$
# Find denials in audit log
ausearch -m AVC -ts recent
sealert -a /var/log/audit/audit.log # human-readable explanations
# Generate allow rule from denial (test in permissive first)
audit2allow -M mymodule < /var/log/audit/audit.log
semodule -i mymodule.pp
# Restore default labels (fixes "wrong context" errors after manual copy)
restorecon -Rv /var/www/html
# Persistent mode change (/etc/selinux/config)
SELINUX=enforcingAppArmor
AppArmor uses path-based profiles. Simpler than SELinux but less expressive.
# Status
aa-status
apparmor_status
# Modes
aa-complain /usr/bin/myapp # log but don't deny (like SELinux permissive)
aa-enforce /usr/bin/myapp # enforce profile
# Generate a profile interactively
aa-genprof /usr/bin/myapp
# → run the application, then answer prompts
# Check what's being blocked
grep DENIED /var/log/syslog | grep apparmor
# Profile location
ls /etc/apparmor.d/
# Example minimal profile
# /etc/apparmor.d/usr.bin.myapp
/usr/bin/myapp {
#include <abstractions/base>
/etc/myapp.conf r,
/var/lib/myapp/** rw,
/proc/self/status r,
deny /etc/shadow r,
}seccomp — Syscall Filtering
seccomp (Secure Computing Mode) filters syscalls a process can make. Used extensively by Docker, Chrome, and systemd.
Memory hookseccomp shrinks the kernel attack surface a process can reach. Every syscall is a door from a process into the kernel, and kernel exploits are reached through syscalls. seccomp is a filter saying "this process may only call these specific syscalls" — so even if the app is compromised, the attacker can't invoke the dangerous syscall their kernel exploit needs (e.g.
ptrace,mount,kexec_load, module loading). Docker's default profile blocks ~44 such syscalls, which is why many container escapes assume that profile is disabled. Think of the three together: capabilities trims what root powers a process has, seccomp trims which syscalls it can make, MAC trims which objects it can touch — three independent walls around a process.
Modes
only read, write, _exit, sigreturn allowed. Almost never used directly.
BPF program decides per-syscall. This is what Docker/Podman use.
# Check if a process has a seccomp filter
grep -i seccomp /proc/$$/status
# Seccomp: 2 → BPF filter mode
# Docker default seccomp profile blocks ~44 syscalls including:
# ptrace, mount, umount2, kexec_load, create_module, init_module,
# delete_module, pivot_root, chroot, clone3 (partial)
# Use a custom profile with Docker
docker run --security-opt seccomp=/path/to/profile.json myimage
# Disable seccomp (testing only)
docker run --security-opt seccomp=unconfined myimage
# systemd service hardening with seccomp
# /etc/systemd/system/myservice.service
[Service]
SystemCallFilter=@system-service # predefined set of safe syscalls
SystemCallFilter=~@privileged # exclude privileged calls
SystemCallErrorNumber=EPERM # return permission error instead of SIGSYSseccomp profile structure (Docker JSON)
{
"defaultAction": "SCMP_ACT_ERRNO",
"architectures": ["SCMP_ARCH_X86_64"],
"syscalls": [
{
"names": ["read", "write", "open", "close", "stat", "mmap", "exit_group"],
"action": "SCMP_ACT_ALLOW"
},
{
"names": ["ptrace", "kexec_load", "init_module"],
"action": "SCMP_ACT_ERRNO"
}
]
}Binary Reduction and Attack Surface
Every binary on the system is a potential exploitation vector. Minimising the set of installed executables reduces the tools available to an attacker post-compromise.
Find what's installed
# Debian/Ubuntu — list all installed packages with sizes
dpkg-query -W --showformat='${Installed-Size}\t${Package}\n' | sort -rn | head -50
# RHEL/CentOS
rpm -qa --queryformat '%{SIZE}\t%{NAME}\n' | sort -rn | head -50
# Count installed packages
dpkg --get-selections | wc -l
rpm -qa | wc -lRemove unnecessary packages
# Common packages to question on a server (keep only what's needed)
apt remove --purge telnet ftp rsh-client rsh-server xinetd talk ntalk \
nis yp-tools tftp tftpd atftpd atftpd-hpa tcpd nfs-kernel-server \
nfs-common rpcbind portmap sendmail postfix exim4 avahi-daemon cups \
bluetooth bluez gcc make perl python3 git wget curl # only remove if truly unused
# After removal, clean orphans
apt autoremove --purgeAudit SUID and SGID binaries
SUID/SGID bits cause a binary to execute with the owner's (often root's) privileges, regardless of who runs it. Every SUID binary is a potential privilege escalation vector.
# Find all SUID binaries
find / -perm -4000 -type f 2>/dev/null | sort
# Find all SGID binaries
find / -perm -2000 -type f 2>/dev/null | sort
# Compare against expected baseline (CIS Benchmark or distro default)
dpkg --verify 2>/dev/null | grep -v '^$' # Debian — reports changed files
# Remove SUID bit from binaries you don't need setuid for
chmod u-s /usr/bin/chsh
chmod u-s /usr/bin/newgrp
chmod u-s /usr/bin/wall
# Keep: sudo, su, passwd, ping (or use capabilities instead)
# Replace SUID with capabilities where possible
# Instead of SUID root on ping:
setcap cap_net_raw+ep /usr/bin/ping
chmod u-s /usr/bin/pingRemove or restrict world-writable directories
# Find world-writable files
find / -xdev -type f -perm -o+w 2>/dev/null
# Find world-writable directories (excluding sticky-bit set ones)
find / -xdev -type d -perm -o+w ! -perm -1000 2>/dev/nullStrip debug symbols from binaries
# Check if a binary has debug symbols
file /usr/bin/sshd
# → ELF 64-bit ... not stripped ← symbols present
# → ELF 64-bit ... stripped ← already stripped
# Strip debug info
strip --strip-unneeded /usr/local/bin/myapp
# Build release binaries already stripped (Go)
go build -ldflags="-s -w" -o myapp .
# Python — compile .py to .pyc and remove source
python3 -OO -m compileall .
# Remove man pages, documentation, and locale data (container optimisation)
find /usr/share/doc -depth -type f ! -name 'copyright' -delete
find /usr/share/man /usr/share/locale -type f -deleteSBOM — Software Bill of Materials
An SBOM is a machine-readable inventory of every component, library, and dependency in a software artifact. Used for vulnerability management, license compliance, and supply chain security.
Formats
| Format | Standard body | Common use |
|---|---|---|
| SPDX | Linux Foundation | open-source, broad tooling |
| CycloneDX | OWASP | security-focused, richer vulnerability data |
| SWID | ISO/IEC 19770 | enterprise software inventory |
Generating an SBOM with syft
# Install syft (Anchore)
curl -sSfL https://raw.githubusercontent.com/anchore/syft/main/install.sh | sh -s -- -b /usr/local/bin
# Generate SBOM for a directory (source scan)
syft dir:/path/to/project -o spdx-json > sbom.spdx.json
syft dir:/path/to/project -o cyclonedx-json > sbom.cdx.json
# Generate SBOM for a container image
syft docker:myapp:latest -o spdx-json > sbom.spdx.json
syft oci-dir:/path/to/extracted/image -o cyclonedx-json > sbom.cdx.json
# Generate SBOM for a Debian/RPM system
syft dir:/ -o spdx-json > host-sbom.spdx.jsonScanning an SBOM for vulnerabilities with grype
# Install grype
curl -sSfL https://raw.githubusercontent.com/anchore/grype/main/install.sh | sh -s -- -b /usr/local/bin
# Scan an SBOM
grype sbom:sbom.spdx.json
# Scan a container image directly (grype fetches its own SBOM internally)
grype myapp:latest
# Only show fixable vulnerabilities at high/critical severity
grype myapp:latest --only-fixed -f high
# Output as JSON for CI gate
grype sbom:sbom.cdx.json -o json | jq '.matches[].vulnerability | select(.severity=="Critical")'Scanning with trivy (all-in-one)
# Filesystem scan (OS packages + language dependencies)
trivy fs /path/to/project
# Container image scan
trivy image myapp:latest
# Generate CycloneDX SBOM and scan in one step
trivy image --format cyclonedx --output sbom.cdx.json myapp:latest
trivy sbom sbom.cdx.json
# Scan for misconfigurations (Dockerfile, Kubernetes manifests, IaC)
trivy config /path/to/manifests/
# Exit code 1 if CRITICAL vulnerabilities found (use in CI)
trivy image --exit-code 1 --severity CRITICAL myapp:latestCI/CD integration pattern
# GitHub Actions — generate SBOM and block on critical CVEs
- name: Generate SBOM
run: |
syft packages dir:. -o cyclonedx-json > sbom.cdx.json
- name: Scan SBOM
run: |
grype sbom:sbom.cdx.json --fail-on critical
- name: Upload SBOM as artifact
uses: actions/upload-artifact@v3
with:
name: sbom
path: sbom.cdx.json
- name: Attest SBOM (cosign)
run: |
cosign attest --predicate sbom.cdx.json --type cyclonedx $IMAGE_REFSBOM in a Dockerfile build
FROM debian:bookworm-slim AS builder
RUN apt-get update && apt-get install -y --no-install-recommends \
build-essential && rm -rf /var/lib/apt/lists/*
COPY . /app
WORKDIR /app
RUN make release
# Generate SBOM at build time
FROM anchore/syft:latest AS sbom
COPY --from=builder /app/bin/myapp /myapp
RUN syft packages /myapp -o spdx-json > /sbom.spdx.json
FROM gcr.io/distroless/base-debian12
COPY --from=builder /app/bin/myapp /myapp
# SBOM stored separately and attached via cosign attestFile Integrity Monitoring (AIDE)
AIDE (Advanced Intrusion Detection Environment) takes a baseline snapshot of filesystem attributes (hashes, permissions, ownership, timestamps) and alerts on any changes.
# Install
apt install aide
# Configure monitored paths (/etc/aide/aide.conf or /etc/aide.conf)
# Default rules already cover /bin, /sbin, /usr, /etc, /root
# Add custom rules:
/opt/myapp NORMAL # monitor with default attributes
!/var/log # exclude (too noisy)
!/proc # exclude virtual filesystem
!/sys
# Initialise baseline (after a known-clean state)
aide --init
mv /var/lib/aide/aide.db.new /var/lib/aide/aide.db
# Run check against baseline
aide --check
# After intentional changes (e.g. package update), update baseline
aide --update
mv /var/lib/aide/aide.db.new /var/lib/aide/aide.db
# Automate daily check
# /etc/cron.daily/aide
#!/bin/bash
/usr/bin/aide --check 2>&1 | mail -s "AIDE report $(hostname)" security@example.comPAM Hardening
PAM (Pluggable Authentication Modules) controls how authentication works for every service that calls pam_*.
# Key PAM config files
/etc/pam.d/sshd
/etc/pam.d/login
/etc/pam.d/sudo
/etc/pam.d/common-auth # Debian/Ubuntu — shared auth stack
/etc/pam.d/system-auth # RHEL/CentOS — shared auth stackEnforce password quality (pam_pwquality)
# /etc/security/pwquality.conf
minlen = 16
minclass = 3 # at least 3 character classes (upper/lower/digits/special)
maxrepeat = 3 # no more than 3 repeated chars in a row
dcredit = -1 # require at least 1 digit
ucredit = -1 # require at least 1 uppercase
lcredit = -1 # require at least 1 lowercase
ocredit = -1 # require at least 1 special character
difok = 8 # at least 8 chars different from old password# /etc/pam.d/common-password (Debian) — add pam_pwquality BEFORE pam_unix
password requisite pam_pwquality.so retry=3
password [success=1 default=ignore] pam_unix.so obscure use_authtok try_first_pass sha512 shadowLock accounts after failed attempts (pam_faillock)
# /etc/security/faillock.conf
deny = 5 # lock after 5 failures
fail_interval = 900 # within 15 minutes
unlock_time = 600 # locked for 10 minutes
# /etc/pam.d/common-auth
auth required pam_faillock.so preauth
auth sufficient pam_unix.so
auth [default=die] pam_faillock.so authfail
auth sufficient pam_faillock.so authsucc
# Check locked accounts
faillock --user username
# Unlock manually
faillock --user username --resetRestrict su to the wheel group
# /etc/pam.d/su
auth required pam_wheel.so use_uid# Only members of group 'wheel' can use su
usermod -aG wheel adminuserauditd — Kernel Audit Framework
auditd listens to the kernel audit subsystem and records syscalls, file accesses, and system events. Unlike syslog, it is harder to tamper with because it runs in kernel space.
# Status
systemctl status auditd
auditctl -s # show runtime status
# View recent events
ausearch -ts recent -i # last 10 minutes, human-readable
aureport --summary # summary of all event typesKey audit rules
# /etc/audit/rules.d/hardening.rules
## Immutable rules (load these last — prevents runtime modification)
-e 2
## Monitor authentication files
-w /etc/passwd -p wa -k identity
-w /etc/shadow -p wa -k identity
-w /etc/group -p wa -k identity
-w /etc/sudoers -p wa -k privilege_escalation
-w /etc/sudoers.d/ -p wa -k privilege_escalation
## Monitor PAM configuration
-w /etc/pam.d/ -p wa -k pam_config
## Monitor SSH config
-w /etc/ssh/sshd_config -p wa -k sshd_config
## Privilege escalation — setuid/setgid syscalls
-a always,exit -F arch=b64 -S setuid -S setgid -S setreuid -S setregid -k privilege_escalation
-a always,exit -F arch=b64 -S setresuid -S setresgid -k privilege_escalation
## Kernel module operations
-w /sbin/insmod -p x -k module_load
-w /sbin/rmmod -p x -k module_load
-w /sbin/modprobe -p x -k module_load
-a always,exit -F arch=b64 -S init_module -S finit_module -k module_load
## Network configuration changes
-a always,exit -F arch=b64 -S sethostname -S setdomainname -k network_change
-w /etc/issue -p wa -k network_change
-w /etc/hosts -p wa -k network_change
-w /etc/network/ -p wa -k network_change
## Execution auditing (noisy — tune for your environment)
-a always,exit -F arch=b64 -S execve -k exec
## File deletion
-a always,exit -F arch=b64 -S unlink -S unlinkat -S rename -S renameat -F auid>=1000 -k file_delete# Load rules
auditctl -R /etc/audit/rules.d/hardening.rules
# or restart auditd
service auditd restart
# Search for specific events
ausearch -k privilege_escalation -ts today
ausearch -k module_load -i
ausearch -f /etc/passwd
# Generate report of failed syscalls
aureport --syscall --failedKernel Hardening Sysctls Reference
Complete reference for /etc/sysctl.d/99-hardening.conf:
## Network — attack surface reduction
net.ipv4.ip_forward = 0 # disable IP routing
net.ipv4.conf.all.accept_source_route = 0 # don't accept source-routed packets
net.ipv4.conf.default.accept_source_route = 0
net.ipv4.conf.all.accept_redirects = 0 # ignore ICMP redirect messages
net.ipv4.conf.default.accept_redirects = 0
net.ipv6.conf.all.accept_redirects = 0
net.ipv4.conf.all.send_redirects = 0 # don't send redirects (not a router)
net.ipv4.conf.all.rp_filter = 1 # reverse-path filter (uRPF strict mode)
net.ipv4.conf.default.rp_filter = 1
net.ipv4.tcp_syncookies = 1 # SYN flood protection
net.ipv4.conf.all.log_martians = 1 # log packets with impossible addresses
net.ipv4.conf.default.log_martians = 1
net.ipv4.icmp_echo_ignore_broadcasts = 1 # don't respond to broadcast pings (smurf amp)
net.ipv4.icmp_ignore_bogus_error_responses = 1
net.ipv6.conf.all.accept_ra = 0 # don't accept IPv6 router advertisements
## Network — hardened TCP
net.ipv4.tcp_timestamps = 0 # disable timestamps (reduces uptime disclosure)
net.ipv4.conf.all.proxy_arp = 0
## Kernel — memory and information exposure
kernel.randomize_va_space = 2 # full ASLR (stack, heap, VDSO, mmap)
kernel.dmesg_restrict = 1 # only root can read dmesg
kernel.kptr_restrict = 2 # hide kernel symbols from /proc/kallsyms
kernel.perf_event_paranoid = 3 # disable perf for non-root
kernel.unprivileged_bpf_disabled = 1 # only root can load eBPF programs
net.core.bpf_jit_harden = 2 # harden BPF JIT (disable kallsyms in JIT)
kernel.yama.ptrace_scope = 1 # only parent can ptrace child
# =2: only root; =3: no ptrace at all
## Filesystem
fs.suid_dumpable = 0 # no core dumps from SUID processes
fs.protected_symlinks = 1 # prevent symlink following attacks in sticky dirs
fs.protected_hardlinks = 1 # prevent hardlink-based attacks
fs.protected_fifos = 2 # prevent FIFO attacks in world-writable dirs
fs.protected_regular = 2 # prevent regular-file attacks in world-writable dirs
## Magic SysRq (disable on production)
kernel.sysrq = 0
## Core dumps — disable entirely
kernel.core_pattern = |/bin/false# Apply without reboot
sysctl -p /etc/sysctl.d/99-hardening.conf
# Verify a specific value
sysctl kernel.randomize_va_spaceSSH Hardening
# /etc/ssh/sshd_config
# Authentication
PermitRootLogin no # never allow root SSH
AuthenticationMethods publickey # key-only; no password auth
PasswordAuthentication no
ChallengeResponseAuthentication no
KerberosAuthentication no
GSSAPIAuthentication no
# Restrict who can log in
AllowGroups sshusers # only members of 'sshusers' group
# or
AllowUsers alice bob deploy@10.0.1.5 # specific users; optional source IP
# Protocol and algorithms
Protocol 2
KexAlgorithms curve25519-sha256,diffie-hellman-group16-sha512,diffie-hellman-group18-sha512
Ciphers chacha20-poly1305@openssh.com,aes256-gcm@openssh.com,aes128-gcm@openssh.com
MACs hmac-sha2-512-etm@openssh.com,hmac-sha2-256-etm@openssh.com
HostKeyAlgorithms ssh-ed25519,rsa-sha2-512,rsa-sha2-256
# Session limits
MaxAuthTries 3
LoginGraceTime 30
MaxSessions 5
ClientAliveInterval 300 # send keepalive every 5 min
ClientAliveCountMax 2 # disconnect after 10 min idle
# Disable unnecessary features
X11Forwarding no
AllowAgentForwarding no
AllowTcpForwarding no # disable unless you need SSH tunnelling
PrintMotd no
PermitEmptyPasswords no
PermitUserEnvironment no # prevent env-based bypasses
StrictModes yes # check file permissions before accepting key auth
# Banner
Banner /etc/ssh/ssh_banner # show legal warning before auth
# Logging
LogLevel VERBOSE # logs key fingerprints (important for audit)
SyslogFacility AUTH# Generate an Ed25519 host key (if not already present)
ssh-keygen -t ed25519 -f /etc/ssh/ssh_host_ed25519_key -N ""
# Remove weak host keys
rm -f /etc/ssh/ssh_host_dsa_key* /etc/ssh/ssh_host_rsa_key*
# Test config before reloading
sshd -t
# Reload
systemctl reload sshdFilesystem Mount Hardening
# /etc/fstab — add mount options to reduce attack surface
# /tmp — never execute binaries here, never create devices, never setuid
tmpfs /tmp tmpfs defaults,noexec,nosuid,nodev,size=512m 0 0
# /var/tmp — same
tmpfs /var/tmp tmpfs defaults,noexec,nosuid,nodev,size=128m 0 0
# /home — users should not be able to create device files or setuid binaries
/dev/sda5 /home ext4 defaults,nosuid,nodev 0 2
# /dev/shm — shared memory: noexec prevents shellcode injection via memfd
tmpfs /dev/shm tmpfs defaults,noexec,nosuid,nodev 0 0
# Separate partition for /var/log to prevent log flooding from filling /
/dev/sda6 /var/log ext4 defaults,noexec,nosuid,nodev 0 2
# Remount /proc with hidepid (hides other users' processes)
proc /proc proc defaults,hidepid=2,gid=proc 0 0
# Add services that need /proc access to 'proc' group:
usermod -aG proc www-dataMemory hookWhat
hidepid=2actually denies an attacker. By default any user can read/proc/<pid>/for every process on the box — including other users' full command lines (/proc/<pid>/cmdline) and environment variables (/proc/<pid>/environ). That's a goldmine: passwords passed on the command line (mysql -p hunter2), API tokens and DB creds stuffed into env vars, and a complete map of what's running to aim a privesc at.hidepid=2makes/proc/<pid>of other users invisible (not just unreadable — they can't even see the PID exists), so a low-priv foothold can no longer shoulder-surf root's processes for secrets. Thenoexeclines above matter for the same reason attackers love writable dirs:noexecon/tmp,/var/tmp, and/dev/shmstops the classic "drop my exploit in/tmpand run it" — they can write the file but not execute it.
# Verify current mount options
mount | grep -E "on /tmp|on /home|on /dev/shm"
findmnt --verifyTroubleshooting Cheat Sheet
System state snapshot
# What's running?
systemctl list-units --state=failed # failed services
systemctl status myservice
journalctl -u myservice -n 100 --no-pager
# What processes are running?
ps auxf # tree format
pstree -ap # compact tree
top / htop
# Who is logged in?
w
last | head -20 # login history
lastb | head -20 # failed loginsNetwork troubleshooting
# Current connections
ss -tulnp # listening ports + PID
ss -tunp # all TCP/UDP + PID
netstat -tulnp # legacy
# Active nftables ruleset
nft list ruleset
# Routing table
ip route show
ip -6 route show
# ARP table
ip neigh show
# Interface statistics
ip -s link show eth0
# DNS resolution
dig +trace example.com # full recursive trace
resolvectl query example.com # systemd-resolved
# Packet capture
tcpdump -i eth0 -nn -s0 'port 443'
tcpdump -i eth0 -nn 'host 1.2.3.4'Security investigation
# SUID binaries (quick check)
find /bin /sbin /usr -perm -4000 -type f 2>/dev/null
# World-writable files in critical directories
find /etc /usr /bin /sbin -xdev -perm -o+w -type f 2>/dev/null
# Recently modified files (last 24h)
find /etc /usr /bin /sbin /lib -newer /tmp/.ref -type f 2>/dev/null
# Loaded kernel modules
lsmod
modinfo <module_name>
# Capabilities on binaries
getcap -r / 2>/dev/null
# Open files / network sockets by process
lsof -i -n -P
lsof -p <pid>
# Check file integrity against package database
rpm -Va 2>/dev/null | grep -v "^....5" # RHEL: files not matching RPM checksums
debsums -c 2>/dev/null # Debian: files not matching dpkg checksums
# Systemd timer and cron jobs
systemctl list-timers --all
crontab -l
ls -la /etc/cron.* /var/spool/cron/
cat /etc/cron.d/*
# Sudoers check
visudo -c # validate syntax
sudo -l -U username # what a user can runKernel and boot
# Kernel version and build
uname -a
cat /proc/version
cat /boot/config-$(uname -r) | grep -E "CONFIG_MODULE_SIG|CONFIG_LOCKDOWN|CONFIG_IMA"
# Boot parameters actually in use
cat /proc/cmdline
# Secure Boot status
mokutil --sb-state
cat /sys/firmware/efi/efivars/SecureBoot-* # raw: byte 4 = 1 means enabled
# Kernel lockdown mode
cat /sys/kernel/security/lockdown
# Dmesg (requires dmesg_restrict=0 or root)
dmesg | grep -E "secure.boot|lockdown|module.sig|ima|selinux|apparmor"
# Check IMA log
cat /sys/kernel/security/ima/ascii_runtime_measurements | head -20Log quick triage
# Authentication failures
journalctl _SYSTEMD_UNIT=sshd.service | grep -i "invalid\|failed\|disconnect"
grep "Failed password\|Invalid user" /var/log/auth.log | awk '{print $9,$11}' | sort | uniq -c | sort -rn
# Privilege escalation
grep "sudo\|su " /var/log/auth.log | grep -v "session opened\|session closed"
ausearch -k privilege_escalation -ts today -i
# Kernel messages
journalctl -k --since "1 hour ago"
dmesg -T | grep -E "OOM|segfault|call trace" | tail -30
# SELinux/AppArmor denials
ausearch -m AVC,USER_AVC -ts today | head -50
grep "apparmor.*DENIED" /var/log/syslog | tail -30Interview Questions and Answers
UEFI firmware verifies the bootloader's signature against its key database. On most Linux systems: UEFI → Shim (signed by Microsoft CA, loaded first) → GRUB (signed by distro key embedded in Shim) → kernel (signed by distro key). Shim can also trust additional keys stored in MOK (Machine Owner Key) — mokutil --import adds them. Each link verifies the next; a compromised bootloader cannot load an unsigned kernel.
Secure Boot is binary: it checks signatures and either allows or blocks a component from loading. It's a gate. Measured Boot records (hashes) every component as it loads into TPM PCR registers — but it doesn't block anything. The measurements create a tamper-evident audit trail. The combination: Secure Boot blocks known-bad components; Measured Boot gives you evidence of what actually ran, enabling TPM-sealed secrets (only released if the expected PCR values match).
PCR (Platform Configuration Register) is a TPM register that stores a hash chain — each extend operation hashes the current value concatenated with new data, making it tamper-evident. PCR 0 = firmware, PCR 7 = Secure Boot state, PCR 10 = IMA measurements. To seal a LUKS key: systemd-cryptenroll --tpm2-device=auto --tpm2-pcrs=0+7 /dev/sda2 — the TPM only releases the key if PCR 0 and PCR 7 have the expected values (i.e., the same firmware and Secure Boot state as when the key was sealed). Rootkit installs change PCR values → key is locked.
dm-verity builds a Merkle tree over a block device — each data block is hashed, then pairs of hashes are hashed together, up to a single root hash stored in a signed superblock. On every read, the kernel verifies the block's hash against the tree. Any modification — even one bit — produces a hash mismatch and makes the block unavailable. Solves: runtime tampering detection for read-only partitions (rootfs, system image). Used in Android verified boot and immutable Linux systems.
kernel.modules_disabled=1 do? Can it be reversed?It prevents loading any new kernel modules — insmod, modprobe, and finit_module() all fail after this sysctl is set. This prevents an attacker with root from loading a kernel rootkit via a .ko file. It is one-way: once set to 1, it cannot be set back to 0 in the same boot session. The kernel enforces this because allowing reversal would defeat the purpose. It persists only until reboot; to make it permanent, set it early in the boot process (initramfs or kernel command line).
lockdown=integrity and lockdown=confidentiality?Both are kernel lockdown modes that restrict what root can do to the kernel. integrity: prevents modifications to the running kernel — no /dev/mem writes, no unsigned modules, no kexec with unsigned kernels, no raw PCI access. An attacker with root still can't patch kernel memory. confidentiality: everything in integrity PLUS prevents reading kernel memory and secrets from userspace — blocks /dev/mem reads, hibernation (which writes RAM to disk), and certain debug interfaces. Confidentiality is a superset.
Layered approach: (1) Secure Boot — kernel is signed, unsigned modules won't load; (2) kernel.modules_disabled=1 — set early in boot to prevent module loading at all; (3) Module signing — CONFIG_MODULE_SIG_FORCE=y — only modules signed with the kernel's embedded key load; (4) Kernel lockdown (lockdown=confidentiality) — prevents memory patching even by root; (5) IMA — measures modules before loading and enforces policy. Combining all five means even a compromised root account cannot alter kernel execution.
nftables is the modern replacement — single kernel subsystem vs iptables' separate ones (ip_tables, ip6tables, arptables). Table: a container for chains, associated with a network family (inet = IPv4+IPv6, ip, ip6). Chain: a sequence of rules with a hook (input/output/forward) and a default policy (accept/drop). Rule: a match expression + verdict (accept, drop, jump). nftables uses a JIT bytecode VM (more efficient), native set/map support, and atomic rule updates. Single nft command for all address families.
SELinuxlabel-based MAC (Mandatory Access Control). Every object and process has a label; policy defines which labels can interact. Very granular, very complex. Default on RHEL/CentOS/Fedora. Handles complex policies well but has a steep learning curve. AppArmor: path-based MAC. Policies reference filesystem paths, not labels. Simpler to write and understand. Default on Ubuntu/Debian/SUSE. Choose SELinux for maximum granularity on RHEL systems; AppArmor for simpler profiles on Ubuntu/containers where path-based access is sufficient.
seccomp (secure computing mode) is a kernel mechanism that filters which system calls a process can make. In strict mode: only read, write, exit, sigreturn. In filter mode (seccomp-BPF): a BPF program decides per-syscall. Docker applies a default seccomp profile that blocks ~44 dangerous syscalls (ptrace, kexec_load, mount, pivot_root, etc.) for all containers unless overridden. This significantly reduces the kernel attack surface even if a container is compromised.
A SUID binary runs with the file owner's UID (usually root) regardless of who executes it. Any SUID binary with a code execution path (file read/write, shell spawn, command execution) can be abused to gain root — see GTFOBins. Find with: find / -perm -4000 -type f 2>/dev/null (SUID), find / -perm -2000 -type f 2>/dev/null (SGID). Remove SUID from any binary that doesn't explicitly need it: chmod u-s /path/to/binary.
An SBOM (Software Bill of Materials) is a machine-readable inventory of all software components in an artifact — packages, versions, licenses, hashes, dependencies. Standard formats: SPDX (ISO standard, JSON/YAML) or CycloneDX (OWASP standard, XML/JSON). CI/CD integration: syft docker:myimage -o spdx-json > sbom.json during image build, then grype sbom:sbom.json --fail-on critical to gate the pipeline on CVE severity. Attach the SBOM as a build artifact and optionally sign it with cosign attest.
AIDE (Advanced Intrusion Detection Environment) is a file integrity monitor. It takes a baseline snapshot of files (hashes, permissions, timestamps, inode metadata) and stores them in a database. Periodic aide --check runs compare current state against the baseline and alerts on discrepancies. Limitations: (1) the database must be stored securely (read-only media or remote) or an attacker can update it; (2) it's reactive — it detects changes after the fact; (3) it doesn't monitor memory or network; (4) any file added after the baseline is invisible until a new baseline is taken.
Key rule categories: (1) Authentication: pam_unix, sshd, /etc/passwd//etc/shadow writes; (2) Privilege escalation: sudo/su executions, setuid/setgid syscalls; (3) File integrity: writes to /etc/, /usr/bin/, /sbin/; (4) Module loading: init_module, finit_module, delete_module syscalls; (5) Network config changes: sethostname, setdomainname; (6) Time manipulation: settimeofday, adjtimex; (7) Make rules immutable at the end: auditctl -e 2.
sysctl parameters would you set to harden a production Linux server?Critical ones: kernel.dmesg_restrict=1 (hide dmesg from unprivileged users), kernel.kptr_restrict=2 (hide kernel pointers), net.ipv4.conf.all.rp_filter=1 (reverse path filtering, prevent spoofing), net.ipv4.tcp_syncookies=1 (SYN flood protection), net.ipv4.conf.all.accept_redirects=0 (no ICMP redirects), kernel.randomize_va_space=2 (ASLR), fs.protected_hardlinks=1, fs.protected_symlinks=1, net.ipv4.ip_forward=0 (unless router), kernel.yama.ptrace_scope=1 (restrict ptrace to parent processes only).
PermitRootLogin no, PasswordAuthentication no (keys only), PubkeyAuthentication yes, AuthorizedKeysFile .ssh/authorized_keys, X11Forwarding no, AllowAgentForwarding no, MaxAuthTries 3, LoginGraceTime 30, AllowUsers <explicit user list> (or AllowGroups), Protocol 2 (SSH2 only, though this is now default), KexAlgorithms restricted to modern curves (no diffie-hellman-group1), Ciphers restricted (no arcfour, 3DES). Use sshd -T | grep -E 'permittoot|passwordauth|x11' to audit.
hidepid=2 on /proc and what attack does it mitigate?hidepid=2 (or hidepid=invisible in newer kernels) is a mount option for procfs that hides other users' /proc/<pid>/ entries — a non-root user can only see their own processes. Without it, any user can read /proc/<pid>/cmdline, /proc/<pid>/environ (may contain secrets like passwords passed as env vars or command arguments), and /proc/<pid>/fd/ (open file descriptors). Set in /etc/fstab: proc /proc proc defaults,nosuid,noexec,nodev,hidepid=2,gid=proc 0 0.
/tmp and why?/tmp is world-writable, making it a common malware staging area. Hardening: mount /tmp as a separate tmpfs with noexec (prevents executing binaries from /tmp), nosuid (prevents SUID escalation from /tmp files), nodev (no device files). In /etc/fstab: tmpfs /tmp tmpfs defaults,noexec,nosuid,nodev,size=2G 0 0. Also apply sticky bit (already default on most distros: chmod +t /tmp) to prevent users from deleting each other's files. This blocks the most common pattern of dropping a shellscript or binary to /tmp and executing it.