Security Notes
Incident Response & Forensics

Digital Forensics Deep Dive

48 min read 14 sections 14 model answers verified 2026-06

Breadth layernotes-security-core-knowledge.md


Order of Volatility

Collect the most volatile evidence first — it disappears when the system is powered off or rebooted.

1. CPU registers, cache                 (lost on context switch)
2. Routing tables, ARP cache, process list, kernel stats, RAM
3. Temporary file systems (/tmp, swap)
4. Disk — non-volatile but mutable
5. Remote logging (syslog, SIEM)
6. Physical media (tapes, CDs)
7. Archival / off-site backups

ImplicationIf an incident involves a running system and you can preserve memory, do it before pulling the plug. Pulling the plug loses RAM contents. (Exception: active ransomware encrypting live data — sometimes you pull the plug anyway.)

Memory hook

capture from fastest-fading to most permanentOrder of volatility = "grab what's evaporating first." RAM holds the crown jewels of a live intrusion — running malware, decryption keys, injected code, network state, unencrypted data — and it's gone the instant you power off. Disk survives a reboot; backups survive a fire. So the cardinal IR sin is "I rebooted the box to see if that fixed it," which incinerates the best evidence. The phrase to remember: registers → RAM → disk → backups; capture live memory before you touch power.

The whole live-acquisition decision in one picture:

flowchart TD
    A([Suspect system]) --> B{Active harm now?<br/>ransomware encrypting,<br/>data deleting}
    B -- Yes --> C[CONTAIN immediately<br/>preservation loses to safety]
    B -- No --> D{Powered on?}
    D -- Yes --> E[Capture RAM FIRST<br/>it evaporates on power-off]
    E --> F[Isolate — DON'T power off<br/>quarantine network, keep it running]
    F --> G[Image / snapshot the disk]
    D -- No --> G
    G --> H[Hash source + image<br/>verify they match]
    H --> I([Analyse the COPY offline<br/>original stays sealed])

    style E fill:#fde68a,stroke:#b45309,color:#000
    style F fill:#fde68a,stroke:#b45309,color:#000
    style I fill:#bbf7d0,stroke:#166534,color:#000
    style C fill:#fecaca,stroke:#991b1b,color:#000

Forensic Soundness on a Live System

You've spotted exactly the right problem. You cannot observe a running system without changing it — this is the digital form of Locard's exchange principle ("every contact leaves a trace"). Every command you type allocates memory, touches the page cache, updates atime/access timestamps, may write to /tmp, and leaves shell-history and process artefacts. And if the box is missing the tools you need, the naive fix — install them — is the worst possible move:

  • A package install writes to disk, overwriting unallocated space and slack — potentially destroying deleted-file evidence you haven't recovered yet.
  • It updates the package database and can upgrade shared libraries — including, as you said, the very library a rootkit trojaned. Post-install, that original malicious file may be gone or its hash changed, and you've lost the artefact.
  • It runs post-install scripts, changes timestamps across the filesystem, and may even phone out — a huge, uncontrolled footprint.

So the principle is minimise and document your footprint — never expand it. You don't install onto the evidence box; you bring your tools to it.

How it's actually done (forensically sound live response):

Don't install anything on the evidence machine.

Full stop.

Run trusted, statically-linked binaries from read-only external media

(a response USB/CD, or a mounted read-only share). Static linking matters for a second reason: a compromised host may have trojaned system libraries or core utilities (a rootkit's fake ls, netstat, libc). A dynamically-linked tool would load those libraries and lie to you. A static binary carries its own code and doesn't trust the host. Invoke it by full path, not via the host's $PATH.

Prefer acquire-then-analyse-elsewhere.

Capture RAM (LiME/WinPmem) and a disk image with minimal interaction, then do the heavy analysis on a separate forensic workstation against the image — so almost nothing runs on the original.

Image disks through a write blocker

so disk acquisition is provably read-only. (RAM is the exception — you can't write-block memory; the acquisition tool itself occupies some RAM. That's accepted; you note the tool and its footprint.)

Document every action

exact commands, tools and their hashes, timestamps, and operator. This converts unavoidable change into accounted-for change.

Will it hold up in court?Yes — courts already accept that live acquisition modifies the system; that's an inherent, well-understood property, not automatic spoliation. Admissibility doesn't require a magically untouched machine — it requires a documented, validated, repeatable methodology by trusted tools, with the changes minimised and explained (in the US this is the Daubert-style "is the method sound and accepted?" test, layered on chain-of-custody). The defence will ask "you changed the system, didn't you?" and the sound answer is: "Yes — by exactly these known actions, logged here, using validated read-only tooling; everything else is provably unchanged by hash." What gets evidence thrown out is the opposite: undocumented tampering, installing software on the original, or working on the original instead of a verified copy.

Memory hook

bring your tools, don't build them on-scene. A live system is a crime scene you're forced to walk through: you can't leave zero prints, so you wear known boots and log every step. Never apt install on the evidence box — that's pouring concrete over the floor you're dusting for prints. Trusted static binaries on read-only media, image first and analyse on a clean workstation, and document the footprint you couldn't avoid. Minimised + documented = admissible; undocumented modification = thrown out.


Prepare Before the Incident: Roles & Access That Must Already Exist

Every acquisition step in this doc quietly assumes a lot is already in place — a forensic account to receive snapshots, a quarantine security group, a break-glass admin, immutable logs. None of it can be built mid-incident. Creating an IAM role or a new account while you're breached is slow, needs the very admins who may be compromised, generates change-noise that tips off an attacker watching CloudTrail, and assumes your SSO still works — when the IdP may be the thing that's down or owned (think Golden SAML). The rule: decide and provision in calm; execute in chaos.

People — the response org (named in the IR plan, with backups)

You assign these roles, not people, in advance — and name primary + backup for each, because the on-call person changes:

Incident Commander (IC)

owns the incident and makes the calls; coordinates, not necessarily the most technical person in the room.

Scribe

keeps the timestamped log of every action and decision (feeds the post-incident report and the chain of custody).

Forensics / Investigation Lead

runs acquisition and analysis (the rest of this doc).

Comms Lead

internal + external messaging (execs, customers, regulators).

Legal / Privacy counsel

breach-notification duties, legal privilege over the investigation, law-enforcement liaison.

Service owners / SMEs

the on-call engineers who actually know the affected systems.

Executive sponsor

can authorise business-disruptive actions and own the consequences.

For a distributed company

a follow-the-sun on-call rotation and a documented escalation tree with timezones and backups, plus one agreed war-room channel/bridge as the single source of truth. Pre-sign an external IR/forensics retainer and know your cyber-insurance contact — you don't want to negotiate a contract while breached.

Access — pre-provisioned, least-privilege, ready to assume

Pre-built constructWhat it isWhy it must exist in advance
Break-glass accountA pre-created, highly-privileged account with credentials sealed in a vault/safe, hardware MFA, and a high-priority alert on any useWhen your IdP/SSO is down or compromised (Golden SAML), your normal admin path is gone — this is the only way back in. You can't mint it after the IdP is owned.
Dedicated forensic account / projectAn isolated, locked-down AWS account / GCP project for analysis, with cross-account trust to receive shared snapshotsStanding up an account + networking + key grants takes days and approvals — impossible mid-incident.
Read-only investigator roleAssume-able role: CloudTrail, Config, GuardDuty, VPC Flow Logs, Describe*/List*, read of log bucketsLets analysts investigate immediately without handing out admin (least privilege under pressure).
Forensic-acquisition roleScoped role: create/copy/share EBS snapshots, run SSM commands, isolate (modify SG, detach ASG), tag resourcesTurns "isolate + snapshot" into one command instead of a permissions scramble while the clock runs.
KMS grants pre-arrangedKey policy/grants letting the forensic account decrypt copied encrypted snapshotsThe classic failure: you share the snapshot, then discover you can't read it because the CMK was never granted cross-account. Encrypted-by-default means this will bite you if unprepared.
Quarantine constructsA deny-all "quarantine" security group, an isolation subnet/VPC, a deny-all NetworkPolicy, a forensic K8s namespace + RBAC + tools image, an isolation SCPIsolation becomes a single pre-tested action, not a design exercise at 3am.
MDM/EDR everywhere + FileVault key escrowEvery endpoint pre-enrolled in Jamf/EDR with live-response and recovery-key escrowYou cannot enrol a remote laptop after it's compromised — and you'll need that escrowed FileVault key to unlock a Mac image.
📖

What "KMS grants pre-arranged" actually means. KMS is AWS's Key Management Service — the thing that holds encryption keys. EC2 disks (EBS volumes) are almost always encrypted at rest with a KMS key, which means their snapshots are encrypted with that same key. Here's the trap that ruins forensics if you don't prepare: copying the disk is easy, but reading it requires permission to use the key — and the key lives in the production account, while you want to analyse in a separate forensic account. Share the snapshot, and the forensic account still sees encrypted bytes it can't decrypt. Worse, AWS won't even let you share a snapshot encrypted with the default aws/ebs key — only one encrypted with a customer-managed key (CMK) can be shared cross-account. So two things must be set up in advance:

Production volumes use a customer-managed KMS key

(not the default aws/ebs key), so snapshots are shareable at all.

That key's policy (or a standing KMS grant) lets the forensic account use it

specifically kms:Decrypt, kms:DescribeKey, kms:GenerateDataKey*, kms:ReEncrypt*, and kms:CreateGrant. (A "grant" is just AWS's scoped, revocable way to delegate use of a key to another principal without rewriting the whole key policy — the clean pre-arranged mechanism.)

How it's used during the incident

prod account                          forensic account
─────────────                         ────────────────
1. create-snapshot (encrypted, CMK)
2. share snapshot ───────────────────► 3. copy-snapshot  --kms-key-id <FORENSIC account's own key>
                                           (re-encrypts the copy under a key the forensic account
                                            fully controls → now self-contained, can revoke prod access)
                                        4. create volume from the copy → attach READ-ONLY → analyse

The point of pre-arranging it: editing a production KMS key policy during an incident is slow, needs approvals, risks locking out live services, and the change shows up in CloudTrail where an attacker may be watching. Do it once, in calm — then step 2→3 "just works" at 3am instead of dead-ending on AccessDenied.

Logging — you can't collect what you never recorded

CloudTrail (org-wide → an object-locked / immutable S3 bucket), VPC Flow Logs, GuardDuty, and Kubernetes audit logs must be on, centralised, and retained before the incident. These off-box, tamper-resistant logs are often your best evidence — and there is no retroactive switch. Retention has to outlast typical attacker dwell time (weeks to months).

Decision rights — pre-authorised so nobody waits at 3am

Who can declare an incident

, and the severity levels that trigger what.

Pre-approved containment authority

e.g. "the SOC may isolate any non-prod host without sign-off" — so responders act instead of hunting for an approver.

Who authorises disruptive actions

taking prod offline, forcing a global session/token revocation, paying a ransom, notifying regulators.

Legal-hold / evidence-preservation authority

, and the regulator + law-enforcement contacts and external-comms approval chain, written down in advance.

Memory hook

decide in calm, execute in chaosThe incident is the worst possible time to be creating an IAM role, standing up a forensic account, hunting the KMS key, or asking "who's even allowed to take prod offline?" Pre-build the roles (human and IAM), the forensic account, the break-glass path, the quarantine SG, the immutable logs, and the decision rights — so response is assembling a kit you already own, not improvising one while the building burns.


Memory Forensics — Volatility 3

Volatility parses raw memory dumps (.raw, .vmem, .dmp, .mem) and extracts artefacts.

Can I figure out everything from a RAM copy?

Short answer: RAM is the single richest source for what's happening right now — and useless for almost everything that happened before or that lives only on disk. A memory image is a photograph of one instant of volatile state, not a history. So you get the live runtime in extraordinary detail, but you are blind to the timeline, to dormant artefacts, and to anything that wasn't loaded into memory at the moment of capture.

What ONLY RAM gives you (capture it or lose it forever):

Secrets in the clear

disk-encryption keys (LUKS/BitLocker), TLS session keys, passwords, API tokens, and decrypted data that's ciphertext on disk. The key is in RAM only while the volume is mounted / the app is running.

Fileless / in-memory-only malware

payloads that never touch disk (PowerShell-in-memory, process-injected code). There's nothing on disk to find; RAM is the only witness.

Unpacked malware & decrypted configs

malware that's packed/encrypted on disk is unpacked in RAM, so you read its real code and C2 config.

True live state

the real process tree (including hidden/unlinked processes a rootkit removed from the OS list — psscan vs pslist), active and recently-closed network connections, injected code (malfind), loaded modules, command lines, environment variables, open handles, clipboard, and in-memory command history.

Rootkit traces

syscall-table hooks and hidden kernel modules visible only in live kernel memory.

What a RAM copy CANNOT tell you (this is the part people miss):

History / a timeline.

RAM is one instant. It can't tell you what ran and exited an hour ago, when the intrusion started, or the sequence of events. That comes from disk MAC times, logs, and a supertimeline.

Disk-resident persistence that isn't currently running.

Dormant cron jobs, scheduled tasks, run keys, systemd units, modified binaries, web shells, authorized_keys — RAM shows them only if loaded now. To find sleeping persistence you need the disk.

Full file contents and the filesystem.

Only files currently open/cached are in memory. Arbitrary files, deleted files, slack and unallocated space — disk forensics, not RAM.

Historical logs.

auth.log, event logs, web access logs live on disk (ideally shipped off-box). RAM holds at most a few recently-buffered lines.

Cold data that was never loaded.

A file the attacker never opened, or encrypted-at-rest data that was never decrypted during the capture, simply isn't in memory — and if its key was never loaded, you can't decrypt it either.

Whatever was swapped out.

Pages paged to swap / the pagefile aren't in the RAM image — memory analysis is incomplete without also grabbing swap.

Anything after the snapshot.

It's frozen at capture time; ongoing activity continues unseen.

Memory that isn't host RAM.

A host RAM image doesn't include device memory — notably GPU VRAM (see GPU-resident malware), NIC buffers, or firmware/UEFI implants.

Memory hook

RAM is the now, disk is the story, logs are the witness. RAM gives you the live intrusion and the secrets that exist nowhere else, but it's a single frozen frame with no past and no view of the disk. That's exactly why the order of volatility says "grab RAM first" and why you still image the disk and pull the logs: memory + disk + logs are complementary, and any one alone leaves a hole. One more caveat — a RAM capture of a running system isn't atomic, so structures can be slightly inconsistent ("page smear"), worst on big-memory hosts.

Acquiring a Memory Image

bash
# Linux — LiME (Loadable Memory Extractor)
sudo insmod lime-$(uname -r).ko "path=/tmp/mem.lime format=lime"

# Windows — WinPmem (open-source), Magnet RAM Capture, DumpIt
winpmem_mini_x64_rc2.exe output.raw

# VM snapshot — hypervisor creates a vmem file automatically
# Freeze VM → copy the .vmem file

Key Volatility 3 Commands

bash
# Identify OS and memory profile
python3 vol.py -f memory.lime banners.Banners
python3 vol.py -f memory.lime windows.info

# Process listing
python3 vol.py -f memory.lime windows.pslist   # process list (no hidden procs)
python3 vol.py -f memory.lime windows.pstree   # parent-child tree
python3 vol.py -f memory.lime windows.psscan   # EPROCESS scan — finds hidden/unlinked processes

# Discrepancy between pslist and psscan → suspicious (rootkit hiding processes)
# 💡 Memory hook: pslist = "walk the official guest list" (follows the OS's own
#    linked list of processes — a rootkit can UNLINK itself to hide). psscan =
#    "sweep the whole building for anyone process-shaped" (scans raw memory for
#    EPROCESS structures, finding the unlinked/hidden ones). Anything in psscan
#    but NOT pslist is a process actively hiding from the OS = red flag.
python3 vol.py -f memory.lime windows.pslist | awk '{print $2}' > pslist.txt
python3 vol.py -f memory.lime windows.psscan | awk '{print $2}' > psscan.txt
diff pslist.txt psscan.txt   # processes in psscan but not pslist = hidden

# Loaded DLLs and injected code
python3 vol.py -f memory.lime windows.dlllist --pid 1234
python3 vol.py -f memory.lime windows.malfind   # find injected PE headers in memory

# Network connections
python3 vol.py -f memory.lime windows.netstat
python3 vol.py -f memory.lime windows.netscan   # finds closed connections too

# Command history
python3 vol.py -f memory.lime windows.cmdline   # command line for each process
python3 vol.py -f memory.lime windows.consoles  # console history (cmd.exe)

# Registry hives
python3 vol.py -f memory.lime windows.registry.hivelist
python3 vol.py -f memory.lime windows.registry.printkey --key "SOFTWARE\Microsoft\Windows\CurrentVersion\Run"

# Dump a specific process
python3 vol.py -f memory.lime windows.dumpfiles --pid 1234

# Strings in process memory
python3 vol.py -f memory.lime windows.strings --pid 1234 | grep -i "password\|c2\|http"

Linux Memory Forensics

bash
python3 vol.py -f memory.lime linux.pslist
python3 vol.py -f memory.lime linux.bash           # bash history from memory
python3 vol.py -f memory.lime linux.netfilter      # netfilter hooks (rootkit indicator)
python3 vol.py -f memory.lime linux.check_syscall  # detect syscall table hooks
python3 vol.py -f memory.lime linux.lsmod          # loaded kernel modules
python3 vol.py -f memory.lime linux.hidden_modules # modules hidden from lsmod

Disk Forensics

Creating Forensic Images

bash
# dd — raw bit-for-bit copy
dd if=/dev/sda of=/mnt/evidence/disk.img bs=4M status=progress

# dcfldd — dd with hashing
dcfldd if=/dev/sda of=disk.img hash=sha256 hashlog=disk.sha256

# ewfacquire — Expert Witness Format (EWF/E01) with compression and metadata
ewfacquire -t evidence/disk -f encase6 /dev/sda

Always hash before and after imagingSHA-256 of source == SHA-256 of image = unmodified copy.

Memory hook

the hash is your "unbroken seal." Hashing the source and the image and showing they match proves the copy is bit-for-bit identical and that you didn't alter the evidence. Re-hashing later and getting the same value proves nothing changed in your custody. In court (or a serious internal investigation) this is everything: a defense lawyer's first move is "how do we know you didn't tamper with it?" — the matching hash, a write-blocker, and a custody log are the answer. Hash + write-block + log = admissible; skip any one = challengeable.

Mounting Images Read-Only

bash
# Calculate partition offset
mmls disk.img  # show partition table and start sectors

# Mount raw image read-only
mount -o ro,loop,offset=$((512*2048)) disk.img /mnt/evidence

# Mount E01 with ewfmount
ewfmount disk.E01 /mnt/ewf
mount -o ro /mnt/ewf/ewf1 /mnt/evidence

Autopsy / Sleuth Kit

bash
# Sleuth Kit command line
mmls disk.img            # partition table
fls -r -o 2048 disk.img  # list all files including deleted (recursive)
icat -o 2048 disk.img 12345 > recovered_file  # extract file by inode

# Deleted file recovery
fls -d -r -o 2048 disk.img  # -d = deleted only

Autopsy (GUI on top of Sleuth Kit): Timeline, keyword search, artifact extraction, hash lookup against NSRL (known-good file database).


Log Analysis and Timeline

Building a Super-Timeline with log2timeline / Plaso

bash
# Create timeline from disk image
log2timeline.py --parsers linux,syslog,bash_history timeline.plaso disk.img

# Windows
log2timeline.py --parsers win7,winevt timeline.plaso disk.img

# Filter and export
psort.py -z UTC -o l2tcsv -w timeline.csv timeline.plaso

# Query by time range
psort.py timeline.plaso "date > '2024-01-01' and date < '2024-01-02'"

Key Windows Log Sources

C:\Windows\System32\winevt\Logs\
  Security.evtx       — logon, object access, privilege use, account management
  System.evtx         — service start/stop, driver load, hardware events
  Application.evtx    — application errors and events
  Microsoft-Windows-PowerShell/Operational.evtx  — PS script block logging
  Microsoft-Windows-Sysmon/Operational.evtx      — Sysmon (System Monitor: a free Microsoft Sysinternals
                                                     tool that logs process creation, network connections,
                                                     and file/registry changes) — only present if installed
  Microsoft-Windows-TaskScheduler/Operational.evtx — scheduled task events

C:\Users\<user>\NTUSER.DAT               — per-user registry hive; recent files, run keys
C:\Windows\System32\config\SAM           — local user accounts and hashes
C:\Windows\Prefetch\                     — execution artefacts (first/last run, files accessed)
C:\Windows\System32\config\SYSTEM        — system hive; services, last boot time

Key Linux Log Sources

/var/log/auth.log          — SSH logins, sudo, PAM authentication
/var/log/syslog            — general system log
/var/log/kern.log          — kernel messages (module loads, hardware)
/var/log/apache2/          — web server access and error logs
/var/log/audit/audit.log   — auditd events (syscalls, file access)
~/.bash_history            — command history (can be cleared/modified)

Quick Triage

bash
# Failed SSH logins
grep "Failed password" /var/log/auth.log | awk '{print $11}' | sort | uniq -c | sort -rn | head

# Successful logins
grep "Accepted password\|Accepted publickey" /var/log/auth.log

# New users or group changes
grep "useradd\|groupadd\|usermod" /var/log/auth.log

# Web server — top IPs
awk '{print $1}' /var/log/apache2/access.log | sort | uniq -c | sort -rn | head

# Files modified in the last 24h
find /etc /usr /home -newer /tmp/reference_time -type f 2>/dev/null

Network Forensics

PCAP Analysis

bash
# Capture
tcpdump -i eth0 -w capture.pcap -s 0

# Extract HTTP
tcpdump -r capture.pcap -A 'tcp port 80' | grep -E "GET|POST|Host:|User-Agent:"

# Extract files from PCAP
tcpflow -r capture.pcap -C

# DNS queries
tcpdump -r capture.pcap -n 'udp port 53'

# Detect long DNS names (DGA indicator)
tshark -r capture.pcap -T fields -e dns.qry.name 'dns' | awk '{print length, $0}' | sort -rn | head

# Detect beaconing — periodic connections to same destination
tshark -r capture.pcap -T fields -e frame.time_epoch -e ip.dst -e tcp.dstport \
  | sort | awk 'prev && $2==prev{print $1-t,$0} {prev=$2; t=$1}'

Zeek — Network Log Analysis

bash
zeek -r capture.pcap local
# Produces:
#   conn.log    — all connections with duration, bytes
#   dns.log     — DNS queries and responses
#   http.log    — HTTP requests (URI, user-agent, response code)
#   ssl.log     — TLS connections (SNI, cert chain)
#   files.log   — transferred files with hashes

Windows Artefacts

Prefetch

C:\Windows\Prefetch\MALWARE.EXE-XXXXXXXX.pf
→ records: executable name, run count, last 8 run times, files accessed during execution
→ parse with: pecmd.exe or WinPrefetchView

Registry Run Keys (Persistence)

HKLM\SOFTWARE\Microsoft\Windows\CurrentVersion\Run
HKCU\SOFTWARE\Microsoft\Windows\CurrentVersion\Run
HKLM\SYSTEM\CurrentControlSet\Services   (services)
HKLM\SOFTWARE\Microsoft\Windows NT\CurrentVersion\Winlogon   (Userinit, Shell)

Shellbags / LNK / Jump Lists

Shellbags

(HKCU\SOFTWARE\Microsoft\Windows\Shell\BagMRU) — folder browsing history; persists even after folder deletion

LNK files

(C:\Users\<user>\AppData\Roaming\Microsoft\Windows\Recent\) — file access records; contain original path + MAC times

Jump Lists

(AutomaticDestinations\) — recent files per application


Linux Artefacts

bash
~/.bash_history         # commands typed; written on session exit
/proc/<pid>/exe         # symlink to binary; shows "(deleted)" if binary removed
/proc/<pid>/cmdline     # full command line of a running process
/proc/<pid>/maps        # memory maps — injected regions have unusual paths

# Recover a deleted running binary
cp /proc/<pid>/exe /tmp/recovered_binary

Cloud Forensics

Last verified2026-06

Cloud forensics is largely log forensics: instead of imaging a disk, you mostly query the platform's own audit trails — and the best part is they live off the box, so a compromised instance can't tamper with them. The AWS pieces used below, in plain English:

⚠️

The catch: you can only collect what you enabled before the incident. Living off the box is the upside; the downside is that most of these sources are not recorded by default, and CloudTrail captures nothing retroactively — it only logs events that happen after a trail exists. The on/off default is flagged on each item below. The headline trap is CloudTrail data events (who read which S3 object, who invoked which Lambda), which are off by default and billed extra — so mid-incident, "what did they actually exfiltrate from the bucket?" is frequently unanswerable because the events were never recorded. Cloud forensic readiness is therefore a pre-incident task (see the Logging — you can't collect what you never recorded section above); stand the pieces up with the cloudtrail-detection lab.

CloudTrail

the account's audit log of API calls: who did what, from which IP, when (e.g. "this role created a snapshot at 14:02 UTC"). Your "who did it" source. Defaults: management events are retained for 90 days in the Event history automatically (that's what lookup-events queries) — but only single-region and short-lived. A trail is what gives you multi-region, long-term, integrity-validated logs in S3, and data events (S3 object reads, Lambda invokes) are OFF by default and cost extra. Create the trail before you need it.

VPC Flow Logs

network connection metadata (source/dest IP, port, byte counts, allowed/denied) for traffic inside your VPC (Virtual Private Cloud = your isolated private network in AWS). No payload — just who talked to whom. Default: OFF — you must enable Flow Logs per VPC/subnet/ENI (to S3 or CloudWatch) ahead of time, or there's no network record to reconstruct lateral movement or exfil from.

GuardDuty

AWS's managed threat-detection service, essentially a built-in intrusion-detection system. It continuously analyses CloudTrail, VPC Flow Logs, and DNS logs and raises findings like "instance credentials used from an IP outside AWS" or "this host contacted a crypto-mining domain." You read its findings during an investigation; you don't deploy or run it yourself. Default: OFF — it must be explicitly enabled, and it only analyses activity from the moment it's turned on (it won't retro-scan last month's logs). It also reads the DNS ("which domain did the box resolve?") and Route 53 Resolver query logs are themselves off by default.

EBS volume / snapshot

an EBS volume is the virtual hard drive attached to an instance; an EBS snapshot is a point-in-time, immutable copy of one. Taking a snapshot doesn't disturb the running instance, which makes it your cloud-native forensic image + write blocker in a single step.

SSM Run Command

AWS Systems Manager (SSM) is a management agent preinstalled on most instances; Run Command lets you execute a command or script on an instance remotely through the AWS API — no SSH or RDP needed (as long as the instance has the right IAM role). It's the clean way to push a memory-capture binary to a box you can't reach directly. It does touch the instance, so you log that you used it.

Security group

an instance's virtual firewall (inbound/outbound allow rules). Swapping it for a deny-all "quarantine" group is how you isolate a host without powering it off.

IAM role

the identity an instance assumes to call AWS APIs; its temporary credentials sit in the instance's metadata/memory, so replacing or revoking the role kills any credentials an attacker stole from the box.

Auto Scaling Group (ASG)

a controller that keeps a set number of instances running and automatically replaces "unhealthy" ones — which is exactly why you must detach an instance from its ASG before isolating it, or AWS will terminate your evidence and spin up a replacement.

AWS

bash
# CloudTrail — management API calls (who did what, from where)
aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=ConsoleLogin

# VPC Flow Logs — network traffic metadata
aws logs filter-log-events --log-group-name /aws/vpc/flowlogs --filter-pattern "REJECT"

# GuardDuty findings
aws guardduty list-findings --detector-id <id>

EC2 forensic isolation flow

  1. Change security group to deny all inbound/outbound (preserve RAM state, don't terminate)
  2. Find the instance's volumes (you start with an instance ID, not a volume ID): aws ec2 describe-volumes --filters Name=attachment.instance-id,Values=i-xxx --query "Volumes[].VolumeId" --output text — snapshot all of them (root + data)
  3. Snapshot each volume: aws ec2 create-snapshot --volume-id vol-xxx (or create-snapshots --instance-specification InstanceId=i-xxx for all at once)
  4. Acquire memory via SSM Run Command → LiME → S3
  5. Analyse snapshot in a separate forensic account

GCP

bash
# Cloud Audit Logs
gcloud logging read "logName=projects/PROJECT/logs/cloudaudit.googleapis.com%2Factivity"

# IAM policy changes
gcloud logging read 'protoPayload.methodName="SetIamPolicy"'

GCP defaults mirror the same gapAdmin Activity audit logs (who changed what) are always on and free, but Data Access audit logs (who read data) are OFF by default for almost every service (BigQuery is the exception) and must be enabled per service — the direct equivalent of CloudTrail's data-events trap. VPC Flow Logs and DNS logging are likewise opt-in.


Acquiring Evidence from Modern Systems (K8s, EC2, macOS)

The textbook ("pull the disk, image with a write blocker") assumes you're standing at the machine. In a geographically distributed company you almost never are — the box is a pod on a node in another region, an EC2 instance, or a laptop two timezones away. Two rules dominate everything below:

Remote-first.

You can't seize it; you push collection to it (MDM, EDR live-response, SSM Run Command) or pull state via API (kubectl, the cloud control plane). Plan acquisition assuming zero physical access.

UTC, always.

Evidence spans regions and timezones, so the first thing you do with any artefact is normalise its timestamps to UTC and record the source's original timezone. A timeline that mixes America/Los_Angeles, Europe/Warsaw, and a container logging in UTC is worse than useless — it invents a false sequence of events. Most cloud/audit logs are already UTC; endpoint logs often aren't. Pin it down per source.

Memory hook

push, snapshot, normaliseYou push tools in (you don't install on the box — see Forensic Soundness), you snapshot for the disk (the cloud-native write blocker), and you normalise to UTC before you correlate anything across sites.

Pick your environment — the three playbooks at a glance:

flowchart TD
    Q([What am I acquiring?]) --> K{Type}
    K -->|K8s pod| P1[PRESERVE: relabel to orphan pod<br/>cordon node, deny-all NetworkPolicy]
    P1 --> P2[kubectl debug ephemeral container<br/>get tools into a distroless pod]
    P2 --> P3[Targeted proc-memory + disk at NODE<br/>node is usually EC2 → EC2 flow]

    K -->|EC2 instance| E1[ISOLATE: quarantine SG<br/>detach from ASG, strip IAM role]
    E1 --> E2[RAM: AVML pushed via SSM<br/>static binary, no SSH, no install]
    E2 --> E3[DISK: EBS snapshot<br/>attach READ-ONLY in forensic account]

    K -->|macOS laptop| M1[FileVault + T2/Silicon:<br/>no dead image, go LIVE + remote]
    M1 --> M2[Drive via MDM/EDR<br/>FileVault key from Jamf escrow]
    M2 --> M3[Logical: log collect, sysdiagnose<br/>TCC, quarantine, LaunchAgents]

    P3 --> L([Off-box logs: CloudTrail / Flow Logs /<br/>cluster audit — collect regardless of host])
    E3 --> L
    M3 --> L

    style P3 fill:#fde68a,stroke:#b45309,color:#000
    style E3 fill:#fde68a,stroke:#b45309,color:#000
    style M3 fill:#fde68a,stroke:#b45309,color:#000
    style L fill:#bbf7d0,stroke:#166534,color:#000

Kubernetes pods

Pods are ephemeral, often distroless (no shell or tools), share the node's kernel, and can be rescheduled or die any second — so speed and "don't destroy it" matter most.

bash
# 1. PRESERVE — never `kubectl delete pod`. Isolate it instead:
kubectl label pod <pod> quarantine=true app-
#   (Relabel trick. ADDS quarantine=true and REMOVES the `app` label (the trailing `-` deletes a label).
#    Why: removing the label the ReplicaSet selects on ORPHANS this pod — the controller stops managing
#    it (so it won't be auto-deleted out from under you) and launches a fresh replacement to keep the
#    service running. You get to keep the original pod frozen for analysis, and it's dropped from its
#    Service so it stops receiving live traffic.)
kubectl cordon <node>
#   (Marks the node "unschedulable" so NO NEW pods get placed on it while you do node-level work.
#    Why this and NOT `kubectl drain`: drain EVICTS/kills the running pods on the node = destroys your
#    evidence. cordon leaves everything running.)
#   + apply a deny-all NetworkPolicy selecting quarantine=true
#     (A NetworkPolicy is Kubernetes' firewall for pods. A deny-all one severs this pod's network —
#      cutting C2 and lateral movement — while leaving it RUNNING so memory and live state survive.)

# 2. CAPTURE ORCHESTRATION STATE (the spec IS evidence — it names the image, mounts, and identity)
kubectl get pod <pod> -o yaml          # full spec: image digest, env vars, volume mounts, node, serviceAccountName
kubectl describe pod <pod>; kubectl get events --field-selector involvedObject.name=<pod>   # recent scheduling/restart events
kubectl logs <pod> -c <container> --previous   # stdout/stderr; --previous = the last CRASHED/terminated container

# 3. GET TOOLS IN WITHOUT MODIFYING THE IMAGE (works even on distroless / no-shell containers)
kubectl debug -it <pod> --image=<your-forensic-image> --target=<container> --share-processes
#   (Injects a temporary "ephemeral container" into the running pod, sharing the target's process
#    namespace. You get a shell + YOUR tools sitting next to the suspect process WITHOUT altering the
#    target's image or filesystem — the modern answer to "the container has no shell or tools".
#    <your-forensic-image> = a tools container you built ahead of time and pushed to your registry.)
  • Memory — go targeted, not the whole node. A full RAM dump of a big EKS node is usually not worth it: ML/GPU worker nodes routinely have 256 GB–2 TB of RAM, so the dump is a file as large as RAM, takes a long time, hammers disk I/O, needs an equally huge evidence volume, and can push the node into memory pressure (evicting workloads). Since a container is just processes on the node, dump only the suspect container's process memory instead: find its PIDs and grab those.
    bash
    crictl inspect <container-id> | grep -i pid          # or: ps -ef | grep <process>  on the node
    sudo gcore -o /mnt/evid/proc <PID>                   # core dump of ONE process (MBs–GBs, not TBs)
    # or read regions directly: /proc/<PID>/maps + /proc/<PID>/mem ; metadata in /proc/<PID>/{cmdline,environ,exe,fd}
    Pair that with cheap live triage that captures most of the volatile value without a giant image — process list, ss -tunap (connections), lsof, loaded modules, and the container's /proc/<pid>/ artefacts. Reserve a full-node AVML/LiME dump for small nodes or when you specifically need in-memory-only secrets/fileless code and have sized the target and the evidence volume first.
  • Container filesystem: kubectl cp for triage, or node-side from the overlay (/var/lib/containerd/.../overlay2/), or a containerd/CRIU checkpoint.
  • Pull the image by digest and analyse it offline.
  • The node is usually an EC2 instance (EKS) → for the heavy lifting (RAM + disk) fall through to the EC2 steps. Fargate gives you no node access — you're limited to logs, kubectl cp, and the ephemeral debug container.
  • Pod order of volatility: ephemeral container memory → node RAM → container writable layer → node EBS → cluster audit logs → registry image. (Deeper: kubernetes/incident-response.md.)

EC2 instances

Don't terminate (destroys instance-store and, by default, the root EBS) and don't stop (loses RAM) until you've captured memory.

bash
# 1. ISOLATE WITHOUT KILLING — keep it running, just sever it
aws ec2 modify-instance-attribute --instance-id i-xxx --groups sg-quarantine
#   (Swaps the instance's SECURITY GROUP — its virtual firewall — for a deny-all/forensic-only one.
#    Result: network-isolated (C2 + lateral movement cut) but still POWERED ON, so RAM is preserved.)
aws autoscaling detach-instances --instance-ids i-xxx --auto-scaling-group-name asg --no-should-decrement-desired-capacity
#   (Removes the box from its AUTO SCALING GROUP. Do this FIRST: the ASG treats an isolated/unhealthy
#    instance as failed and would TERMINATE your evidence and launch a replacement. The
#    --no-should-decrement... flag tells it to launch that replacement, so production stays healthy
#    while your instance sits detached and untouched. Also remove it from any load-balancer target group.)
#   + cut THIS instance's IAM identity — surgically, at the instance↔profile association (NOT the role):
aws ec2 replace-iam-instance-profile-association --association-id iip-assoc-xxx \
  --iam-instance-profile Name=quarantine-deny-all       # swaps the profile on ONE instance only
#     (find the association id: describe-iam-instance-profile-associations --filters
#      Name=instance-id,Values=i-xxx. This affects ONLY this box — other instances sharing the same
#      role are untouched. Do NOT attach a Deny to the shared ROLE itself — that breaks everyone using it.)

# 2. MEMORY — capture from inside with a STATIC binary (no install), pushed via SSM (no SSH)
aws ssm send-command --instance-ids i-xxx --document-name AWS-RunShellScript \
  --parameters 'commands=["/mnt/evid/avml /mnt/evid/mem.lime"]'      # then ship the dump to S3
#   (AVML = Microsoft's single static Linux memory grabber: one self-contained binary, so it honours
#    the no-install rule. Write the dump to an attached evidence volume. Windows hosts: WinPmem or DumpIt.)
#   SIZE REALITY: a full dump is a file as big as RAM. On a 512GB+ ML/GPU box that's often infeasible —
#   prefer a TARGETED per-process core (gcore <PID>) + live triage, and reserve full RAM for small hosts.

# 3. DISK — first FIND the instance's volumes (you only have an instance ID i-xxx, not a vol-id yet)
aws ec2 describe-volumes --filters Name=attachment.instance-id,Values=i-xxx \
  --query "Volumes[].{Vol:VolumeId,Device:Attachments[0].Device,Root:Attachments[0].DeleteOnTermination}" --output table
#   (lists every EBS volume attached to that instance, its device name e.g. /dev/xvda, and whether it
#    dies on terminate. Snapshot ALL of them — root AND data volumes — not just the obvious one. You can
#    also read them off the instance: describe-instances ... BlockDeviceMappings[].Ebs.VolumeId)

# then the cloud write blocker is an EBS SNAPSHOT (point-in-time, immutable, doesn't touch the instance)
aws ec2 create-snapshot --volume-id vol-xxx --description "evidence i-xxx <case>"
#   → create a NEW volume from the snapshot, attach it READ-ONLY to a hardened forensic instance in an
#     ISOLATED VPC / dedicated forensic account. Record snapshot IDs + SHA-256 hashes for custody.
#   Snapshotting many volumes at once? `create-snapshots` (plural) --instance-specification snapshots
#   every attached volume in one consistent call.
  • The richest, immutable, off-box evidence is the logs: CloudTrail (API/role use), VPC Flow Logs, GuardDuty, CloudWatch, SSM session history. Collect these regardless of instance state. (Deeper: cloud/aws/incident-response.md.)
  • Using SSM to run your tools does modify the instance — that's fine, document it (it's the cleanest remote push). Analyse the snapshot in the forensic account, never on the victim.
📖

"Won't cutting the IAM role break every other instance?" — only if you do it wrong. Instances do share roles all the time — an Auto Scaling Group launches N identical boxes all using the same role, and that's normal. The trick is what you cut, and where:

Right (surgical)

change the instance↔instance-profile association for the one compromised instance (replace-iam-instance-profile-association). That swaps the identity on that box only; the other instances sharing the role are completely untouched. This is the isolation move.

Wrong (collateral damage)

attaching a Deny to the shared role, or running IAM's "revoke active sessions" on it, hits every instance using that role. (They'd recover by re-fetching fresh creds from IMDS, but you'd cause an outage — and you still wouldn't have surgically isolated the attacker.)

So you do not need a unique role per instance. But the incident does expose why fine-grained, per-workload identity is the goal: a shared role means stolen creds carry that role's full permissions and a wider blast radius, and attribution in CloudTrail is muddier ("which of the 40 boxes did this?"). The clean design:

EC2

scope roles per service/workload, least-privilege, not one fat role for everything.

Kubernetes

don't let pods inherit the node's instance role via IMDS (block it with a hop limit / NetworkPolicy). Give each pod its own role with IRSA (IAM Roles for Service Accounts) or EKS Pod Identity — a projected service-account token lets each workload assume a scoped role independent of the node and of other pods. Then you can revoke or isolate one workload's identity with zero collateral damage — which is exactly the problem you just spotted.

Now what? Two ways to actually analyse that snapshot. A snapshot is just an immutable point-in-time backup sitting in AWS — you can't read it directly. You either turn it back into a volume and attach it to a forensic machine, or you pull it down to a local box as a file. Both keep the original instance untouched.

Path A — attach to a forensic instance in the cloud (the default; uses your pre-created forensic role):

bash
# In the FORENSIC account, using the pre-provisioned forensic role:
# (1) make the copy self-contained — re-encrypt under the forensic account's OWN KMS key
aws ec2 copy-snapshot --source-snapshot-id snap-xxx --source-region us-east-1 \
  --kms-key-id <forensic-account-key> --encrypted --description "evidence copy"
# (2) turn the snapshot into a volume — MUST be in the same AZ as the forensic instance
aws ec2 create-volume --snapshot-id snap-copy --availability-zone us-east-1a --volume-type gp3
# (3) attach it to a running, hardened, isolated forensic EC2 instance
aws ec2 attach-volume --volume-id vol-new --instance-id i-forensic --device /dev/sdf
# (4) on the forensic instance, mount READ-ONLY — and avoid journal replay (that would WRITE to evidence)
lsblk                                                         # find the device, e.g. /dev/nvme1n1p1
sudo mount -o ro,noexec,nosuid,nodev,norecovery /dev/nvme1n1p1 /mnt/evidence
#   ro = read-only · norecovery = don't replay the filesystem journal (replay = a write!) ·
#   noexec/nosuid/nodev = never run or trust anything on the evidence disk

Then run your analysis tooling (Sleuth Kit, log2timeline, mac_apt, AIDE) against /mnt/evidence on the forensic instance — never the victim. This is the fast, cheap, isolated default: no data leaves the controlled forensic account.

Path B — take it offline and analyse on your own Linux box:

AWS won't let you "download a snapshot" directly, so you either: use the open-source coldsnap tool (AWS Labs) which reads the snapshot via the EBS direct APIs and writes a raw disk image locally — coldsnap download snap-xxx evidence.img — or attach the volume to a throwaway instance, dd it to a file, and copy that to S3 / down to your lab. Now evidence.img is an ordinary raw image you treat exactly like any other:

bash
sha256sum evidence.img                                       # custody hash
sudo losetup -r -f --show evidence.img                       # -r = read-only loop device → /dev/loop0
sudo mount -o ro,norecovery /dev/loop0p1 /mnt/evidence       # or use Sleuth Kit (mmls/fls/icat) without mounting

Do this when policy or chain-of-custody requires evidence in a physical lab, or you want it fully off-cloud. The trade-off is egress: large volumes are slow and costly to pull down, so in-cloud (Path A) is the usual choice and offline (Path B) is for when you specifically need the image in hand.

Memory hook

Either way, the rules are identical to a physical disk. Work on a copy (the snapshot is your immutable master), mount it read-only with norecovery so you never replay a journal into the evidence, analyse on a separate forensic machine, and hash + log everything. The snapshot is the cloud's write-blocked master image; a volume-from-snapshot or a coldsnap image is your working copy.

macOS (remote laptops)

Modern Macs make the classic dead-disk image effectively impossible: FileVault encrypts the whole disk, and the T2 / Apple Silicon secure storage means the internal SSD is encrypted and soldered with a sealed System Volume. So macOS DFIR is live, logical, and remote.

  • You won't seize the laptop → drive collection through MDM (Jamf) or EDR live-response: push a collector, pull the results. Jamf also escrows the FileVault recovery key — get it from there if you ever need to unlock an image.
  • Memory: full RAM capture is unreliable on recent macOS (SIP, Apple Silicon broke osxpmem); realistically you do live triage, not a RAM dump.
  • Logical collection (run a static collector, minimal footprint):
bash
sudo log collect --output /evidence/unified.logarchive   # Unified Logs — the central modern macOS log store
#   (`log collect` is BUILT IN to macOS — no tool to source. It packages the unified logs into a
#    .logarchive you copy off and read later with `log show`.)
sudo sysdiagnose -f /evidence/                           # also BUILT IN — broad system-state snapshot
# Optional collectors you pre-stage (see sourcing box below): UAC or AutoMacTC to bundle artefacts,
# osquery for live state; parse the collected data offline with mac_apt.
Key artefacts

Unified Logs; TCC.db (which apps were granted camera/mic/files/Full-Disk-Access — superb for spotting malware); quarantine (LSQuarantineEventsV2 + the com.apple.quarantine xattr — download provenance); persistence: LaunchAgents/LaunchDaemons, login items, configuration profiles; plus knowledgeC.db, FSEvents, Spotlight, browser history.

If you must disk-image

do it while the Mac is unlocked/live via Apple Silicon "Share Disk" (or Target Disk Mode on Intel) to another Mac — but it needs the FileVault password/recovery key, and imaging the live volume modifies it, so document it.

🧰

Where to source these tools — and pre-stage them. Download or build every one of these ahead of time onto your forensic kit (a response USB, a "forensic AMI", a tools container image, your analysis workstation). Never fetch them onto — or from — the victim machine during an incident: that's both an install-on-evidence violation (Forensic Soundness) and an OPSEC tell. All are free / open-source:

  • AVML (Linux memory, single static binary — easiest) — github.com/microsoft/avml, grab the prebuilt release binary.
  • LiME (Linux memory, kernel module) — github.com/504ensicsLabs/LiME. Catch: it must be compiled against the target's exact kernel version, which is why AVML's prebuilt static binary usually wins under pressure.
  • WinPmem (Windows memory) — github.com/Velocidex/WinPmem; DumpIt — Magnet Forensics.
  • Volatility 3 (memory analysis, runs on your workstation) — github.com/volatilityfoundation/volatility3.
  • UAC (Unix/Linux/macOS/ESXi triage collector) — github.com/tclahr/uac; AutoMacTC (macOS) — github.com/CrowdStrike/automactc; mac_apt (macOS artefact parser) — github.com/ydkhatri/mac_apt; osquery (live state, all OSes) — osquery.io.
  • osxpmem (macOS memory, often broken on recent versions) — part of github.com/Velocidex/c-aff4.
  • Don't want a shopping list? The SANS SIFT Workstation (free VM) ships most Linux/analysis tooling pre-installed — a ready-made workstation for the "analyse the snapshot elsewhere" step.

Malware Triage

Static Analysis

bash
# File identification
file malware.exe
strings malware.exe | grep -E "http|cmd|powershell|regsvr32"

# PE entropy — high entropy section → packed/encrypted
python3 -c "
import pefile, math
pe = pefile.PE('malware.exe')
for s in pe.sections:
    data = s.get_data()
    freq = [data.count(bytes([i]))/len(data) for i in range(256)]
    entropy = -sum(f * math.log2(f) for f in freq if f > 0)
    print(s.Name.decode().strip(), round(entropy,2))
"

# Hash and VirusTotal lookup
sha256sum malware.exe
curl -X POST https://www.virustotal.com/api/v3/files -H "x-apikey: KEY" -F "file=@malware.exe"

# YARA
yara -r rules/ malware.exe

Dynamic Analysis (Sandboxing)

Cuckoo Sandbox

open-source, self-hosted

ANY.RUN

interactive online sandbox

Hybrid Analysis

free tier

Key behaviours to observe:

  • Process creation (especially cmd.exe, powershell.exe, wscript.exe, regsvr32.exe)
  • Network connections (C2 IPs/domains, DNS lookups for DGA-style names)
  • Registry persistence keys written
  • Process injection (CreateRemoteThread, VirtualAllocEx, WriteProcessMemory)

Chain of Custody

Forensic evidence must be handled correctly to be admissible:

Hash before and after

SHA-256 of source == SHA-256 of image

Write-protect originals

hardware write blockers for physical disks

Log every touch

who, when, what action, with what tool

Work on copies

all analysis on verified copies; originals sealed

Tamper-evident packaging

for physical media

The integrity chain — where a single break makes it "challengeable":

flowchart LR
    A([Original evidence]) --> B[Hash source<br/>SHA-256]
    B --> C[Write-block / snapshot<br/>read-only acquisition]
    C --> D[Create image]
    D --> E{Hash image matches<br/>hash source?}
    E -- No --> X[STOP — acquisition<br/>is not faithful]
    E -- Yes --> F[Seal original<br/>analyse the COPY only]
    F --> G[Log every handoff<br/>who · when · what · tool]
    G --> H([Admissible:<br/>hash + write-block + custody log])

    style E fill:#fde68a,stroke:#b45309,color:#000
    style X fill:#fecaca,stroke:#991b1b,color:#000
    style H fill:#bbf7d0,stroke:#166534,color:#000
Memory hook

The seal analogyThe matching hash is a tamper-evident seal: break any link — no source hash, working on the original, an unlogged handoff — and the defence's "how do we know you didn't alter it?" has no answer. Hash + write-block + custody log = admissible; miss one = challengeable.


Interview Questions

Q
What is the order of volatility and why does it matter?
Model answer

It's the principle of collecting evidence from most ephemeral to most permanent, because the volatile stuff disappears when you lose power or reboot. The order is roughly CPU registers and cache, then RAM and live state like the process list and network connections, then temporary filesystems and swap, then disk, then remote logs, then backups. It matters because RAM holds the crown jewels of a live intrusion — running malware, encryption keys, injected code, unencrypted data, C2 connections — and it's gone instantly on shutdown. So if a system is live and I can safely capture memory, I do that before touching power. The classic mistake is rebooting a suspected-compromised box, which destroys the best evidence. The exception is active harm like ransomware encrypting data, where containment may trump preservation.

Q
Walk me through how you'd respond to a potentially compromised Linux server.
Model answer

First, don't reboot or "clean up" — preserve evidence in order of volatility. If it's live and safe, I capture memory with something like LiME, then snapshot the disk. I isolate it at the network layer rather than powering off, so I keep RAM and don't tip a dead-man's switch. Then I investigate: in memory, the process tree, network connections, and bash history; on disk, auth.log for who logged in and escalated, auditd for what they did, /proc for running processes (including deleted binaries via /proc/pid/exe), cron/systemd/authorized_keys for persistence, and recently modified files. I build a timeline, identify the entry vector and blast radius, and check whether credentials or the cloud role were stolen. Containment and eradication follow once I've scoped it, and I recover from a known-good image rather than trusting the box. Throughout I hash evidence and keep a custody log.

Q
What is Volatility and what can you extract from a memory image?
Model answer

Volatility is the standard open-source memory-forensics framework — it parses a raw RAM dump and reconstructs OS structures. From memory you get things you can't reliably get from disk: the live process list and parent-child tree, hidden or unlinked processes via psscan, network connections including closed ones, command lines and console history, injected code via malfind, loaded DLLs/kernel modules, registry keys cached in memory, and often credentials, encryption keys, or decrypted malware configs. On Linux it can detect syscall-table hooks and hidden kernel modules — rootkit indicators. It's the tool of choice because memory captures the true runtime state, including things malware tried to hide from the OS.

Q
How do you create a forensically sound disk image and why hash it?
Model answer

I make a bit-for-bit copy using a write blocker on the original so the acquisition can't modify it, with a tool like dcfldd or ewfacquire that images and hashes in one pass. Before and after, I compute a cryptographic hash — SHA-256 — of both the source and the image. Matching hashes prove the copy is identical to the original and that I didn't alter the evidence; re-hashing later proves nothing changed in my custody. Then all analysis happens on verified copies, never the original, which stays sealed. The hashing is what makes the evidence defensible — it's the cryptographic seal that answers "how do we know this wasn't tampered with."

Q
What are Windows Prefetch files and what do they tell you?
Model answer

Prefetch files are a Windows performance feature that records when programs run so they load faster next time, and they're a goldmine for forensics. Each .pf file tells you an executable existed and ran, how many times, the last several run times, and which files it accessed during execution. For an investigation that's evidence of execution — proof that a given malware or tool actually ran on the host and when — which is often the question that matters. A missing Prefetch entry for a known-installed program, or one for a binary that's since been deleted, is itself notable. The caveat is Prefetch can be disabled, especially on SSDs/servers.

Q
What is log2timeline/Plaso and what does a supertimeline answer?
Model answer

Plaso, driven by log2timeline, ingests an entire disk image and extracts timestamps from every artifact it understands — filesystem MAC times, event logs, registry, browser history, Prefetch, bash history — and merges them into one unified, sortable "supertimeline." It answers the core investigative question: what happened, in what order, around the time of the incident. Instead of pivoting between dozens of artifact types, you get a single chronological view that reveals the sequence — initial access, then execution, then persistence, then lateral movement. The trade-off is volume: a supertimeline is huge, so you filter to the relevant window and known-bad indicators to make it usable.

Q
How would you detect beaconing in a PCAP?
Model answer

Beaconing is a C2 implant phoning home on a schedule, so the tell is regularity even when the payload is encrypted. I'd group connections by source and destination, compute the time intervals between successive connections to the same destination, and look for low variance — near-constant intervals — over many connections, allowing for jitter. Consistent small payload sizes and long-lived low-volume flows strengthen it. Then I exclude legitimate periodic traffic like NTP, software updates, and telemetry, and investigate what's left. Tools like Zeek's conn.log or RITA automate this beacon-scoring. It's a behavioral detection — it catches the beaconing pattern regardless of the specific C2 framework.

Q
What AWS services provide forensic visibility, and what does each log?
Model answer

CloudTrail is the primary one — it records management API calls, the who-did-what-from-where audit trail, and with data events enabled it also logs S3 object reads and Lambda invocations, though those are off by default and a common evidence gap. VPC Flow Logs capture network connection metadata for traffic analysis. GuardDuty provides threat detections derived from those logs plus DNS. For deeper config history, AWS Config records resource state over time, and CloudTrail Lake or Athena lets you query it all. The big caveat I'd raise is that forensic readiness in AWS is a preparation problem — if CloudTrail data events and Flow Logs weren't enabled before the incident, that evidence simply doesn't exist after the fact.

Q
What's the difference between pslist and psscan in Volatility?
Model answer

pslist walks the operating system's own doubly-linked list of active processes — it shows what the OS will admit is running. psscan instead scans raw memory for process structures (EPROCESS) directly, regardless of whether they're in that linked list. The difference matters for rootkit detection: a process can hide by unlinking itself from the OS list (DKOM — direct kernel object manipulation), so it vanishes from pslist but still exists in memory and shows up in psscan. So anything that appears in psscan but not pslist is a process actively hiding from the OS — a strong indicator of a rootkit or stealthy malware. psscan can also recover terminated processes whose structures haven't been overwritten yet.

Q
Explain chain of custody and why it matters.
Model answer

Chain of custody is the documented, unbroken record of who handled a piece of evidence, when, and what they did with it, from acquisition through analysis to storage. In practice that means hashing evidence at acquisition, using write blockers on originals, logging every access, working only on verified copies while the original stays sealed, and tamper-evident handling for physical media. It matters because evidence is only useful if it's trustworthy: in legal proceedings, or any serious investigation, the integrity of the evidence will be challenged, and a clean custody chain plus matching hashes is what proves it wasn't altered or planted. A gap in custody — an unlogged handoff, a missing hash, working on the original — is exactly what gets evidence thrown out.

Q
If the evidence machine doesn't have the tools you need, why not just install them — and how do you avoid contaminating it?
Model answer

You never install onto the evidence box, because the install itself is contamination: it writes to disk and overwrites unallocated space, possibly destroying deleted-file evidence; it updates the package database and can upgrade shared libraries — including the very library a rootkit trojaned, so you'd lose the artefact; and it changes timestamps everywhere. The sound approach is to bring your tools to the system, not build them on it: run trusted, statically-linked binaries from read-only external media, invoked by full path. Static linking matters twice over, because a compromised host may have trojaned core utilities or libc, and a dynamically-linked tool would load those and lie to you. Better still, acquire memory and a disk image with minimal interaction and do the analysis on a separate forensic workstation. You can't touch a live system without changing it — that's Locard's principle — so the standard isn't zero footprint, it's minimised and fully documented footprint. That's what holds up in court: a defensible methodology where every change is accounted for, versus undocumented modification, which is what gets evidence excluded.

Q
How do you acquire evidence from a Kubernetes pod, an EC2 instance, and a macOS laptop when the team is distributed and you have no physical access?
Model answer

Everything is remote, so I push collection to the asset or pull state via API, and I normalise every timestamp to UTC up front because the fleet spans timezones and a mixed-timezone timeline invents a false sequence. For a pod, I preserve rather than delete — isolate it with a deny-all NetworkPolicy and drop its Service labels, cordon the node but never drain it, capture the spec and logs with kubectl, and get tools in via an ephemeral debug container sharing the target's namespaces, which works even on distroless images; the container's memory I grab from the node, and since the node is usually an EC2 instance I image it there. For EC2, I isolate without killing — quarantine security group, detach from the auto-scaling group so it isn't replaced, neuter the IAM role — then capture RAM from inside with a static binary like AVML pushed over SSM, and for disk I take an EBS snapshot, which is the cloud write blocker, and attach it read-only to a forensic instance in an isolated account. For macOS, FileVault plus the T2/Apple Silicon sealed storage make dead imaging impractical, so it's live logical collection through MDM or EDR — Unified Logs via log collect, sysdiagnose, TCC and quarantine and persistence artefacts — and I pull the FileVault recovery key from Jamf escrow if I ever need to unlock an image. Throughout, the immutable off-box logs — CloudTrail, VPC Flow Logs, cluster audit logs — are often the best evidence and I collect them regardless of the host's state.

Q
What needs to be in place before an incident so you're not creating roles and access on the fly?
Model answer

The principle is decide and provision in calm, execute in chaos — because creating an IAM role or an account mid-incident is slow, needs the very admins who may be compromised, makes change-noise that tips off an attacker watching CloudTrail, and assumes your SSO still works when the IdP may be exactly what's down or owned. On the people side I want the response roles assigned in advance with primary and backup names — incident commander, scribe, forensics lead, comms, legal, and the service owners who know the systems — plus a follow-the-sun on-call rotation and escalation tree for a distributed company, and a pre-signed external IR retainer. On the access side: a break-glass account with sealed credentials and alerting for when SSO is gone; a dedicated, isolated forensic account that already has cross-account trust to receive snapshots and the KMS grants to decrypt them, because the classic failure is sharing an encrypted snapshot you then can't read; pre-built least-privilege roles, a read-only investigator role and a forensic-acquisition role; and quarantine constructs like a deny-all security group and isolation network policy ready to apply in one command. Logging — CloudTrail to an immutable bucket, flow logs, GuardDuty, Kubernetes audit — has to be on and retained beforehand, since there's no retroactive switch. And the decision rights need pre-authorising: who declares an incident, who can isolate without sign-off, and who approves disruptive actions like taking prod offline, so nobody is hunting for an approver at 3am.

Q
Can you figure out everything from a RAM capture? What can't it tell you?
Model answer

No — RAM is the richest source for what's happening right now, but it's a single frozen instant of volatile state, not a history. It uniquely gives me the things that exist nowhere else: encryption keys and decrypted data, fileless or injected malware that never touches disk, unpacked malware and its live C2 config, the true process tree including rootkit-hidden processes, and active network connections. What it can't give me is the timeline — what ran and exited before the capture — which comes from disk MAC times and logs; dormant on-disk persistence like cron jobs, run keys, or web shells that aren't currently running; full file contents and deleted or unallocated disk data; historical logs; cold data that was never loaded into memory, including encrypted data whose key was never present; and anything paged out to swap, which isn't in the image unless I grab swap too. It also doesn't cover device memory like GPU VRAM. That's the whole reason for order of volatility: grab RAM first because it evaporates, but still image the disk and pull the logs, because memory, disk, and logs are complementary and any one alone leaves a hole. On the people side I want the response roles assigned in advance with primary and backup names — incident commander, scribe, forensics lead, comms, legal, and the service owners who know the systems — plus a follow-the-sun on-call rotation and escalation tree for a distributed company, and a pre-signed external IR retainer. On the access side: a break-glass account with sealed credentials and alerting for when SSO is gone; a dedicated, isolated forensic account that already has cross-account trust to receive snapshots and the KMS grants to decrypt them, because the classic failure is sharing an encrypted snapshot you then can't read; pre-built least-privilege roles, a read-only investigator role and a forensic-acquisition role; and quarantine constructs like a deny-all security group and isolation network policy ready to apply in one command. Logging — CloudTrail to an immutable bucket, flow logs, GuardDuty, Kubernetes audit — has to be on and retained beforehand, since there's no retroactive switch. And the decision rights need pre-authorising: who declares an incident, who can isolate without sign-off, and who approves disruptive actions like taking prod offline, so nobody is hunting for an approver at 3am.