Security Notes
Malware & Exploitation

Malware Reverse Engineering — Linux Basics and Pattern Extraction

This file covers the practical workflow for analysing suspicious Linux binaries: triage, static analysis, pattern extraction, dynamic analysis, and decompilation. Everything here is executable from a Linux analysis VM with no special licences.

55 min read 15 sections 15 model answers

Breadth layernotes-security-core-knowledge.md Also see: EDR Evasion · Anti-Debugging · Digital Forensics


Why We Do This — The Analyst's Mindset

When you find a suspicious binary on a Linux system, you have one primary job: answer the five questions before the malware does more damage.

1. What IS this?        → file type, architecture, linked libraries, packed or not?
2. What does it WANT?   → network connections, files it reads/writes, commands it runs?
3. How does it PERSIST? → cron, systemd, rc.local, LD_PRELOAD, kernel module?
4. What did it STEAL?   → credentials, SSH keys, environment variables, clipboard?
5. Where does it PHONE HOME? → C2 IPs, domains, beaconing interval, protocol?

Each phase of analysis is designed to answer a different subset of those questions:

PhaseMethodAnswers
Triagefile, xxd, entropy, stringsWhat is it? Is it worth deeper analysis?
Staticreadelf, nm, objdump, checksecArchitecture, capabilities, anti-analysis features
Pattern extractionstrings + grep, regexIOCs — C2 addresses, persistence paths, credentials
Dynamicstrace, ltraceWhat it actually does when running (bypasses obfuscation)
InstrumentationGDB, FridaHow it works internally; extract decrypted data
DecompilationGhidra, Binary NinjaUnderstand complex logic; find kill switches or config

Static analysis vs dynamic analysis — why both?

Static

(analysing without running): safe, no risk of infection or C2 contact, works even on encrypted binaries (you can see the structure). Limitation: if strings are XOR-encrypted, strings finds nothing.

Dynamic

(running the binary and watching): reveals exactly what happens, bypasses all obfuscation — encrypted strings are decrypted before use, so you see them. Limitation: requires isolation (never run malware on a non-isolated machine), and the malware may detect your analysis environment and not exhibit its real behaviour.

The workflow is not linearStatic findings inform what to look for in dynamic analysis. Dynamic findings reveal which functions to decompile. You loop between phases as you learn more.

Memory hook

static is reading the recipe, dynamic is tasting the dish. Static analysis reads the binary without running it — safe, no infection risk, works on the structure even when you can't read the contents — but it's defeated by obfuscation (XOR-encrypted strings show nothing to strings). Dynamic analysis runs it and watches — the malware must decrypt its strings and reveal its C2 before it can use them, so obfuscation melts away — but it requires isolation (never run malware on a connected box) and the malware may detect the lab and play dead (hence anti-debugging). They're complementary: static tells you where to look, dynamic shows you what actually happens. Mnemonic: static = safe but fooled by encryption; dynamic = truthful but needs a cage. And tie it back to the five questions — every command you run is in service of what is it / what does it want / how does it persist / what did it steal / where does it phone home, which is also the shape of the IOCs you hand to the detection team.

One rule above all elsenever run an untrusted binary on a machine connected to your production network or any network the C2 can reach. Use an isolated VM with an internal-only network, or a dedicated sandbox. (See also: Analyst OPSEC Failures — DNS queries from your sandbox can alert the attacker.)


Triage — First 60 Seconds

GoalIn 60 seconds, produce three outputs: (1) a hash for threat intel lookups, (2) a verdict on whether the binary is packed/encrypted, (3) enough string snippets to decide what kind of malware this might be. This determines whether you go straight to static analysis or need to unpack first.

Decision tree after triage

file → "ELF" + low entropy (<6.5) → proceed to static analysis
file → "ELF" + high entropy (>7.0) → likely packed → find OEP, dump, then static analysis
file → "data" (unknown)            → binary blob, shellcode, or encrypted payload
file → "ASCII text"                → script — read it directly
file → anything containing "UPX"  → UPX packed → try `upx -d` first

Before running anything, answer: what is this file? Is it packed? Is it network-capable? Only then decide how to proceed.

bash
# 1. Identify the file type
file suspicious_binary
# → ELF 64-bit LSB executable, x86-64, dynamically linked, stripped
# → ELF 64-bit LSB executable, x86-64, statically linked, stripped
# → data  ← unknown format, possibly packed/encrypted
# → ASCII text  ← shell script, check shebang

# 2. Hash it for VirusTotal lookup
sha256sum suspicious_binary
md5sum suspicious_binary     # legacy, but VT still accepts it
# Paste the hash at virustotal.com or use the API:
curl -s --request GET \
  "https://www.virustotal.com/api/v3/files/<sha256>" \
  --header "x-apikey: $VT_API_KEY" | jq '.data.attributes.last_analysis_stats'

# 3. Check for recognizable magic bytes
xxd suspicious_binary | head -4
# 7f 45 4c 46  → ELF
# 4d 5a        → PE (Windows binary — unusual on Linux, check for Wine targets)
# 23 21        → shebang (#!)
# 1f 8b        → gzip
# 50 4b        → ZIP / JAR / APK

# 4. Entropy check (packed/encrypted = high entropy)
python3 -c "
import sys, math, collections
data = open(sys.argv[1], 'rb').read()
freq = collections.Counter(data)
entropy = -sum((c/len(data)) * math.log2(c/len(data)) for c in freq.values())
print(f'Overall entropy: {entropy:.4f} / 8.0')
print('Likely packed/encrypted' if entropy > 7.0 else 'Probably not packed')
" suspicious_binary

# 5. Quick string preview
strings -n 8 suspicious_binary | head -40

Static Analysis Tools Reference

Goal of static analysisAnswer as many questions as possible without running the binary. You learn what the binary is capable of (imported functions), what hard-coded IOCs it contains (strings), how it was compiled (stripped/not, mitigations, architecture), and whether it's packed. All of this without ever giving the malware a chance to phone home.

Order of operations in static analysis

  1. file → type and basic properties
  2. readelf -h → architecture, entry point, type (EXEC vs DYN)
  3. readelf -S → sections, find suspicious flags or names
  4. readelf -d → imported libraries
  5. nm -D → imported function names (capabilities)
  6. strings → IOCs, config, strings
  7. checksec → security mitigations summary
  8. binwalk -E → entropy, detect packing

file — Type Identification

bash
file binary
# Key fields:
#   ELF 64-bit / 32-bit      → architecture
#   LSB / MSB                 → little-endian / big-endian
#   executable / shared obj / relocatable → file role
#   dynamically linked        → will load .so files at runtime
#   statically linked         → everything compiled in — no external .so dependencies
#   (uses shared libs)        → older phrasing for dynamically linked
#   stripped                  → symbol table removed — function names gone
#   not stripped              → debug symbols present — function names visible in gdb/objdump

strings — Extract Printable Strings

strings scans for sequences of printable characters. It is the fastest way to extract IOCs, error messages, hard-coded configuration, and obfuscated fragments.

bash
# Default: ASCII, minimum 4 chars
strings binary

# Increase minimum length to reduce noise (8 chars filters most junk)
strings -n 8 binary

# Include wide (UTF-16LE) strings — important for Windows malware cross-compiled for Linux
# or malware targeting Wine environments
strings -el binary            # little-endian UTF-16
strings -eb binary            # big-endian UTF-16

# Scan the entire file (not just sections — catches data in overlay/appended data)
strings -a binary

# Show file offset for each string (useful for Ghidra cross-reference)
strings -a -t x binary        # hex offset
strings -a -t d binary        # decimal offset

# Filter for specific patterns immediately
strings -n 8 binary | grep -iE 'https?://'
strings -n 8 binary | grep -iE '[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}'
strings -n 8 binary | grep -iE 'password|passwd|secret|token|apikey|api_key'
strings -n 8 binary | grep -iE '\.(sh|py|pl|php|rb)$'
strings -n 8 binary | grep -iE '/etc/|/tmp/|/var/|/proc/'
strings -n 8 binary | grep -iE 'execve|system|popen|fork|wget|curl'

Pitfallstrings only finds printable strings. Encrypted or XOR-obfuscated strings produce no output. Use dynamic analysis or entropy analysis to detect them.


readelf — ELF Header and Sections

bash
# ELF file header: architecture, entry point, class, endianness
readelf -h binary

# Section headers: name, type, size, offset, flags
readelf -S binary
# Important sections:
#   .text    → executable code
#   .data    → initialized global variables
#   .rodata  → read-only data (hard-coded strings, constants)
#   .bss     → uninitialized globals (zero at startup)
#   .plt     → Procedure Linkage Table (external function stubs)
#   .got     → Global Offset Table (resolved addresses of external functions)
#   .dynamic → dynamic linking metadata
#   .symtab  → symbol table (may be absent in stripped binaries)
#   .debug_* → debug information (DWARF)

# Program headers (segments): how OS maps the binary into memory
readelf -l binary

# Dynamic section: shared library dependencies, RPATH, symbol versions
readelf -d binary
# Look for NEEDED entries (imported .so files) and RPATH (unusual rpath = suspicious)

# Symbol table: function names (stripped = only external/dynamic symbols)
readelf -s binary
readelf --syms binary

# Relocation entries: which external functions are called and from where
readelf -r binary

# All headers in one shot
readelf -a binary | less

# Check for SUID bit and capabilities in binary metadata
readelf -n binary   # notes section (may contain OS ABI, build ID)

objdump — Disassembly

bash
# Disassemble the .text section (code)
objdump -d binary

# Disassemble all sections (catches code in unusual sections)
objdump -D binary

# Use Intel syntax (more readable than AT&T default)
objdump -d -M intel binary

# Show source line information (only if not stripped)
objdump -d -S binary

# Dump a specific section as hex
objdump -s -j .rodata binary

# Show the full symbol table
objdump -t binary

# Show dynamic symbol table (imported functions)
objdump -T binary

# Find all CALL instructions (who calls what)
objdump -d binary | grep 'call\|jmp' | head -50

nm — Symbol Table

bash
# List all symbols
nm binary

# Show dynamic symbols only (what's imported from .so files)
nm -D binary

# Sort by address
nm -n binary

# Demangle C++ names
nm --demangle binary

# Check if stripped (no output or minimal output means stripped)
nm binary 2>&1 | head -5
# "no symbols" → stripped binary — function names unavailable in gdb without debug info

# Example output interpretation:
# 0000000000401234 T main        → T = .text section, defined here
# 0000000000000000 U puts@@GLIBC_2.2.5  → U = undefined (imported from libc)
# 0000000000403000 D some_global → D = .data section (initialized global)
# 0000000000000000 w __gmon_start__ → w = weak symbol

ldd — Shared Library Dependencies

bash
# Show all shared libraries the binary will load at runtime
ldd binary

# Example output:
#   linux-vdso.so.1 (0x00007ffd3a1fc000)
#   libpthread.so.0 → /lib/x86_64-linux-gnu/libpthread.so.0
#   libc.so.6 → /lib/x86_64-linux-gnu/libc.so.6

# What the library list tells you:
#   libssl.so / libcrypto.so  → TLS/crypto capability
#   libcurl.so               → HTTP client capability
#   libpthread.so            → multithreading (evasion, process injection)
#   libdl.so                 → dlopen() — dynamic library loading at runtime
#   libpam.so                → PAM auth hooks (credential theft)
#   no interesting libs      → statically linked or custom resolver

# NEVER run ldd on an untrusted binary on a real system.
# ldd works by setting LD_TRACE_LOADED_OBJECTS=1 and executing the binary.
# Use readelf -d instead for untrusted samples:
readelf -d binary | grep NEEDED

checksec — Security Mitigation Audit

bash
# Install: apt install checksec   or   pip install checksec
checksec --file binary

# Or using pwntools:
pwn checksec binary

# Example output:
# RELRO:    Full RELRO      ← GOT is read-only after startup
# Stack:    Canary found    ← stack smashing protector enabled
# NX:       NX enabled      ← non-executable stack
# PIE:      PIE enabled     ← position-independent executable (ASLR applies)
# RPATH:    No RPATH        ← no custom library search path
# Symbols:  No Symbols      ← stripped

# What to look for in malware:
# No PIE = fixed load address → easier to exploit or ROP
# No Canary = easier buffer overflow exploitation
# RPATH present = suspicious custom library search path (sideloading)

binwalk — Embedded Files and Entropy

bash
# Install: apt install binwalk
binwalk binary

# Detect embedded files (inside packers, droppers)
binwalk binary
# Outputs: offset, description
# 0x1234    gzip compressed data
# 0x5678    ELF 64-bit executable
# 0xABCD    PNG image

# Extract embedded files
binwalk -e binary
# Creates a _binary.extracted/ directory

# Entropy analysis — visualize packed regions
binwalk -E binary
# High-entropy regions (near 8.0) = encrypted/compressed
# Transitions from low to high entropy = packer header/stub → payload boundary

# Recursive extraction (handles nested archives)
binwalk -e -M binary

ELF Structure Analysis

GoalUnderstand the binary's layout before touching any code. The ELF format is a container — it tells you where the code is, what libraries it needs, what security mitigations are in place, and (if not stripped) what all the functions are named. Reading it correctly saves hours of disassembly.

The two views of an ELF binary:

ELF has two overlapping but different views of the same file. This confuses many analysts at first:

Section view

(linker's view): how the binary is divided logically during compilation — .text for code, .data for variables, .rodata for constants, etc. Used by the linker and by analysis tools like objdump and readelf -S. Not needed at runtime.

Segment view

(loader's view): how the OS kernel maps the file into memory when the program runs — LOAD segments define which byte ranges get mapped and with what permissions. This is what the kernel actually reads. Each segment typically spans several sections.

Think of it like this: sections are the architect's floor plan; segments are the walls the builder actually constructs.

ELF Binary Layout (on disk):
┌─────────────────────────┐  offset 0x00
│  ELF Header (64 bytes)  │  magic, type, arch, entry point address, offsets to tables
├─────────────────────────┤  offset 0x40 (64 bytes in)
│  Program Headers        │  segment table — read by kernel loader
│  (Segment table)        │  LOAD, DYNAMIC, INTERP, GNU_STACK, etc.
├─────────────────────────┤
│  .interp                │  path to dynamic linker: /lib64/ld-linux-x86-64.so.2
│  .note.gnu.build-id     │  build ID (unique hash of the binary — useful for correlation)
│  .gnu.hash / .hash      │  symbol hash tables for fast dynamic lookup
│  .dynsym                │  dynamic symbol table (imports — NEVER stripped)
│  .dynstr                │  string table for dynamic symbol names
│  .gnu.version           │  symbol version info (which glibc version needed)
│  .rela.dyn / .rela.plt  │  relocation entries (where to patch addresses at load time)
│  .plt                   │  Procedure Linkage Table — stubs for external function calls
│  .plt.got               │  PLT stubs using GOT
│  .text                  │  ← THE CODE. This is where functions live.
│  .rodata                │  read-only data: string literals, constant tables
│  .eh_frame              │  exception handling / stack unwinding data
│  .data                  │  initialized global/static variables (read-write)
│  .bss                   │  uninitialized globals (zero-filled by OS at load time, no disk space)
│  .got                   │  Global Offset Table — resolved addresses of external functions
│  .got.plt               │  GOT entries for PLT lazy binding
│  .dynamic               │  dynamic linking metadata — library names, flags, addresses
│  .symtab                │  FULL symbol table (stripped in release builds)
│  .strtab                │  string table for .symtab names
│  .shstrtab              │  string table for section names themselves
│  .debug_info            │  DWARF debug info (present only in debug builds)
├─────────────────────────┤
│  Section Header Table   │  at end of file — offset stored in ELF header
└─────────────────────────┘

Reading the ELF File Header (readelf -h)

The ELF header is the first 64 bytes of every ELF file. It contains the metadata the kernel reads to decide how to load the program. Here is a real readelf -h output from a typical malicious ELF, with every field explained:

$ readelf -h suspicious.elf

ELF Header:
  Magic:   7f 45 4c 46 02 01 01 00 00 00 00 00 00 00 00 00
           ↑↑ ↑↑↑↑↑↑↑↑↑ ↑↑ ↑↑ ↑↑ ↑↑ ↑↑
           │  E  L  F    │  │  │  └─ padding zeros (9 bytes)
           │             │  │  └──── OS/ABI: 0x00 = System V (Linux uses this even though it's Linux ABI)
           │             │  └─────── ELF version: always 1
           │             └────────── Data encoding: 01 = little-endian (LSB), 02 = big-endian
           └──────────────────────── Class: 02 = ELF64 (64-bit), 01 = ELF32

  ┌─ FOCUS: Class tells you 32-bit vs 64-bit. Affects all pointer sizes, struct offsets.
  │  Malware targeting embedded devices or 32-bit systems shows 01 here.
  └─────────────────────────────────────────────────────────────────────────────────────

  Class:                             ELF64
  Data:                              2's complement, little endian
  Version:                           1 (current)
  OS/ABI:                            UNIX - System V
                                     ↑ Nearly all Linux binaries say "UNIX - System V" even though
                                       they're clearly Linux. Only a few use "Linux" ABI explicitly.
                                       MUSL-compiled malware often shows UNIX - System V.

  ABI Version:                       0

  Type:                              EXEC (Executable file)
                                     ↑ FOCUS: This field is critical.
                                     ┌─────────────────────────────────────────────────────────────
                                     │  EXEC = standard executable with a fixed load address (no PIE)
                                     │  DYN  = position-independent executable (PIE enabled) OR shared library
                                     │         Modern hardened binaries and most distro packages use DYN.
                                     │         If this is malware compiled as DYN, it has PIE (ASLR applies).
                                     │         If EXEC, it loads at a FIXED address every time — easier to ROP.
                                     │  REL  = relocatable object file (.o) — not a complete program
                                     │  CORE = core dump — may contain memory of a crashed program
                                     └─────────────────────────────────────────────────────────────

  Machine:                           Advanced Micro Devices X86-64
                                     ↑ Architecture. Other values you'll see:
                                       ARM (32-bit), AArch64 (64-bit ARM — phones, Raspberry Pi, AWS Graviton),
                                       MIPS (routers/IoT), PowerPC, RISC-V.
                                       Botnet malware often ships multi-arch builds.

  Version:                           0x1

  Entry point address:               0x401080
                                     ↑ FOCUS: This is where execution begins after the dynamic linker
                                       has finished loading. On a non-stripped binary this is usually
                                       _start (which calls __libc_start_main which calls main).
                                       On a PACKED binary the entry point is the packer stub, NOT main.
                                       The OEP (Original Entry Point) is somewhere else, found after unpacking.
                                       In Ghidra: press G → type 0x401080 to jump here immediately.

  Start of program headers:          64 (bytes into file)
                                     ↑ Segment table starts right after the ELF header (at offset 64).
                                       Always 64 for standard ELF64.

  Start of section headers:          28672 (bytes into file)
                                     ↑ Section header table offset. In a stripped binary this might be 0
                                       (section table stripped). The binary still runs — the kernel only
                                       needs program headers, not section headers.

  Flags:                             0x0
                                     ↑ Architecture-specific flags. x86-64 always 0. ARM uses this for
                                       Thumb mode, MIPS for ABI version, etc.

  Size of this header:               64 (bytes)
  Size of program headers:           56 (bytes)   ← each segment descriptor is 56 bytes
  Number of program headers:         11            ← 11 segments in this binary
  Size of section headers:           64 (bytes)
  Number of section headers:         28            ← 28 sections
  Section header string table index: 27            ← section #27 contains the section names

What to extract immediately from the ELF header:

FieldWhat it tells you
Type: EXECFixed load address — no ASLR on the binary itself → easier ROP chains
Type: DYNPIE — ASLR applies → need a leak to bypass
Entry point: 0x4xxxxxFor EXEC: entry is in .text. Anything else is suspicious.
Entry point: 0x0Malformed or position-independent — check program headers
Machine: AArch64ARM 64-bit — IoT / mobile malware
Machine: MIPSRouter malware (Mirai family)
Number of sections: 0Stripped section table — objdump -S won't work, use program headers

Reading the Section Headers (readelf -S)

Sections are the logical divisions of the file. readelf -S shows all of them with their type, offset, size, and permissions. Here is a real output with the security-relevant sections highlighted:

$ readelf -S --wide suspicious.elf

There are 28 section headers, starting at offset 0x7000:

Section Headers:
  [Nr] Name              Type            Address          Off    Size   ES Flg Lk Inf Al
  [ 0]                   NULL            0000000000000000 000000 000000 00      0   0  0
  [ 1] .interp           PROGBITS        0000000000400238 000238 00001c 00   A  0   0  1
       ↑ FOCUS: Contains the path to the dynamic linker as a string.
         Normal: /lib64/ld-linux-x86-64.so.2
         Suspicious: a custom path → malware may use a custom loader / LD_PRELOAD trick

  [ 2] .note.ABI-tag     NOTE            0000000000400254 000254 000020 00   A  0   0  4
  [ 3] .note.gnu.build-id NOTE           0000000000400274 000274 000024 00   A  0   0  4
       ↑ Build ID: a SHA1 hash of the binary computed at link time. Useful for correlation
         across VirusTotal and other intel sources. Extract with: readelf -n binary

  [ 4] .gnu.hash         GNU_HASH        0000000000400298 000298 000030 00   A  5   0  8
  [ 5] .dynsym           DYNSYM          00000000004002c8 0002c8 0000f0 18   A  6   1  8
       ↑ FOCUS: Dynamic symbol table. These are the IMPORTED functions.
         This section is NEVER stripped — it's required at runtime for dynamic linking.
         `nm -D binary | grep ' U '` reads this. The imported function list tells you
         what the binary CAN do: socket+connect = network, execve = exec, open+read = fs.

  [ 6] .dynstr           STRTAB          00000000004003b8 0003b8 000090 00   A  0   0  1
       ↑ String table for .dynsym names (function and library name strings)

  [ 7] .gnu.version      VERSYM          0000000000400448 000448 000014 02   A  5   0  2
  [ 8] .gnu.version_r    VERNEED         0000000000400460 000460 000030 00   A  6   1  8
       ↑ Which version of glibc each function requires. "GLIBC_2.17" = compatible with
         anything from 2017+. Useful for dating when the binary was built.

  [ 9] .rela.dyn         RELA            0000000000400490 000490 000018 18   A  5   0  8
  [10] .rela.plt         RELA            00000000004004a8 0004a8 0000c0 18  AI  5  23  8
       ↑ Relocation tables. Tell the dynamic linker where to patch function addresses.
         .rela.plt entries correspond 1:1 with PLT stubs — each entry is an imported function.

  [11] .init             PROGBITS        0000000000400568 000568 00001a 00  AX  0   0  4
       ↑ Code that runs BEFORE main() — constructor functions. Rootkits may hide here.

  [12] .plt              PROGBITS        0000000000400590 000590 0000a0 10  AX  0   0 16
       ↑ Procedure Linkage Table. For each imported function (connect, execve, etc.)
         there is a small stub here. First call → resolves address via dynamic linker,
         patches .got.plt. Subsequent calls → jump directly. Seeing these stubs in
         disassembly is normal. Their names come from .dynsym.

  [13] .plt.got          PROGBITS        0000000000400630 000630 000008 08  AX  0   0  8

  [14] .text             PROGBITS        0000000000400640 000640 0003d2 00  AX  0   0 16
       ↑ FOCUS: The actual code. Flags "AX" = Allocated (in memory) + eXecutable.
         Normal: this is the ONLY section with X flag.
         Suspicious: .data or .bss also has X flag → shellcode / self-modifying code.
         Suspicious: .text has very high entropy (>7.0) → packed/encrypted code stub.
         Suspicious: .text is tiny (<0x100 bytes) → likely a packer stub (real code elsewhere).
         Size matters: compare .text size against total file size. If .text is 5% of the file
         and the rest is high-entropy .data → the real code is compressed in .data.

  [15] .fini             PROGBITS        0000000000400a14 000a14 000009 00  AX  0   0  4
       ↑ Code that runs AFTER main() returns — destructor functions.

  [16] .rodata           PROGBITS        0000000000400a20 000a20 000048 00   A  0   0  4
       ↑ Read-only data. String literals live here on non-stripped, non-obfuscated binaries.
         `strings` on the binary extracts these. If .rodata is tiny or encrypted → obfuscation.

  [17] .eh_frame_hdr     PROGBITS        0000000000400a68 000a68 000024 00   A  0   0  4
  [18] .eh_frame         PROGBITS        0000000000400a90 000a90 000070 00   A  0   0  8
       ↑ Exception handling frame data. Required for C++ exceptions and backtraces.
         Present even in C code compiled with recent GCC. Can be used by analysts to
         find function boundaries in stripped binaries.

  [19] .init_array       INIT_ARRAY      0000000000601df8 001df8 000008 08  WA  0   0  8
       ↑ Array of constructor function pointers — run before main(). TLS callbacks live here
         in ELF (the Linux equivalent of the Windows TLS callback anti-debug technique).
         `readelf -S | grep init_array` then `readelf -x 19 binary` to dump the addresses.

  [20] .fini_array       FINI_ARRAY      0000000000601e00 001e00 000008 08  WA  0   0  8

  [21] .dynamic          DYNAMIC         0000000000601e08 001e08 0001d0 10  WA  6   0  8
       ↑ FOCUS: Dynamic linking metadata. Contains: library names (DT_NEEDED),
         address of symbol table (DT_SYMTAB), GOT address (DT_PLTGOT), init/fini addresses.
         Malware packed as a shared library often modifies DT_INIT to point to the unpacker stub.
         `readelf -d binary` reads this section. Key tags:
           DT_NEEDED  = shared library dependency (one per imported .so)
           DT_RPATH / DT_RUNPATH = custom library search path (suspicious if non-standard)
           DT_INIT    = address of _init() function (runs before main)
           DT_FINI    = address of _fini() function (runs after main)

  [22] .got              PROGBITS        0000000000601fd8 001fd8 000010 08  WA  0   0  8
  [23] .got.plt          PROGBITS        0000000000601fe8 001fe8 000060 08  WA  0   0  8
       ↑ FOCUS: Global Offset Table. Contains the resolved addresses of imported functions.
         After dynamic linking, connect@got.plt contains the real address of connect() in libc.
         GOT overwrite is a classic exploitation technique — overwrite a GOT entry to redirect
         a function call. FULL RELRO (from checksec) makes .got.plt read-only after load,
         preventing this attack. PARTIAL RELRO only protects .got, not .got.plt.

  [24] .data             PROGBITS        0000000000602048 002048 000010 00  WA  0   0  8
       ↑ Initialized global/static variables. Encrypted config blobs often live here.
         A large .data section with high entropy = stored encrypted payload.

  [25] .bss              NOBITS          0000000000602058 002058 000008 00  WA  0   0  8
       ↑ Uninitialized globals. Not stored on disk (NOBITS = no disk space).
         Zero-filled by OS at load time.

  [26] .symtab           SYMTAB          0000000000000000 003058 000660 18     27  49  8
       ↑ FOCUS: Full symbol table — present only in NON-STRIPPED binaries.
         Contains ALL function and variable names, including internal ones.
         Stripped binaries: this section is absent (or size 0).
         If present: gold mine for analysis — every function is named.
         `nm binary` reads this. `nm -D binary` reads .dynsym instead.

  [27] .strtab           STRTAB          0000000000000000 0036b8 000310 00      0   0  1
  [28] .shstrtab         STRTAB          0000000000000000 0039c8 000102 00      0   0  1

Section flags decoded

FlagMeaningSecurity relevance
AAlloc — section is loaded into memory at runtimeSections without A are debug/metadata only
XeXecute — memory region is executableONLY .text should have this; extras = shellcode
WWrite — memory region is writable.data/.bss/.got are W; .text being W = self-modifying code
IInfo linkInternal ELF bookkeeping
MMerge — duplicate entries can be removed
SStrings — contains null-terminated strings

Red flags in section headers

bash
# Any section that is both Writable AND eXecutable
readelf -S binary | grep -E 'WX|AXW'
# → indicates self-modifying code or injected shellcode region

# .text section has extremely high entropy
# Indicates the packer stub is there but real code is encrypted inside

# Section named .text but type is not PROGBITS
readelf -S binary | awk '/\.text/{print}'

# Missing .dynamic section
readelf -S binary | grep -c DYNAMIC
# → 0 means statically linked (common in malware for maximum portability)

# Suspicious section names
readelf -S binary | grep -iE 'UPX|vmp|enigma|themida|packed'

Reading the Program Headers (readelf -l)

Program headers describe segments — how the kernel actually maps the binary into memory. Here is a real readelf -l output annotated:

$ readelf -l --wide suspicious.elf

Elf file type is EXEC (Executable file)
Entry point 0x401080
There are 11 program headers, starting at offset 64

Program Headers:
  Type           Offset   VirtAddr           PhysAddr           FileSiz  MemSiz   Flg Align
  PHDR           0x000040 0x0000000000400040 0x0000000000400040 0x000268 0x000268 R   0x8
  ↑ PHDR: The program header table itself. Self-referential.

  INTERP         0x000238 0x0000000000400238 0x0000000000400238 0x00001c 0x00001c R   0x1
  ↑ FOCUS: Points to the dynamic linker path (reads the .interp section).
    Normal: /lib64/ld-linux-x86-64.so.2
    If this is missing: statically linked binary (no dynamic loader needed).
    Contents: readelf --string-dump=.interp binary

  LOAD           0x000000 0x0000000000400000 0x0000000000400000 0x000b00 0x000b00 R E 0x200000
  ↑ FOCUS: LOAD segments are what the kernel actually maps into memory.
    This one: Flags = "R E" = Read + Execute. This is the CODE segment.
    Contains: .text, .rodata, .plt, .init, .fini
    VirtAddr 0x400000 = the base load address for EXEC binaries (fixed).
    FileSiz = MemSiz = 0x000b00: no extra zeroed memory needed.

  LOAD           0x001df8 0x0000000000601df8 0x0000000000601df8 0x000268 0x000270 RW  0x200000
  ↑ FOCUS: Second LOAD segment. Flags = "RW" = Read + Write. This is the DATA segment.
    Contains: .data, .bss, .got, .got.plt, .dynamic
    MemSiz (0x000270) > FileSiz (0x000268) by 8 bytes → those 8 bytes are .bss (zeroed by OS).
    No X flag here = Writable data is not executable. GOOD.
    If you see a LOAD segment with "RWE" (Read+Write+Execute) → VERY SUSPICIOUS.
    Self-modifying code or shellcode that writes and executes itself lives in RWE regions.

  DYNAMIC        0x001e08 0x0000000000601e08 0x0000000000601e08 0x0001d0 0x0001d0 RW  0x8
  ↑ Points to the .dynamic section (dynamic linker metadata).

  NOTE           0x000254 0x0000000000400254 0x0000000000400254 0x000044 0x000044 R   0x4
  ↑ Build notes (.note.ABI-tag, .note.gnu.build-id).

  GNU_EH_FRAME   0x000a68 0x0000000000400a68 0x0000000000400a68 0x000024 0x000024 R   0x4
  ↑ Exception handling frame info. Used by debuggers and libgcc for stack unwinding.

  GNU_STACK      0x000000 0x0000000000000000 0x0000000000000000 0x000000 0x000000 RW  0x10
  ↑ FOCUS: Stack permissions. Flags "RW" = Read + Write only (no execute).
    This is correct — the stack should NEVER be executable (NX/DEP protection).
    If you see "RWE" here → NX is disabled → stack-based shellcode can execute.
    Equivalent to checksec showing "NX: disabled".

  GNU_RELRO      0x001df8 0x0000000000601df8 0x0000000000601df8 0x000208 0x000208 R   0x1
  ↑ FOCUS: RELRO (Relocation Read-Only) marking. After dynamic linking finishes,
    the memory range in this segment is made read-only (mprotect to R only).
    PARTIAL RELRO: only .got (non-PLT) is in this segment — .got.plt remains writable.
    FULL RELRO: both .got and .got.plt are in this segment → GOT is fully read-only.
    Absence of GNU_RELRO → no RELRO → GOT is always writable → trivial GOT overwrite.

 Section to Segment mapping:
  Segment Sections...
   00
   01     .interp
   02     .interp .note.ABI-tag .note.gnu.build-id .gnu.hash .dynsym .dynstr
          .gnu.version .gnu.version_r .rela.dyn .rela.plt .init .plt .plt.got
          .text .fini .rodata .eh_frame_hdr .eh_frame
   03     .init_array .fini_array .dynamic .got .got.plt .data .bss
   04     .dynamic
   05     .note.ABI-tag .note.gnu.build-id
   06     .eh_frame_hdr
   07
   08     .init_array .fini_array .dynamic .got

Quick checklist from program headers

bash
# 1. Any RWE LOAD segment? (writable AND executable)
readelf -l binary | grep 'LOAD' | grep 'RWE\| RW.*E\| R.*WE'
# If yes: shellcode / unpacker stub — serious red flag

# 2. Stack executable?
readelf -l binary | grep 'GNU_STACK'
# Should show: RW  not RWE
# RWE = NX disabled = checksec "NX: disabled"

# 3. RELRO present?
readelf -l binary | grep 'GNU_RELRO'
# Absent = no RELRO = GOT overwrite possible

# 4. Entry point inside the expected LOAD segment?
readelf -h binary | grep 'Entry'
readelf -l binary | grep 'LOAD'
# Entry should be inside the R E LOAD segment (code), not the RW segment (data)
# Entry in a data segment → unpacker stub running from a data region → packed

Stripped vs Non-Stripped

bash
# Non-stripped: rich symbol table — gdb shows function names
readelf -s binary | grep -c FUNC   # dozens or hundreds of lines
nm binary | head -20

# Stripped: only dynamic symbols remain
file binary
# → stripped

nm binary
# → nm: binary: no symbols

# Dynamic symbols (imported functions) are never stripped — they are required at runtime
nm -D binary | grep ' U '   # undefined = imported from .so
# Example: printf@@GLIBC_2.2.5, socket@@GLIBC_2.2.5, connect@@GLIBC_2.2.5

Key insightEven stripped binaries expose their imported function names via dynamic symbols. socket, connect, bind, recv, send → network capability. popen, execve, system → command execution. open, read, write, unlink → filesystem operations.


Sections That Should Not Exist

bash
readelf -S binary | grep -E 'Name|Flags'

# Suspicious section names:
# .UPX0, .UPX1, .UPX2   → UPX packer
# .packed                 → custom packer
# .enigma, .vmp          → VMProtect / Enigma
# No section names at all → heavily stripped or custom format

# Missing sections that should exist:
# No .text section → code is in an unnamed section (anti-analysis)
# No .dynamic section → statically linked (common in malware for portability)

Pattern Extraction — Strings and IOCs

Network Indicators

bash
# URLs and URIs
strings -n 8 binary | grep -oE 'https?://[a-zA-Z0-9./_?=&%-]+'
strings -n 8 binary | grep -oE 'ftp://[a-zA-Z0-9./_?=&%-]+'

# IP addresses (IPv4)
strings -n 7 binary | grep -oE '\b([0-9]{1,3}\.){3}[0-9]{1,3}\b' \
  | grep -v '^127\.' | grep -v '^0\.'

# IPv6
strings binary | grep -oE '([0-9a-fA-F]{1,4}:){7}[0-9a-fA-F]{1,4}'

# Domain names (catch C2 domains)
strings -n 8 binary | grep -oE '[a-zA-Z0-9.-]+\.(com|net|org|io|ru|cn|cc|tk|top|xyz|info)'

# Port numbers in context (look for socket configuration)
strings binary | grep -oE ':[0-9]{2,5}\b'

# User-Agent strings (HTTP C2 often has a hardcoded UA)
strings binary | grep -i 'User-Agent\|Mozilla\|curl\|python-requests'

Credentials and Secrets

bash
# Hardcoded passwords and keys
strings -n 8 binary | grep -iE 'password|passwd|passw|secret|apikey|api_key|access_key|auth_token'

# Base64-encoded blobs (potential embedded keys or second-stage payloads)
strings binary | grep -oE '[A-Za-z0-9+/]{20,}={0,2}' | while read b64; do
    echo "$b64" | base64 -d 2>/dev/null | strings -n 4
done

# SSH key headers
strings binary | grep -E 'BEGIN (RSA|EC|OPENSSH|PRIVATE|PUBLIC) KEY'

# AWS credentials pattern
strings binary | grep -oE 'AKIA[A-Z0-9]{16}'
strings binary | grep -oE '[a-zA-Z0-9/+]{40}'

Filesystem Indicators

bash
# Persistence paths
strings binary | grep -E '/etc/(cron|rc|init|systemd|profile|bash|passwd|shadow|ld\.so)'
strings binary | grep -E '(~|/home|/root)/\.(bashrc|profile|ssh|config|local)'
strings binary | grep -E '/var/(spool|run|tmp|log)'
strings binary | grep -E '\.service$|\.timer$|\.socket$'   # systemd unit persistence

# Suspicious temp paths
strings binary | grep -E '/tmp/\.|/dev/shm/\.'    # hidden files in temp directories

# /proc filesystem access (rootkit, credential theft)
strings binary | grep -E '/proc/[0-9]*/|/proc/self/'
strings binary | grep -E '/proc/net/|/proc/sys/'

Command Execution Patterns

bash
# Shell commands embedded in binary
strings binary | grep -E '^(bash|sh|zsh|python|perl|ruby|php|node) '
strings binary | grep -E 'execve|system\(|popen\(|fork\(\)'
strings binary | grep -E 'chmod [0-7]{3,4}|chown root|setuid|setgid'

# Reverse shell indicators
strings binary | grep -E '/dev/tcp/|/dev/udp/'
strings binary | grep -E 'bash -i|nc -e|ncat|socat'
strings binary | grep -iE 'reverse.?shell|bind.?shell'

# Privilege escalation
strings binary | grep -E 'sudo|su -|NOPASSWD|visudo'
strings binary | grep -E 'LD_PRELOAD|LD_LIBRARY_PATH'

# C2 command keywords (common RAT command strings)
strings binary | grep -iE '\b(shell|exec|upload|download|screenshot|keylog|persist)\b'

Crypto and Ransom Indicators

bash
# Ransomware patterns
strings binary | grep -iE 'encrypt|decrypt|AES|RSA|ransom|bitcoin|wallet|tor'
strings binary | grep -iE '\.encrypted$|\.locked$|\.crypto$'
strings binary | grep -iE 'README|HELP|HOW.TO.DECRYPT|RECOVERY'

# Cryptocurrency miner patterns
strings binary | grep -iE 'stratum\+tcp|mining.pool|xmrig|monero|hashrate'
strings binary | grep -iE 'nicehash|pool\.minexmr\.com|cryptonight'

Obfuscation and Packing Detection

High-Entropy Detection — Per Section

bash
python3 << 'EOF'
import pefile, math, sys

def section_entropy(data):
    if not data: return 0
    freq = {}
    for b in data: freq[b] = freq.get(b, 0) + 1
    return -sum((c/len(data)) * math.log2(c/len(data)) for c in freq.values())

try:
    pe = pefile.PE(sys.argv[1])
    print(f"{'Section':<20} {'Size':>10} {'Entropy':>8}  Verdict")
    print("-" * 55)
    for s in pe.sections:
        data = s.get_data()
        ent = section_entropy(data)
        verdict = "PACKED/ENCRYPTED" if ent > 7.0 else ("suspicious" if ent > 6.5 else "ok")
        print(f"{s.Name.decode().strip():<20} {len(data):>10}  {ent:>7.3f}  {verdict}")
except Exception as e:
    # For ELF, use a different approach
    import struct
    with open(sys.argv[1], 'rb') as f:
        data = f.read()
    chunk_size = 4096
    print(f"{'Offset':<12} {'Entropy':>8}  Verdict")
    for i in range(0, len(data), chunk_size):
        chunk = data[i:i+chunk_size]
        ent = section_entropy(chunk)
        if ent > 6.5:
            print(f"0x{i:<10x} {ent:>7.3f}  {'HIGH ENTROPY'}") 
EOF suspicious_binary

UPX Detection and Unpacking

UPX is the most common open-source packer. Detecting and unpacking it is straightforward:

bash
# Detection
strings binary | grep -iE 'UPX[0-9!]'
readelf -S binary | grep -E 'UPX[0-9]'
upx -t binary     # test if valid UPX

# Unpack in place
upx -d -o binary_unpacked binary

# If UPX headers are stripped (common anti-analysis measure):
# The attacker removed "UPX0"/"UPX1" magic strings to prevent automatic unpacking
# Solution: run the binary in a debugger to the OEP, then dump memory
# x64dbg: run to entry, then Plugins → Scylla → Dump + Fix IAT

Identifying the Real Entry Point (OEP)

When a packed binary runs, the stub decrypts the payload and jumps to the Original Entry Point (OEP). Finding it lets you dump the unpacked binary.

bash
# Strategy 1: Look for a "tail jump" (jmp to OEP after unpacking loop)
objdump -d packed_binary | tail -30
# Common pattern: loop ending with jmp [eax] or jmp [rax] — that's the OEP jump

# Strategy 2: hardware breakpoint on VirtualAlloc return
# The stub allocates memory, writes the real binary, then jumps to it.
# Break when VirtualAlloc returns, note the allocation address.
# Set a hardware execute breakpoint at that address.
# Run → hit the breakpoint at the OEP.

Dynamic Analysis — strace and ltrace

GoalObserve what the malware actually does when running — network connections, files created, processes spawned, credentials read. Dynamic analysis bypasses all static obfuscation: XOR-encrypted strings are decrypted before use and appear in function arguments you can observe; packed code is unpacked into memory and executes normally.

Why dynamic is essential alongside staticA binary with zero interesting strings in static analysis may have a full C2 URL and command list that appears only at runtime after decryption. strace captures every system call — it doesn't matter how obfuscated the code is, because system calls are the boundary between userspace and kernel and cannot be avoided.

The isolation requirementdynamic analysis means running the malware, so it must happen somewhere it can't hurt you:

Network

Host-only or no networking, or a sinkholed network where every DNS name resolves to 127.0.0.1, so C2 calls go nowhere.

Snapshot

Take one before every run and revert afterwards. Never reuse a "dirty" VM.

No shared folders

Nothing mounted between the VM and the host — a shared folder is a path out of the sandbox.

Fresh identity

Different MAC address and hostname each run: malware fingerprints analysis environments and goes quiet.

strace — System Call Tracer

bash
# Trace all syscalls with timestamps
strace -tt ./binary

# Follow child processes (forked daemons, shell spawns)
strace -f ./binary 2>&1 | tee strace.log

# Trace a specific category of syscalls
strace -e trace=network ./binary       # socket, connect, bind, recv, send
strace -e trace=file ./binary          # open, read, write, stat, unlink
strace -e trace=process ./binary       # execve, fork, clone, exit
strace -e trace=signal ./binary        # signal, sigaction
strace -e trace=ipc ./binary           # shared memory, semaphores

# Count syscalls (profiling)
strace -c ./binary

# Attach to a running process
strace -p <pid>
strace -fp <pid>   # follow children too

Key patterns to look for in strace output:

bash
# C2 connection attempt
connect(3, {sa_family=AF_INET, sin_port=htons(4444), sin_addr=inet_addr("1.2.3.4")}, 16)

# File creation (dropper)
openat(AT_FDCWD, "/tmp/.cache/.x", O_WRONLY|O_CREAT|O_TRUNC, 0755)
write(3, "\x7fELF\x02\x01\x01\x00", 64)   # writing a new ELF to disk

# Privilege escalation attempt
setuid(0)   → EPERM (failed — not running as root yet)
# or
openat(AT_FDCWD, "/etc/sudoers", O_RDONLY) = -1 EACCES

# Credential access
openat(AT_FDCWD, "/etc/shadow", O_RDONLY) = -1 EACCES
openat(AT_FDCWD, "/proc/1/environ", O_RDONLY) = -1 EPERM

# Shell spawn
execve("/bin/bash", ["/bin/bash", "-i"], 0x7fff... /* 23 vars */)

# Persistence (systemd service or cron)
openat(AT_FDCWD, "/etc/systemd/system/malware.service", O_WRONLY|O_CREAT)

# Network listen (bind shell or local port for lateral movement)
bind(3, {sa_family=AF_INET, sin_port=htons(1337), sin_addr=inet_addr("0.0.0.0")}, 16)
listen(3, 5)

# Anti-debug check
ptrace(PTRACE_TRACEME, 0, NULL, NULL) = -1 EPERM  ← already traced → self-detects
openat(AT_FDCWD, "/proc/self/status", O_RDONLY)   ← reading TracerPid

ltrace — Library Call Tracer

ltrace intercepts calls to shared libraries. Unlike strace (kernel boundary), ltrace works at the libc boundary — you see printf, malloc, strcmp, connect etc.

bash
# Basic trace
ltrace ./binary

# Follow child processes
ltrace -f ./binary

# Filter for specific functions
ltrace -e 'strcmp+strcpy+strcat+malloc+free+connect' ./binary

# Show timestamps
ltrace -tt ./binary

# Attach to running process
ltrace -p <pid>

Key patterns in ltrace output

bash
# String comparison (often reveals decrypted passwords being checked)
strcmp("admin", "admin")  = 0          # match
strcmp("correct_password", "wrongpass") = non-zero

# Memory operations (can reveal decrypted strings mid-execution)
malloc(256)  = 0x55a1b2c3
# ... binary copies decrypted string here ...
puts("https://c2.evil.com/beacon")     # now visible in plaintext

# File operations
fopen("/etc/passwd", "r") = 0x55a1...
fread(buf, 1, 4096, 0x55a1...) = 1024  # reading credentials

# Network operations (higher level than strace's connect)
getaddrinfo("c2.evil.com", "443", ...)   # DNS resolution
connect(sockfd, {AF_INET, "1.2.3.4", 443}, ...)

# Crypto operations (if using OpenSSL / libcrypto)
EVP_EncryptInit_ex(ctx, EVP_aes_256_cbc(), ...)   # AES encryption starting

ltrace vs strace trade-offstrace is always available and harder to evade (syscall level). ltrace requires debug symbol hooks and can be defeated by statically-linked binaries (no shared lib calls to intercept). Use both.


GDB — Debugging Basics

GoalStep through the binary's execution at instruction level, inspect and modify memory and registers in real time, extract decrypted data, and bypass anti-debug checks. GDB is what you reach for when strace/ltrace tells you what happened but not how — or when you need to intercept a decrypt function and read the plaintext before it gets used.

When to use GDB over strace

  • You need to see inside a function call (strace only shows the syscall boundary)
  • You want to extract a decrypted string or key from a register or memory region
  • The binary is packed and you need to find the OEP (set a breakpoint on mprotect or mmap)
  • You need to bypass an anti-debug check (patch ptrace return value, skip a JNZ)
bash
# Start debugging
gdb ./binary
gdb -q ./binary    # quiet mode (no banner)

# Attach to running process
gdb -p <pid>

# Better experience: install pwndbg or gdb-peda
pip install pwndbg   # then: echo "source /path/to/pwndbg/gdbinit.py" >> ~/.gdbinit

Essential GDB Commands

# Execution control
run [args]           → start the program
run < input_file     → start with stdin redirected
start                → run and break at main()
continue / c         → resume execution
next / n             → step over (don't enter function calls)
step / s             → step into function calls
finish               → run until current function returns
until <line>         → run until source line

# Breakpoints
break main           → break at function main
break *0x401234      → break at absolute address
break *main+0x20     → break 32 bytes into main
info breakpoints     → list all breakpoints
delete 1             → delete breakpoint #1
disable 1            → disable without deleting
condition 1 rax==0   → conditional breakpoint

# Hardware breakpoints (no INT 3 — evades code-checksum anti-debug)
hbreak *0x401234          → hardware execute breakpoint
watch *0x603018           → hardware write watchpoint (fires when address written)
rwatch *0x603018          → hardware read watchpoint

# Examining memory
x/10i $rip           → disassemble 10 instructions from current RIP
x/20x $rsp           → dump 20 hex words from stack pointer
x/s 0x402000         → print string at address
x/b *0x401234        → read single byte
print $rax           → print register value
info registers       → show all registers
info registers rax rbx rsp rip   → specific registers

# Memory search
find 0x400000, 0x500000, "password"    → search for string in range
find /b 0x400000, 0x500000, 0x90, 0x90 → search for NOP NOP bytes

# Modify execution
set $rax = 0         → overwrite register
set *0x601020 = 0    → overwrite memory
set {int}0x601020 = 1337

# Shared libraries and symbols
info sharedlibrary   → loaded .so files and their address ranges
info functions       → list all known function symbols
info symbols 0x401234 → find nearest symbol to address

# Backtrace (call stack)
backtrace / bt       → show call stack
frame 2              → switch to frame 2

GDB Scripting for Anti-Debug Bypass

python
# gdb_bypass.py — auto-bypass ptrace PTRACE_TRACEME
# Run: gdb -x gdb_bypass.py ./binary

import gdb

class PtraceBypass(gdb.Breakpoint):
    def stop(self):
        # When ptrace is called, check if it's PTRACE_TRACEME (arg0 == 0)
        ptrace_request = int(gdb.parse_and_eval("$rdi"))
        if ptrace_request == 0:  # PTRACE_TRACEME
            gdb.execute("set $rax = 0")    # fake success
            gdb.execute("return")          # skip the actual syscall
            print("[*] ptrace(PTRACE_TRACEME) intercepted — returning 0")
        return False   # don't stop

PtraceBypass("ptrace", internal=False)
gdb.execute("run")

Frida — Dynamic Instrumentation

Frida injects a JavaScript engine into any process without recompilation. Use it to trace function calls, hook specific functions, and extract data from memory at runtime.

bash
# Install
pip install frida frida-tools

# List running processes
frida-ps

# Attach and run a script
frida -l script.js <process_name_or_pid>

# Spawn a new process under Frida
frida -l script.js -f ./binary --no-pause

# Trace all calls to a specific function
frida-trace -i "malloc" ./binary
frida-trace -i "connect" -i "send" -i "recv" ./binary
frida-trace -i "fopen" -i "fread" ./binary

# Trace all calls to any function matching a pattern
frida-trace -I "libc*" ./binary    # trace everything in libc

Frida Script Examples

javascript
// Hook connect() to log all network connections
Interceptor.attach(Module.getExportByName(null, "connect"), {
    onEnter: function(args) {
        // args[1] = sockaddr struct
        var sa_family = Memory.readU16(args[1]);
        if (sa_family == 2) {  // AF_INET
            var port = Memory.readU16(args[1].add(2));
            port = ((port & 0xFF) << 8) | ((port >> 8) & 0xFF);  // ntohs
            var ip_bytes = Memory.readByteArray(args[1].add(4), 4);
            var ip = Array.from(new Uint8Array(ip_bytes)).join(".");
            console.log("[connect] " + ip + ":" + port);
        }
    }
});
javascript
// Hook fopen to log all file opens
Interceptor.attach(Module.getExportByName(null, "fopen"), {
    onEnter: function(args) {
        var path = Memory.readUtf8String(args[0]);
        var mode = Memory.readUtf8String(args[1]);
        console.log("[fopen] " + path + " (" + mode + ")");
    }
});
javascript
// Intercept and log XOR decryption to extract decrypted strings
// Assumes decrypt_string(char* output, const char* input, size_t len, char key)
var decrypt_addr = ptr("0x401234");   // address from Ghidra/objdump
Interceptor.attach(decrypt_addr, {
    onEnter: function(args) {
        this.output_ptr = args[0];
        this.len = parseInt(args[2]);
    },
    onLeave: function(retval) {
        // After function returns, read the decrypted string
        var decrypted = Memory.readUtf8String(this.output_ptr, this.len);
        console.log("[decrypt_string] → " + decrypted);
    }
});
javascript
// Universal: dump all string arguments to any function (spray-and-pray)
["strcmp", "strcpy", "strcat", "puts", "printf"].forEach(function(fn) {
    var addr = Module.getExportByName(null, fn);
    if (addr) {
        Interceptor.attach(addr, {
            onEnter: function(args) {
                try {
                    var s = Memory.readUtf8String(args[0]);
                    if (s && s.length > 3) console.log("[" + fn + "] " + s);
                } catch(e) {}
            }
        });
    }
});

YARA — Writing Detection Rules

YARA is a pattern-matching engine for identifying malware. Rules describe byte patterns, string patterns, and structural conditions.

bash
# Install
apt install yara
pip install yara-python

# Scan a file
yara rule.yar suspicious_binary

# Scan a directory recursively
yara -r rules/ /path/to/samples/

# Scan running processes
yara -p 50 rule.yar   # uses /proc/<pid>/mem

Writing a Rule from Analysis

yara
rule Suspicious_ELF_C2_Beacon {
    meta:
        description = "ELF binary with hardcoded C2 IP and XOR string obfuscation pattern"
        author = "analyst"
        severity = "high"

    strings:
        // Network indicators found via strings
        $c2_ip    = "192.168.10.100" ascii
        $c2_path  = "/beacon" ascii
        $ua       = "Mozilla/5.0 (compatible; malbot/1.0)" ascii

        // XOR decryption loop pattern (x86-64 bytes)
        // lea rsi, [rip+offset]  ; xor [rsi+rcx], al ; loop
        $xor_loop = { 48 8D 35 ?? ?? ?? ?? 30 04 0E 48 FF C1 }

        // Anti-debug: ptrace PTRACE_TRACEME check
        $ptrace_check = { B8 65 00 00 00 0F 05 48 85 C0 }  // syscall 101 = ptrace

        // Persistence path
        $persist   = "/etc/cron.d/updater" ascii
        $persist2  = "/tmp/.x" ascii nocase

        // Interesting imported symbols
        $import_execve = "execve" ascii
        $import_conn   = "connect" ascii

    condition:
        // Must be an ELF binary
        uint32(0) == 0x464C457F
        // At least 2 network or execution indicators
        and 2 of ($c2_ip, $c2_path, $ua, $import_execve, $import_conn)
        // Anti-debug or xor loop
        and ($ptrace_check or $xor_loop)
        // And at least one persistence indicator
        and 1 of ($persist, $persist2)
}

Rule for Cryptominer Detection

yara
rule Cryptominer_XMRig_Pattern {
    meta:
        description = "XMRig or XMRig-compatible miner"

    strings:
        $pool1    = "stratum+tcp://" ascii
        $pool2    = "pool.minexmr.com" ascii
        $pool3    = "mining.pool" ascii
        $algo1    = "cryptonight" ascii nocase
        $algo2    = "randomx" ascii nocase
        $wallet   = "4" ascii     // Monero wallet starts with 4
        $xmrig    = "xmrig" ascii nocase
        $hashrate = "H/s" ascii
        $config   = "--donate-level" ascii

    condition:
        uint32(0) == 0x464C457F
        and (
            2 of ($pool1, $pool2, $pool3)
            or ($xmrig and 1 of ($algo1, $algo2))
            or ($hashrate and $config)
        )
}

Rule from Byte Pattern (Shellcode)

yara
rule Reverse_Shell_Shellcode {
    meta:
        description = "x86-64 reverse shell shellcode pattern (execve /bin/sh)"

    strings:
        // Classic x64 execve("/bin/sh", ...) syscall
        // push 0x68 ; mov rbx, 0x732f2f2f6e69622f ; push rbx ; push rsp ; pop rdi
        $execve_binsh = {
            6A 68
            48 B8 2F 62 69 6E 2F 73 68 00
            50 54 5F
        }

        // syscall 59 (execve)
        $syscall_execve = { B8 3B 00 00 00 0F 05 }

        // /bin/sh string
        $binsh = "/bin/sh" ascii

    condition:
        all of them
        or ($execve_binsh and $syscall_execve)
}

Decompilation Tools

ToolTypeStrengthsHow to use
GhidraDecompiler + RE platformFree, NSA-developed, excellent decompiler, scripting (Python/Java), collaborativeghidraRun → import file → auto-analyse
IDA ProDisassembler + decompilerIndustry standard, best FLIRT signatures, Hex-Rays decompilerCommercial; IDA Free available
Radare2 (r2)Disassembler + analysisCLI, scriptable, supports 100+ archs, carve from memoryr2 -A binary
CutterGUI for Radare2Same power as r2, visual call graph, Ghidra decompiler plugincutter binary
Binary NinjaDisassembler + decompilerClean UI, HLIL/MLIL IR, Python scriptingCommercial; personal licence available
RetDecDecompilerFree online/CLI, outputs C pseudocoderetdec-decompiler binary
angrSymbolic executionPath exploration, vulnerability finding, constraint solvingPython library

Ghidra Quick Reference

bash
# Run headless analysis and export decompiled C
analyzeHeadless /path/to/project MyProject \
  -import suspicious_binary \
  -postScript DecompileHeadless.py \
  -scriptPath /path/to/scripts/ \
  -deleteProject

# Inside Ghidra GUI key shortcuts:
# G         → go to address
# L         → set label (rename symbol)
# F         → open function at cursor
# Ctrl+F    → search text in listing
# Ctrl+L    → search labels/symbols
# X         → cross references (who calls this function)
# P         → disassemble at cursor
# Ctrl+Alt+K → define function at cursor
# Right-click → Rename Variable (in decompiler view)

Radare2 Quick Reference

bash
r2 -A binary      # open and auto-analyse
r2 -d binary      # open in debugger mode

# Inside r2:
aa                # analyse all (equivalent to -A flag)
afl               # list all functions
pdf @ main        # print disassembly of function main
pdf @ sym.func_name
pdg @ main        # print decompiled code (needs r2ghidra plugin)
iz                # list strings in data sections
izz               # list strings in entire binary
iE                # list exports
ii                # list imports
is                # list symbols
axt @ sym.connect # cross-references to connect (who calls it)
/r 0x401234       # find cross-references to address
s sym.main; pdf   # seek to main, disassemble
VV                # visual call graph (ASCII art)
q                 # quit

Malware Family Patterns

Cryptominer Indicators

bash
# String patterns
strings binary | grep -iE 'stratum\+tcp|xmrig|minerd|cryptonight|randomx'
strings binary | grep -iE 'hashrate|difficulty|pool_address|wallet_address'

# Behavioral: CPU usage skyrockets
# strace shows network connection to pool, followed by continuous send/recv
strace -e trace=network ./miner 2>&1 | grep -E 'connect|send|recv'

# Process name masquerading as kernel worker
strings binary | grep -E 'kworker|kdevtmpfs|migration|ksoftirqd'

Botnet / RAT Indicators

bash
# C2 beaconing pattern in strace
# Regular connect() + send() + sleep() loop
grep -E '(connect|send|recv|sleep)' strace.log | head -30

# Command handling keywords
strings binary | grep -iE '\b(shell|exec|cmd|upload|download|screenshot)\b'
strings binary | grep -iE 'reverse.?shell|bind.?shell|pty|tty'

# Persistence mechanisms
strings binary | grep -E '(crontab|rc\.local|\.bashrc|\.profile|systemd|service)'
strings binary | grep -E 'chattr\s*\+i'   # make self immutable

# Process duplication / daemonization
# strace shows: fork() → setsid() → fork() again (double-fork daemon pattern)
grep -E 'fork|setsid|chdir.*root' strace.log

Rootkit Indicators

bash
# LD_PRELOAD-based rootkit
strings binary | grep 'LD_PRELOAD\|/etc/ld.so.preload'
strings binary | grep -E 'readdir|getdents|opendir'   # hiding files by intercepting these

# Hooking patterns
strings binary | grep -E 'dlsym|dlopen|RTLD_NEXT'   # RTLD_NEXT = classic LD_PRELOAD hook
# RTLD_NEXT tells dlsym to find the NEXT occurrence of the function (the real libc one)

# Self-hiding
strings binary | grep -E 'list_del|kobject_del|THIS_MODULE'   # LKM self-hiding

# /proc manipulation
strings binary | grep '/proc/net/tcp\|/proc/pid\|/proc/self'

Full Workflow Example

bash
# --- PHASE 1: TRIAGE ---
file sample.elf
# → ELF 64-bit LSB executable, x86-64, dynamically linked, stripped

sha256sum sample.elf   # check VT

xxd sample.elf | head -2
# → 7f 45 4c 46 02 01 01 00  (ELF)

python3 entropy.py sample.elf
# → Overall entropy: 6.82  (moderately high — possible packing)

# --- PHASE 2: STATIC ---
strings -n 8 sample.elf | grep -iE 'http|tcp|connect|exec|/tmp|passwd'
# → stratum+tcp://pool.minexmr.com:4444
# → /tmp/.x
# → kworker/0:1H

readelf -d sample.elf | grep NEEDED
# → libpthread, libc — standard; no libssl (may use embedded crypto)

nm -D sample.elf | grep ' U '
# → connect, socket, send, recv, fork, setsid, execve, prctl

checksec --file sample.elf
# → NX: enabled, PIE: disabled, No Canary, stripped

# Verdict: likely a cryptominer masquerading as a kernel worker thread
# using stratum protocol. Daemonises itself (fork+setsid). Writes to /tmp.

# --- PHASE 3: DYNAMIC (isolated VM) ---
strace -f ./sample.elf 2>&1 | tee strace.log &
sleep 5; kill %1

grep 'connect\|execve\|openat.*tmp\|prctl' strace.log
# → connect(3, {AF_INET, 195.148.127.141, 4444}, 16)  = 0
# → prctl(PR_SET_NAME, "kworker/0:1H")
# → openat("/tmp/.x", O_WRONLY|O_CREAT|O_TRUNC, 0755)

# Confirmed: network connection to miner pool, process rename, drop to /tmp

# --- PHASE 4: YARA ---
cat > miner.yar << 'EOF'
rule XMRig_Miner {
    strings:
        $pool = "stratum+tcp://" ascii
        $proc = "kworker" ascii
        $tmp  = "/tmp/." ascii
    condition:
        uint32(0) == 0x464C457F and all of them
}
EOF
yara miner.yar sample.elf
# → XMRig_Miner sample.elf

# --- PHASE 5: DECOMPILATION (Ghidra) ---
# Import → auto-analyse → search for "connect" in symbol tree
# Trace callers → find beacon loop → rename functions → export IOCs

Analyst OPSEC Failures

Everything below is something an analyst does during normal investigation that inadvertently signals to the attacker that their malware has been found. Sophisticated threat actors — APT groups, ransomware operators, some crimeware authors — actively monitor for these signals and use them to burn infrastructure, update tooling, or escalate before containment is complete.

The central principle: every action that reaches attacker-controlled infrastructure or a public platform the attacker can monitor is a potential burn.


Threat Intelligence Lookups

VirusTotal hash queries

VT logs every API and web query. Threat actors with VirusTotal Intelligence (a paid tier) can set up YARA or hash-based Hunting notifications — they receive an email when someone searches for or submits their sample. A hash lookup tells the attacker:

  • Their implant was discovered and is under investigation
  • Approximate time of discovery (they can correlate with C2 beaconing logs to identify the victim)
  • If you use a corporate VT account: your organisation
OPSEC failure:   curl -s "https://www.virustotal.com/api/v3/files/<sha256>" -H "x-apikey: $VT_KEY"
Safer:           Query private feeds first — CrowdStrike Intelligence, Recorded Future, Mandiant,
                 internal MISP. Only query VT after containment is complete and you accept the burn.

File uploads to public sandboxes

Uploading the binary (not just the hash) to Any.run, Hybrid-Analysis, or Joe Sandbox with public visibility is worse than a hash query. The full analysis report becomes searchable by anyone. Attackers can search their own malware name or hash and see your report — which may include:

  • The original filename (may contain the victim's hostname or path)
  • Your submitter country
  • The exact time you ran it
  • Network indicators you triggered, revealing how far you traced the campaign
OPSEC failure:   Drag malware onto hybrid-analysis.com → submit as "Public"
Safer:           Use private submission (Any.run has private mode, Joe Sandbox Enterprise does too).
                 Better: self-hosted Cuckoo/CAPE sandbox — nothing leaves your network.

WHOIS and certificate lookups via public resolvers

WHOIS queries themselves are not dangerous (registrars don't notify registrants). However, certificate transparency log searches (crt.sh, Censys, Shodan certificate search) can reveal that a specific cert was queried — and Shodan/Censys notify some customers of searches against their infrastructure.


DNS and Network Exposure

Resolving C2 domains from your analysis network

When malware runs in your sandbox and makes a DNS request for c2.example[.]com, that query leaves your network via your DNS resolver. The authoritative DNS server (which the attacker controls or can read logs from) sees:

  • The query
  • Your resolver's IP (often the corporate NAT or cloud egress IP)
  • Timestamp

Even if you only look up the domain manually (dig c2.example.com) from your analyst machine, that query is visible.

OPSEC failure:   dig malware-c2.attacker-domain.com
                 curl -I http://185.220.101.12/beacon   # "just checking if it's up"
Safer:           Use passive DNS databases — Farsight DNSDB, Cisco Umbrella Investigate,
                 VirusTotal passive DNS (check TTL/resolution history without making live queries).
                 Run dynamic analysis in a network-isolated sandbox that sinkhole-resolves DNS.

Directly connecting to C2 infrastructure to verify it's alive

Connecting to the C2 server from an analyst machine — even a single TCP SYN — gives the attacker:

  • Your IP address (confirms investigation)
  • User-Agent if HTTP (analyst tools have recognisable UAs)
  • Timing

C2 frameworks like Cobalt Strike, Sliver, and Mythic log all inbound connections. Some operators have tripwires: if a non-beacon IP connects to a listener port, they get an alert.

OPSEC failure:   nmap -sV 185.220.101.12        # port scan from corp IP
                 curl http://185.220.101.12/     # verify HTTP server
Safer:           Check from passive sources: Shodan/Censys historical data, Censys IPv4 banners,
                 Greynoise context. Never touch attacker infra from a corp or identifiable IP.
                 If you must probe: use a disposable VPS in a neutral cloud region with no links to your org.

Scanning attacker infrastructure from a corporate IP

Even if you use nmap rather than curl, a scan from a corporate IP range against a threat actor's server tells them exactly what they need to know. Ransomware groups have posted in forums about detecting investigative scans within hours of deployment.


Sandbox and Analysis Environment Fingerprinting

Re-using analysis VMs with consistent fingerprints

Many malware families perform sandbox detection by checking for consistent environment characteristics:

  • Hostname seen in a previous run (DESKTOP-ANALYSIS01, ANY.RUN, CUCKOO)
  • Same MAC address prefix (VMware: 00:0C:29, VirtualBox: 08:00:27)
  • Specific usernames (analyst, malware, sandbox, virus)
  • Absence of real user artefacts (no browser history, no documents, no real installed apps, only fresh install)
  • Screen resolution exactly 1024×768 or 800×600 (common sandbox defaults)
  • CPU core count of 1 or 2 (analysts economise on VM resources; real machines have 8–16)
  • Uptime less than 5 minutes (sandbox spins up, runs sample, tears down)

When malware detects these patterns it runs a clean code path — your analysis captures nothing meaningful, and you've spent time on a decoy.

OPSEC failure:   Run every sample in the same "MalwareAnalysis-Win10" snapshot
Safer:           Rotate VM names, MACs, usernames. Use 8 vCPUs. Pre-populate with fake documents,
                 browser cookies, and plausible user history. Use long-running snapshots (uptime > 7 days).
                 Tools: FlareVM deployed fresh per engagement, VMCloak for automated VM customisation.

Running malware on an analyst machine outside a sandbox

Accidentally double-clicking a sample outside an isolated VM results in a live C2 beacon from your real machine. The attacker now has your actual IP, username, hostname, and OS version. This happens during triage ("let me just check what this does quickly") more often than people admit.


Timing and Coordination Failures

Querying VT or sandbox at the same time the C2 beacon goes silent

If you find the implant, quarantine the host, and immediately query VT — all in the same 5-minute window — a competent attacker correlates:

  1. Beacon from victim IP stops → host was quarantined
  2. VT query for their hash arrives 3 minutes later → analyst found the file

They now know which victim, which host, and the investigation timeline. They may have other implants on the same network that they haven't burned yet.

Safer:           Complete network-level containment, scope the full compromise (lateral movement,
                 persistence mechanisms, all beaconing hosts) BEFORE doing any public lookups.
                 VT queries should happen after you've isolated everything and accepted the attacker
                 knows.

Blocking C2 at the firewall before mapping full scope

Blocking the known C2 IP at the perimeter is satisfying but premature. The moment the beacon dies, the attacker knows:

  • Which victim network lost connectivity to the C2
  • That detection has occurred
  • They may pivot to secondary implants, lateral movement targets, or exfil that was queued
OPSEC failure:   Add C2 IP to deny list in the firewall immediately upon discovery
Safer:           Monitor first. Let the beacon run (in isolation if possible) while you:
                   1. Map all hosts beaconing to the same C2
                   2. Identify any lateral movement or secondary persistence
                   3. Check for data staging / exfil activity
                 Block everything simultaneously once you have a complete picture.

Publishing IOCs before containment

Sharing indicators (IP, domain, hash, YARA rule) to a public MISP instance, Twitter, a threat intel ISAC, or a blog post while the attacker still has active access:

  • They see their IOCs published and immediately know the campaign is burned
  • They rotate infrastructure within hours
  • They may accelerate exfil or destruction before you can evict them
OPSEC failure:   Post "new Lazarus IOCs" to Twitter while still in active IR
Safer:           Share IOCs via TLP:RED or TLP:AMBER in closed-circle ISACs only.
                 Public disclosure only after full eviction or infrastructure has already decayed.

Phishing and Email Investigation

Forwarding phishing emails to analysis tools without stripping headers

Email headers contain a full routing trace from the victim's mail server — including your mail infrastructure's IPs and server names. Forwarding a phishing email to an external analysis service exposes this metadata to the service (and potentially to the attacker if the service is compromised or the attacker operates it).

Opening phishing emails or HTML attachments without isolation

HTML emails and attachments commonly contain tracking pixels — 1×1 transparent images at a unique URL. When your email client (or browser) loads the image, the attacker's web server logs:

  • Your IP address
  • User-Agent (reveals email client, browser, OS version)
  • Timestamp of when you opened it
OPSEC failure:   Open phishing.html in a browser to "see what it looks like"
                 Forward original .eml to an external email security vendor for analysis
Safer:           Open HTML files in a text editor or use `cat`/`strings` first.
                 For full rendering: use an isolated VM with no real network access (or sinkholed DNS).
                 Strip headers before forwarding: use `munpack`, Sublime Text, or mail header parsers
                 that don't make external network calls.

OSINT and Attribution Research

Performing OSINT on attacker personas from a corporate IP

Looking up attacker handles, email addresses, or usernames from LinkedIn, GitHub, Telegram, or forums from a corporate IP or on a logged-in account can:

  • Alert the attacker (honeytokens — some actors set canary links in their profiles)
  • Reveal your organisation's interest to the platform, which may notify the account
  • Tie your investigation to your identity if the attacker reviews their profile visitors
OPSEC failure:   Google "APT29 alias" from a corp laptop while logged into Google
                 Search an attacker's GitHub username from your work machine
Safer:           Use a sterile browser profile with no signed-in accounts.
                 Use a VPN or Tor exit node with no link to your organisation.
                 Use dedicated OSINT VMs (e.g., OSINT Kombine or a clean Tails session).

Attributing publicly before eviction

Publishing a full attribution report ("this is Fancy Bear") while the threat actor still has access to your environment or partner networks is a strategic failure. The actor reads the report, validates what you know and don't know, and adapts. Nation-state groups have cited published IR reports in their own internal retooling decisions.


Communication Channel Compromise

If the attacker has access to your email, Slack, Teams, or ticketing system (all common in persistent access scenarios), then any IR communication on those channels is visible to them in real time. They can read your containment plan, know which hosts are being reimaged, and move laterally to hosts not yet on your radar.

Safer:           Establish an out-of-band communication channel for IR:
                   - Separate Signal group with only IR responders (not on corp infrastructure)
                   - New Slack workspace outside the company org
                   - Phone calls / in-person for critical steps
                 Assume all internal systems are compromised until proven otherwise.

Summary: What to Do Before Any Public Lookup

1. Scope first            — identify ALL beaconing hosts, lateral movement, persistence mechanisms
2. Contain simultaneously — isolate everything at once, not incrementally
3. Out-of-band comms      — use a channel the attacker cannot read
4. Private intel only     — internal MISP, closed-feed threat intel, self-hosted sandbox
5. Accept the burn        — only then query VT, publish IOCs, notify ISAC

Interview Questions and Answers

Q
Walk me through how you would triage an unknown Linux binary in the first 5 minutes.
Model answer

file → sha256sum → entropy check → strings -n 8 | head -40 → readelf -d | grep NEEDED → nm -D | grep ' U ' → checksec. First establish the file type and hash for threat intel, then check entropy (>7.0 = likely packed, deal with unpacking before static analysis is useful). Quick string preview reveals obvious IOCs even without unpacking.


Q
What does the file command tell you, and what does "stripped" mean for analysis?
Model answer

file reads magic bytes and the ELF header to identify architecture (x86-64, ARM, MIPS), linking mode (dynamically/statically linked), and whether the binary is stripped. "Stripped" means the .symtab symbol table was removed — GDB and objdump see only hex addresses instead of function names, making analysis significantly harder. Dynamic symbols (.dynsym) are never stripped.


Q
How do you use strings to extract IOCs? What are its limitations?
Model answer

strings -n 8 binary | grep -iE 'https?://|[0-9]{1,3}\.[0-9]{1,3}|/etc/|/tmp/|password' extracts printable ASCII sequences that are at least 8 chars long. Critical limitation: any XOR-encrypted, RC4-encrypted, or otherwise obfuscated strings produce no useful output. Use dynamic analysis (strace/ltrace/Frida) to catch those at runtime when they're decrypted into memory before use.


Q
What is the difference between readelf -S and readelf -l?
Model answer

-S shows sections — the linker's logical view (.text, .data, .rodata, etc.) used by analysis tools. -l shows segments (program headers) — the kernel loader's view of how the file maps into memory at runtime, with permissions (R, W, X flags). Sections don't exist at runtime; segments do. A packed binary's -l may show an RWE LOAD segment — writable and executable — which is a major red flag.


Q
What does nm -D binary | grep ' U ' tell you, and why is it useful even on stripped binaries?
Model answer

It shows undefined (imported) dynamic symbols — functions the binary calls from shared libraries. These cannot be stripped because the dynamic linker needs them at runtime. Even a fully stripped binary with no function names exposes its capabilities here: connect + send + recv = network capable; execve = can run programs; ptrace = anti-debug or tracing; openat /etc/shadow pattern would appear in strace.


Q
Why should you never run ldd on an untrusted binary? What should you use instead?
Model answer

ldd works by setting LD_TRACE_LOADED_OBJECTS=1 and actually executing the binary. A malicious binary can detect this environment variable and run harmful code — C2 callbacks, persistence installation, or data destruction. Use readelf -d binary | grep NEEDED instead — it reads the .dynamic section directly without executing anything.


Q
How do you detect that a binary is packed before running it?
Model answer

Four signals: (1) high entropy — overall entropy >7.0 or .text section >7.2 (normal code is 4.5–6.0); (2) minimal import table — only VirtualAlloc, ExitProcess, LoadLibrary, GetProcAddress; (3) UPX magic strings — strings | grep UPX or readelf -S | grep UPX0; (4) tiny .text section relative to a large high-entropy .data or unnamed section — the stub is in .text, real code is compressed in .data.


Q
Explain the difference between strace and ltrace. When would you use each?
Model answer

strace intercepts at the kernel syscall boundary — open, connect, execve. Always works, unavoidable, works on statically linked binaries. ltrace intercepts at the libc function boundary — strcmp, malloc, puts, fopen. More readable (shows function names) but defeated by statically linked binaries (no shared lib calls to intercept). Use strace first for network/file/process activity; ltrace for string comparisons and config parsing to catch decrypted values.


Q
What strace output would indicate a reverse shell? A cryptominer? A credential stealer?
Model answer

Reverse shellconnect() to remote IP → dup2(fd, 0) / dup2(fd, 1) / dup2(fd, 2) (redirect stdin/stdout/stderr to socket) → execve("/bin/sh", ...). Miner: connect() to mining pool IP → tight alternating send()/recv() loop. Credential stealer: openat("/etc/shadow", O_RDONLY) or /proc/*/environ reads, or PAM library calls, followed by a network connect() to exfil.


Q
How would you use Frida to extract decrypted strings from a binary that XOR-encrypts them at runtime?
Model answer

Find the decrypt function address in Ghidra (look for XOR loop pattern). Then: Interceptor.attach(ptr("0xADDRESS"), { onEnter: function(args) { this.out = args[0]; this.len = parseInt(args[2]); }, onLeave: function() { console.log(Memory.readUtf8String(this.out, this.len)); } }). By the time onLeave fires, the output buffer contains plaintext.


Q
Write a YARA rule for a Linux ELF that connects to a hardcoded IP on port 4444 and spawns /bin/sh.
Model answer
yara
rule ELF_Reverse_Shell {   strings:     $elf  = { 7F 45 4C 46 }     $ip   = "1.2.3.4" ascii     $sh   = "/bin/sh" ascii     $exec = { B8 3B 00 00 00 0F 05 }  // execve syscall   condition:     $elf at 0 and $ip and $sh and $exec }

Key: uint32(0) == 0x464C457F or $elf at 0 anchors to the ELF magic. Real rule would combine port 4444 context with connect import.


Q
What does high entropy in a .text section tell you? What is the threshold?
Model answer

High entropy in .text means the "code" is actually compressed or encrypted data — you're looking at a packer stub, not real assembly. Normal x86-64 code entropy is 4.5–6.0 because instructions use a non-uniform byte distribution. Threshold: >7.0 is suspicious, >7.2 in .text is essentially certain to be packed. The real code lives elsewhere (usually .data) in a compressed blob.


Q
How would you unpack a UPX-packed binary using a debugger?
Model answer

(1) Break on mprotect or VirtualAlloc — the stub allocates executable memory for the real binary. Note the returned address. (2) Set a hardware execute breakpoint at that address. (3) Resume — the stub decompresses the real binary into the allocation, then jumps to it. (4) Hardware breakpoint fires at the OEP. (5) Dump memory with Scylla (x64dbg) or gdb's dump binary memory out.bin START END. Try upx -d first — only falls back to debugger if UPX headers were stripped.


Q
What is the difference between Ghidra and radare2? When would you reach for each?
Model answer

Ghidrafree, NSA-developed, excellent C decompiler (often better than IDA's), GUI with call graph, best for complex stripped binaries and collaborative analysis. radare2: CLI-first, fast startup, supports 100+ architectures, best for automation/scripting, embedded firmware (MIPS/ARM IoT), quick terminal-based analysis. For RE work on a campaign, Ghidra. For a quick "what does this do?" on an ARM router binary from a terminal, radare2.


Q
How do you detect an LD_PRELOAD-based rootkit using static analysis?
Model answer

Look for: (1) RTLD_NEXT in strings output — the classic pattern for hooking libc functions (get the real function pointer, then wrap it); (2) dlsym/dlopen imports; (3) hooks on readdir, getdents64, opendir (for file hiding) or getpwnam, pam_authenticate (for credential theft); (4) /etc/ld.so.preload path in strings (persistence mechanism). nm -D rootkit.so | grep 'readdir\|getdents\|fopen' reveals which libc functions are being replaced.