Malware Reverse Engineering — Linux Basics and Pattern Extraction
This file covers the practical workflow for analysing suspicious Linux binaries: triage, static analysis, pattern extraction, dynamic analysis, and decompilation. Everything here is executable from a Linux analysis VM with no special licences.
Breadth layernotes-security-core-knowledge.md Also see: EDR Evasion · Anti-Debugging · Digital Forensics
Why We Do This — The Analyst's Mindset
When you find a suspicious binary on a Linux system, you have one primary job: answer the five questions before the malware does more damage.
1. What IS this? → file type, architecture, linked libraries, packed or not?
2. What does it WANT? → network connections, files it reads/writes, commands it runs?
3. How does it PERSIST? → cron, systemd, rc.local, LD_PRELOAD, kernel module?
4. What did it STEAL? → credentials, SSH keys, environment variables, clipboard?
5. Where does it PHONE HOME? → C2 IPs, domains, beaconing interval, protocol?Each phase of analysis is designed to answer a different subset of those questions:
| Phase | Method | Answers |
|---|---|---|
| Triage | file, xxd, entropy, strings | What is it? Is it worth deeper analysis? |
| Static | readelf, nm, objdump, checksec | Architecture, capabilities, anti-analysis features |
| Pattern extraction | strings + grep, regex | IOCs — C2 addresses, persistence paths, credentials |
| Dynamic | strace, ltrace | What it actually does when running (bypasses obfuscation) |
| Instrumentation | GDB, Frida | How it works internally; extract decrypted data |
| Decompilation | Ghidra, Binary Ninja | Understand complex logic; find kill switches or config |
Static analysis vs dynamic analysis — why both?
(analysing without running): safe, no risk of infection or C2 contact, works even on encrypted binaries (you can see the structure). Limitation: if strings are XOR-encrypted, strings finds nothing.
(running the binary and watching): reveals exactly what happens, bypasses all obfuscation — encrypted strings are decrypted before use, so you see them. Limitation: requires isolation (never run malware on a non-isolated machine), and the malware may detect your analysis environment and not exhibit its real behaviour.
The workflow is not linearStatic findings inform what to look for in dynamic analysis. Dynamic findings reveal which functions to decompile. You loop between phases as you learn more.
Memory hookstatic is reading the recipe, dynamic is tasting the dish. Static analysis reads the binary without running it — safe, no infection risk, works on the structure even when you can't read the contents — but it's defeated by obfuscation (XOR-encrypted strings show nothing to
strings). Dynamic analysis runs it and watches — the malware must decrypt its strings and reveal its C2 before it can use them, so obfuscation melts away — but it requires isolation (never run malware on a connected box) and the malware may detect the lab and play dead (hence anti-debugging). They're complementary: static tells you where to look, dynamic shows you what actually happens. Mnemonic: static = safe but fooled by encryption; dynamic = truthful but needs a cage. And tie it back to the five questions — every command you run is in service of what is it / what does it want / how does it persist / what did it steal / where does it phone home, which is also the shape of the IOCs you hand to the detection team.
One rule above all elsenever run an untrusted binary on a machine connected to your production network or any network the C2 can reach. Use an isolated VM with an internal-only network, or a dedicated sandbox. (See also: Analyst OPSEC Failures — DNS queries from your sandbox can alert the attacker.)
Triage — First 60 Seconds
GoalIn 60 seconds, produce three outputs: (1) a hash for threat intel lookups, (2) a verdict on whether the binary is packed/encrypted, (3) enough string snippets to decide what kind of malware this might be. This determines whether you go straight to static analysis or need to unpack first.
Decision tree after triage
file → "ELF" + low entropy (<6.5) → proceed to static analysis
file → "ELF" + high entropy (>7.0) → likely packed → find OEP, dump, then static analysis
file → "data" (unknown) → binary blob, shellcode, or encrypted payload
file → "ASCII text" → script — read it directly
file → anything containing "UPX" → UPX packed → try `upx -d` firstBefore running anything, answer: what is this file? Is it packed? Is it network-capable? Only then decide how to proceed.
# 1. Identify the file type
file suspicious_binary
# → ELF 64-bit LSB executable, x86-64, dynamically linked, stripped
# → ELF 64-bit LSB executable, x86-64, statically linked, stripped
# → data ← unknown format, possibly packed/encrypted
# → ASCII text ← shell script, check shebang
# 2. Hash it for VirusTotal lookup
sha256sum suspicious_binary
md5sum suspicious_binary # legacy, but VT still accepts it
# Paste the hash at virustotal.com or use the API:
curl -s --request GET \
"https://www.virustotal.com/api/v3/files/<sha256>" \
--header "x-apikey: $VT_API_KEY" | jq '.data.attributes.last_analysis_stats'
# 3. Check for recognizable magic bytes
xxd suspicious_binary | head -4
# 7f 45 4c 46 → ELF
# 4d 5a → PE (Windows binary — unusual on Linux, check for Wine targets)
# 23 21 → shebang (#!)
# 1f 8b → gzip
# 50 4b → ZIP / JAR / APK
# 4. Entropy check (packed/encrypted = high entropy)
python3 -c "
import sys, math, collections
data = open(sys.argv[1], 'rb').read()
freq = collections.Counter(data)
entropy = -sum((c/len(data)) * math.log2(c/len(data)) for c in freq.values())
print(f'Overall entropy: {entropy:.4f} / 8.0')
print('Likely packed/encrypted' if entropy > 7.0 else 'Probably not packed')
" suspicious_binary
# 5. Quick string preview
strings -n 8 suspicious_binary | head -40Static Analysis Tools Reference
Goal of static analysisAnswer as many questions as possible without running the binary. You learn what the binary is capable of (imported functions), what hard-coded IOCs it contains (strings), how it was compiled (stripped/not, mitigations, architecture), and whether it's packed. All of this without ever giving the malware a chance to phone home.
Order of operations in static analysis
file→ type and basic propertiesreadelf -h→ architecture, entry point, type (EXEC vs DYN)readelf -S→ sections, find suspicious flags or namesreadelf -d→ imported librariesnm -D→ imported function names (capabilities)strings→ IOCs, config, stringschecksec→ security mitigations summarybinwalk -E→ entropy, detect packing
file — Type Identification
file binary
# Key fields:
# ELF 64-bit / 32-bit → architecture
# LSB / MSB → little-endian / big-endian
# executable / shared obj / relocatable → file role
# dynamically linked → will load .so files at runtime
# statically linked → everything compiled in — no external .so dependencies
# (uses shared libs) → older phrasing for dynamically linked
# stripped → symbol table removed — function names gone
# not stripped → debug symbols present — function names visible in gdb/objdumpstrings — Extract Printable Strings
strings scans for sequences of printable characters. It is the fastest way to extract IOCs, error messages, hard-coded configuration, and obfuscated fragments.
# Default: ASCII, minimum 4 chars
strings binary
# Increase minimum length to reduce noise (8 chars filters most junk)
strings -n 8 binary
# Include wide (UTF-16LE) strings — important for Windows malware cross-compiled for Linux
# or malware targeting Wine environments
strings -el binary # little-endian UTF-16
strings -eb binary # big-endian UTF-16
# Scan the entire file (not just sections — catches data in overlay/appended data)
strings -a binary
# Show file offset for each string (useful for Ghidra cross-reference)
strings -a -t x binary # hex offset
strings -a -t d binary # decimal offset
# Filter for specific patterns immediately
strings -n 8 binary | grep -iE 'https?://'
strings -n 8 binary | grep -iE '[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}'
strings -n 8 binary | grep -iE 'password|passwd|secret|token|apikey|api_key'
strings -n 8 binary | grep -iE '\.(sh|py|pl|php|rb)$'
strings -n 8 binary | grep -iE '/etc/|/tmp/|/var/|/proc/'
strings -n 8 binary | grep -iE 'execve|system|popen|fork|wget|curl'Pitfallstrings only finds printable strings. Encrypted or XOR-obfuscated strings produce no output. Use dynamic analysis or entropy analysis to detect them.
readelf — ELF Header and Sections
# ELF file header: architecture, entry point, class, endianness
readelf -h binary
# Section headers: name, type, size, offset, flags
readelf -S binary
# Important sections:
# .text → executable code
# .data → initialized global variables
# .rodata → read-only data (hard-coded strings, constants)
# .bss → uninitialized globals (zero at startup)
# .plt → Procedure Linkage Table (external function stubs)
# .got → Global Offset Table (resolved addresses of external functions)
# .dynamic → dynamic linking metadata
# .symtab → symbol table (may be absent in stripped binaries)
# .debug_* → debug information (DWARF)
# Program headers (segments): how OS maps the binary into memory
readelf -l binary
# Dynamic section: shared library dependencies, RPATH, symbol versions
readelf -d binary
# Look for NEEDED entries (imported .so files) and RPATH (unusual rpath = suspicious)
# Symbol table: function names (stripped = only external/dynamic symbols)
readelf -s binary
readelf --syms binary
# Relocation entries: which external functions are called and from where
readelf -r binary
# All headers in one shot
readelf -a binary | less
# Check for SUID bit and capabilities in binary metadata
readelf -n binary # notes section (may contain OS ABI, build ID)objdump — Disassembly
# Disassemble the .text section (code)
objdump -d binary
# Disassemble all sections (catches code in unusual sections)
objdump -D binary
# Use Intel syntax (more readable than AT&T default)
objdump -d -M intel binary
# Show source line information (only if not stripped)
objdump -d -S binary
# Dump a specific section as hex
objdump -s -j .rodata binary
# Show the full symbol table
objdump -t binary
# Show dynamic symbol table (imported functions)
objdump -T binary
# Find all CALL instructions (who calls what)
objdump -d binary | grep 'call\|jmp' | head -50nm — Symbol Table
# List all symbols
nm binary
# Show dynamic symbols only (what's imported from .so files)
nm -D binary
# Sort by address
nm -n binary
# Demangle C++ names
nm --demangle binary
# Check if stripped (no output or minimal output means stripped)
nm binary 2>&1 | head -5
# "no symbols" → stripped binary — function names unavailable in gdb without debug info
# Example output interpretation:
# 0000000000401234 T main → T = .text section, defined here
# 0000000000000000 U puts@@GLIBC_2.2.5 → U = undefined (imported from libc)
# 0000000000403000 D some_global → D = .data section (initialized global)
# 0000000000000000 w __gmon_start__ → w = weak symbolldd — Shared Library Dependencies
# Show all shared libraries the binary will load at runtime
ldd binary
# Example output:
# linux-vdso.so.1 (0x00007ffd3a1fc000)
# libpthread.so.0 → /lib/x86_64-linux-gnu/libpthread.so.0
# libc.so.6 → /lib/x86_64-linux-gnu/libc.so.6
# What the library list tells you:
# libssl.so / libcrypto.so → TLS/crypto capability
# libcurl.so → HTTP client capability
# libpthread.so → multithreading (evasion, process injection)
# libdl.so → dlopen() — dynamic library loading at runtime
# libpam.so → PAM auth hooks (credential theft)
# no interesting libs → statically linked or custom resolver
# NEVER run ldd on an untrusted binary on a real system.
# ldd works by setting LD_TRACE_LOADED_OBJECTS=1 and executing the binary.
# Use readelf -d instead for untrusted samples:
readelf -d binary | grep NEEDEDchecksec — Security Mitigation Audit
# Install: apt install checksec or pip install checksec
checksec --file binary
# Or using pwntools:
pwn checksec binary
# Example output:
# RELRO: Full RELRO ← GOT is read-only after startup
# Stack: Canary found ← stack smashing protector enabled
# NX: NX enabled ← non-executable stack
# PIE: PIE enabled ← position-independent executable (ASLR applies)
# RPATH: No RPATH ← no custom library search path
# Symbols: No Symbols ← stripped
# What to look for in malware:
# No PIE = fixed load address → easier to exploit or ROP
# No Canary = easier buffer overflow exploitation
# RPATH present = suspicious custom library search path (sideloading)binwalk — Embedded Files and Entropy
# Install: apt install binwalk
binwalk binary
# Detect embedded files (inside packers, droppers)
binwalk binary
# Outputs: offset, description
# 0x1234 gzip compressed data
# 0x5678 ELF 64-bit executable
# 0xABCD PNG image
# Extract embedded files
binwalk -e binary
# Creates a _binary.extracted/ directory
# Entropy analysis — visualize packed regions
binwalk -E binary
# High-entropy regions (near 8.0) = encrypted/compressed
# Transitions from low to high entropy = packer header/stub → payload boundary
# Recursive extraction (handles nested archives)
binwalk -e -M binaryELF Structure Analysis
GoalUnderstand the binary's layout before touching any code. The ELF format is a container — it tells you where the code is, what libraries it needs, what security mitigations are in place, and (if not stripped) what all the functions are named. Reading it correctly saves hours of disassembly.
The two views of an ELF binary:
ELF has two overlapping but different views of the same file. This confuses many analysts at first:
(linker's view): how the binary is divided logically during compilation — .text for code, .data for variables, .rodata for constants, etc. Used by the linker and by analysis tools like objdump and readelf -S. Not needed at runtime.
(loader's view): how the OS kernel maps the file into memory when the program runs — LOAD segments define which byte ranges get mapped and with what permissions. This is what the kernel actually reads. Each segment typically spans several sections.
Think of it like this: sections are the architect's floor plan; segments are the walls the builder actually constructs.
ELF Binary Layout (on disk): ┌─────────────────────────┐ offset 0x00 │ ELF Header (64 bytes) │ magic, type, arch, entry point address, offsets to tables ├─────────────────────────┤ offset 0x40 (64 bytes in) │ Program Headers │ segment table — read by kernel loader │ (Segment table) │ LOAD, DYNAMIC, INTERP, GNU_STACK, etc. ├─────────────────────────┤ │ .interp │ path to dynamic linker: /lib64/ld-linux-x86-64.so.2 │ .note.gnu.build-id │ build ID (unique hash of the binary — useful for correlation) │ .gnu.hash / .hash │ symbol hash tables for fast dynamic lookup │ .dynsym │ dynamic symbol table (imports — NEVER stripped) │ .dynstr │ string table for dynamic symbol names │ .gnu.version │ symbol version info (which glibc version needed) │ .rela.dyn / .rela.plt │ relocation entries (where to patch addresses at load time) │ .plt │ Procedure Linkage Table — stubs for external function calls │ .plt.got │ PLT stubs using GOT │ .text │ ← THE CODE. This is where functions live. │ .rodata │ read-only data: string literals, constant tables │ .eh_frame │ exception handling / stack unwinding data │ .data │ initialized global/static variables (read-write) │ .bss │ uninitialized globals (zero-filled by OS at load time, no disk space) │ .got │ Global Offset Table — resolved addresses of external functions │ .got.plt │ GOT entries for PLT lazy binding │ .dynamic │ dynamic linking metadata — library names, flags, addresses │ .symtab │ FULL symbol table (stripped in release builds) │ .strtab │ string table for .symtab names │ .shstrtab │ string table for section names themselves │ .debug_info │ DWARF debug info (present only in debug builds) ├─────────────────────────┤ │ Section Header Table │ at end of file — offset stored in ELF header └─────────────────────────┘
Reading the ELF File Header (readelf -h)
The ELF header is the first 64 bytes of every ELF file. It contains the metadata the kernel reads to decide how to load the program. Here is a real readelf -h output from a typical malicious ELF, with every field explained:
$ readelf -h suspicious.elf
ELF Header:
Magic: 7f 45 4c 46 02 01 01 00 00 00 00 00 00 00 00 00
↑↑ ↑↑↑↑↑↑↑↑↑ ↑↑ ↑↑ ↑↑ ↑↑ ↑↑
│ E L F │ │ │ └─ padding zeros (9 bytes)
│ │ │ └──── OS/ABI: 0x00 = System V (Linux uses this even though it's Linux ABI)
│ │ └─────── ELF version: always 1
│ └────────── Data encoding: 01 = little-endian (LSB), 02 = big-endian
└──────────────────────── Class: 02 = ELF64 (64-bit), 01 = ELF32
┌─ FOCUS: Class tells you 32-bit vs 64-bit. Affects all pointer sizes, struct offsets.
│ Malware targeting embedded devices or 32-bit systems shows 01 here.
└─────────────────────────────────────────────────────────────────────────────────────
Class: ELF64
Data: 2's complement, little endian
Version: 1 (current)
OS/ABI: UNIX - System V
↑ Nearly all Linux binaries say "UNIX - System V" even though
they're clearly Linux. Only a few use "Linux" ABI explicitly.
MUSL-compiled malware often shows UNIX - System V.
ABI Version: 0
Type: EXEC (Executable file)
↑ FOCUS: This field is critical.
┌─────────────────────────────────────────────────────────────
│ EXEC = standard executable with a fixed load address (no PIE)
│ DYN = position-independent executable (PIE enabled) OR shared library
│ Modern hardened binaries and most distro packages use DYN.
│ If this is malware compiled as DYN, it has PIE (ASLR applies).
│ If EXEC, it loads at a FIXED address every time — easier to ROP.
│ REL = relocatable object file (.o) — not a complete program
│ CORE = core dump — may contain memory of a crashed program
└─────────────────────────────────────────────────────────────
Machine: Advanced Micro Devices X86-64
↑ Architecture. Other values you'll see:
ARM (32-bit), AArch64 (64-bit ARM — phones, Raspberry Pi, AWS Graviton),
MIPS (routers/IoT), PowerPC, RISC-V.
Botnet malware often ships multi-arch builds.
Version: 0x1
Entry point address: 0x401080
↑ FOCUS: This is where execution begins after the dynamic linker
has finished loading. On a non-stripped binary this is usually
_start (which calls __libc_start_main which calls main).
On a PACKED binary the entry point is the packer stub, NOT main.
The OEP (Original Entry Point) is somewhere else, found after unpacking.
In Ghidra: press G → type 0x401080 to jump here immediately.
Start of program headers: 64 (bytes into file)
↑ Segment table starts right after the ELF header (at offset 64).
Always 64 for standard ELF64.
Start of section headers: 28672 (bytes into file)
↑ Section header table offset. In a stripped binary this might be 0
(section table stripped). The binary still runs — the kernel only
needs program headers, not section headers.
Flags: 0x0
↑ Architecture-specific flags. x86-64 always 0. ARM uses this for
Thumb mode, MIPS for ABI version, etc.
Size of this header: 64 (bytes)
Size of program headers: 56 (bytes) ← each segment descriptor is 56 bytes
Number of program headers: 11 ← 11 segments in this binary
Size of section headers: 64 (bytes)
Number of section headers: 28 ← 28 sections
Section header string table index: 27 ← section #27 contains the section namesWhat to extract immediately from the ELF header:
| Field | What it tells you |
|---|---|
Type: EXEC | Fixed load address — no ASLR on the binary itself → easier ROP chains |
Type: DYN | PIE — ASLR applies → need a leak to bypass |
Entry point: 0x4xxxxx | For EXEC: entry is in .text. Anything else is suspicious. |
Entry point: 0x0 | Malformed or position-independent — check program headers |
Machine: AArch64 | ARM 64-bit — IoT / mobile malware |
Machine: MIPS | Router malware (Mirai family) |
Number of sections: 0 | Stripped section table — objdump -S won't work, use program headers |
Reading the Section Headers (readelf -S)
Sections are the logical divisions of the file. readelf -S shows all of them with their type, offset, size, and permissions. Here is a real output with the security-relevant sections highlighted:
$ readelf -S --wide suspicious.elf
There are 28 section headers, starting at offset 0x7000:
Section Headers:
[Nr] Name Type Address Off Size ES Flg Lk Inf Al
[ 0] NULL 0000000000000000 000000 000000 00 0 0 0
[ 1] .interp PROGBITS 0000000000400238 000238 00001c 00 A 0 0 1
↑ FOCUS: Contains the path to the dynamic linker as a string.
Normal: /lib64/ld-linux-x86-64.so.2
Suspicious: a custom path → malware may use a custom loader / LD_PRELOAD trick
[ 2] .note.ABI-tag NOTE 0000000000400254 000254 000020 00 A 0 0 4
[ 3] .note.gnu.build-id NOTE 0000000000400274 000274 000024 00 A 0 0 4
↑ Build ID: a SHA1 hash of the binary computed at link time. Useful for correlation
across VirusTotal and other intel sources. Extract with: readelf -n binary
[ 4] .gnu.hash GNU_HASH 0000000000400298 000298 000030 00 A 5 0 8
[ 5] .dynsym DYNSYM 00000000004002c8 0002c8 0000f0 18 A 6 1 8
↑ FOCUS: Dynamic symbol table. These are the IMPORTED functions.
This section is NEVER stripped — it's required at runtime for dynamic linking.
`nm -D binary | grep ' U '` reads this. The imported function list tells you
what the binary CAN do: socket+connect = network, execve = exec, open+read = fs.
[ 6] .dynstr STRTAB 00000000004003b8 0003b8 000090 00 A 0 0 1
↑ String table for .dynsym names (function and library name strings)
[ 7] .gnu.version VERSYM 0000000000400448 000448 000014 02 A 5 0 2
[ 8] .gnu.version_r VERNEED 0000000000400460 000460 000030 00 A 6 1 8
↑ Which version of glibc each function requires. "GLIBC_2.17" = compatible with
anything from 2017+. Useful for dating when the binary was built.
[ 9] .rela.dyn RELA 0000000000400490 000490 000018 18 A 5 0 8
[10] .rela.plt RELA 00000000004004a8 0004a8 0000c0 18 AI 5 23 8
↑ Relocation tables. Tell the dynamic linker where to patch function addresses.
.rela.plt entries correspond 1:1 with PLT stubs — each entry is an imported function.
[11] .init PROGBITS 0000000000400568 000568 00001a 00 AX 0 0 4
↑ Code that runs BEFORE main() — constructor functions. Rootkits may hide here.
[12] .plt PROGBITS 0000000000400590 000590 0000a0 10 AX 0 0 16
↑ Procedure Linkage Table. For each imported function (connect, execve, etc.)
there is a small stub here. First call → resolves address via dynamic linker,
patches .got.plt. Subsequent calls → jump directly. Seeing these stubs in
disassembly is normal. Their names come from .dynsym.
[13] .plt.got PROGBITS 0000000000400630 000630 000008 08 AX 0 0 8
[14] .text PROGBITS 0000000000400640 000640 0003d2 00 AX 0 0 16
↑ FOCUS: The actual code. Flags "AX" = Allocated (in memory) + eXecutable.
Normal: this is the ONLY section with X flag.
Suspicious: .data or .bss also has X flag → shellcode / self-modifying code.
Suspicious: .text has very high entropy (>7.0) → packed/encrypted code stub.
Suspicious: .text is tiny (<0x100 bytes) → likely a packer stub (real code elsewhere).
Size matters: compare .text size against total file size. If .text is 5% of the file
and the rest is high-entropy .data → the real code is compressed in .data.
[15] .fini PROGBITS 0000000000400a14 000a14 000009 00 AX 0 0 4
↑ Code that runs AFTER main() returns — destructor functions.
[16] .rodata PROGBITS 0000000000400a20 000a20 000048 00 A 0 0 4
↑ Read-only data. String literals live here on non-stripped, non-obfuscated binaries.
`strings` on the binary extracts these. If .rodata is tiny or encrypted → obfuscation.
[17] .eh_frame_hdr PROGBITS 0000000000400a68 000a68 000024 00 A 0 0 4
[18] .eh_frame PROGBITS 0000000000400a90 000a90 000070 00 A 0 0 8
↑ Exception handling frame data. Required for C++ exceptions and backtraces.
Present even in C code compiled with recent GCC. Can be used by analysts to
find function boundaries in stripped binaries.
[19] .init_array INIT_ARRAY 0000000000601df8 001df8 000008 08 WA 0 0 8
↑ Array of constructor function pointers — run before main(). TLS callbacks live here
in ELF (the Linux equivalent of the Windows TLS callback anti-debug technique).
`readelf -S | grep init_array` then `readelf -x 19 binary` to dump the addresses.
[20] .fini_array FINI_ARRAY 0000000000601e00 001e00 000008 08 WA 0 0 8
[21] .dynamic DYNAMIC 0000000000601e08 001e08 0001d0 10 WA 6 0 8
↑ FOCUS: Dynamic linking metadata. Contains: library names (DT_NEEDED),
address of symbol table (DT_SYMTAB), GOT address (DT_PLTGOT), init/fini addresses.
Malware packed as a shared library often modifies DT_INIT to point to the unpacker stub.
`readelf -d binary` reads this section. Key tags:
DT_NEEDED = shared library dependency (one per imported .so)
DT_RPATH / DT_RUNPATH = custom library search path (suspicious if non-standard)
DT_INIT = address of _init() function (runs before main)
DT_FINI = address of _fini() function (runs after main)
[22] .got PROGBITS 0000000000601fd8 001fd8 000010 08 WA 0 0 8
[23] .got.plt PROGBITS 0000000000601fe8 001fe8 000060 08 WA 0 0 8
↑ FOCUS: Global Offset Table. Contains the resolved addresses of imported functions.
After dynamic linking, connect@got.plt contains the real address of connect() in libc.
GOT overwrite is a classic exploitation technique — overwrite a GOT entry to redirect
a function call. FULL RELRO (from checksec) makes .got.plt read-only after load,
preventing this attack. PARTIAL RELRO only protects .got, not .got.plt.
[24] .data PROGBITS 0000000000602048 002048 000010 00 WA 0 0 8
↑ Initialized global/static variables. Encrypted config blobs often live here.
A large .data section with high entropy = stored encrypted payload.
[25] .bss NOBITS 0000000000602058 002058 000008 00 WA 0 0 8
↑ Uninitialized globals. Not stored on disk (NOBITS = no disk space).
Zero-filled by OS at load time.
[26] .symtab SYMTAB 0000000000000000 003058 000660 18 27 49 8
↑ FOCUS: Full symbol table — present only in NON-STRIPPED binaries.
Contains ALL function and variable names, including internal ones.
Stripped binaries: this section is absent (or size 0).
If present: gold mine for analysis — every function is named.
`nm binary` reads this. `nm -D binary` reads .dynsym instead.
[27] .strtab STRTAB 0000000000000000 0036b8 000310 00 0 0 1
[28] .shstrtab STRTAB 0000000000000000 0039c8 000102 00 0 0 1Section flags decoded
| Flag | Meaning | Security relevance |
|---|---|---|
A | Alloc — section is loaded into memory at runtime | Sections without A are debug/metadata only |
X | eXecute — memory region is executable | ONLY .text should have this; extras = shellcode |
W | Write — memory region is writable | .data/.bss/.got are W; .text being W = self-modifying code |
I | Info link | Internal ELF bookkeeping |
M | Merge — duplicate entries can be removed | |
S | Strings — contains null-terminated strings |
Red flags in section headers
# Any section that is both Writable AND eXecutable
readelf -S binary | grep -E 'WX|AXW'
# → indicates self-modifying code or injected shellcode region
# .text section has extremely high entropy
# Indicates the packer stub is there but real code is encrypted inside
# Section named .text but type is not PROGBITS
readelf -S binary | awk '/\.text/{print}'
# Missing .dynamic section
readelf -S binary | grep -c DYNAMIC
# → 0 means statically linked (common in malware for maximum portability)
# Suspicious section names
readelf -S binary | grep -iE 'UPX|vmp|enigma|themida|packed'Reading the Program Headers (readelf -l)
Program headers describe segments — how the kernel actually maps the binary into memory. Here is a real readelf -l output annotated:
$ readelf -l --wide suspicious.elf
Elf file type is EXEC (Executable file)
Entry point 0x401080
There are 11 program headers, starting at offset 64
Program Headers:
Type Offset VirtAddr PhysAddr FileSiz MemSiz Flg Align
PHDR 0x000040 0x0000000000400040 0x0000000000400040 0x000268 0x000268 R 0x8
↑ PHDR: The program header table itself. Self-referential.
INTERP 0x000238 0x0000000000400238 0x0000000000400238 0x00001c 0x00001c R 0x1
↑ FOCUS: Points to the dynamic linker path (reads the .interp section).
Normal: /lib64/ld-linux-x86-64.so.2
If this is missing: statically linked binary (no dynamic loader needed).
Contents: readelf --string-dump=.interp binary
LOAD 0x000000 0x0000000000400000 0x0000000000400000 0x000b00 0x000b00 R E 0x200000
↑ FOCUS: LOAD segments are what the kernel actually maps into memory.
This one: Flags = "R E" = Read + Execute. This is the CODE segment.
Contains: .text, .rodata, .plt, .init, .fini
VirtAddr 0x400000 = the base load address for EXEC binaries (fixed).
FileSiz = MemSiz = 0x000b00: no extra zeroed memory needed.
LOAD 0x001df8 0x0000000000601df8 0x0000000000601df8 0x000268 0x000270 RW 0x200000
↑ FOCUS: Second LOAD segment. Flags = "RW" = Read + Write. This is the DATA segment.
Contains: .data, .bss, .got, .got.plt, .dynamic
MemSiz (0x000270) > FileSiz (0x000268) by 8 bytes → those 8 bytes are .bss (zeroed by OS).
No X flag here = Writable data is not executable. GOOD.
If you see a LOAD segment with "RWE" (Read+Write+Execute) → VERY SUSPICIOUS.
Self-modifying code or shellcode that writes and executes itself lives in RWE regions.
DYNAMIC 0x001e08 0x0000000000601e08 0x0000000000601e08 0x0001d0 0x0001d0 RW 0x8
↑ Points to the .dynamic section (dynamic linker metadata).
NOTE 0x000254 0x0000000000400254 0x0000000000400254 0x000044 0x000044 R 0x4
↑ Build notes (.note.ABI-tag, .note.gnu.build-id).
GNU_EH_FRAME 0x000a68 0x0000000000400a68 0x0000000000400a68 0x000024 0x000024 R 0x4
↑ Exception handling frame info. Used by debuggers and libgcc for stack unwinding.
GNU_STACK 0x000000 0x0000000000000000 0x0000000000000000 0x000000 0x000000 RW 0x10
↑ FOCUS: Stack permissions. Flags "RW" = Read + Write only (no execute).
This is correct — the stack should NEVER be executable (NX/DEP protection).
If you see "RWE" here → NX is disabled → stack-based shellcode can execute.
Equivalent to checksec showing "NX: disabled".
GNU_RELRO 0x001df8 0x0000000000601df8 0x0000000000601df8 0x000208 0x000208 R 0x1
↑ FOCUS: RELRO (Relocation Read-Only) marking. After dynamic linking finishes,
the memory range in this segment is made read-only (mprotect to R only).
PARTIAL RELRO: only .got (non-PLT) is in this segment — .got.plt remains writable.
FULL RELRO: both .got and .got.plt are in this segment → GOT is fully read-only.
Absence of GNU_RELRO → no RELRO → GOT is always writable → trivial GOT overwrite.
Section to Segment mapping:
Segment Sections...
00
01 .interp
02 .interp .note.ABI-tag .note.gnu.build-id .gnu.hash .dynsym .dynstr
.gnu.version .gnu.version_r .rela.dyn .rela.plt .init .plt .plt.got
.text .fini .rodata .eh_frame_hdr .eh_frame
03 .init_array .fini_array .dynamic .got .got.plt .data .bss
04 .dynamic
05 .note.ABI-tag .note.gnu.build-id
06 .eh_frame_hdr
07
08 .init_array .fini_array .dynamic .gotQuick checklist from program headers
# 1. Any RWE LOAD segment? (writable AND executable)
readelf -l binary | grep 'LOAD' | grep 'RWE\| RW.*E\| R.*WE'
# If yes: shellcode / unpacker stub — serious red flag
# 2. Stack executable?
readelf -l binary | grep 'GNU_STACK'
# Should show: RW not RWE
# RWE = NX disabled = checksec "NX: disabled"
# 3. RELRO present?
readelf -l binary | grep 'GNU_RELRO'
# Absent = no RELRO = GOT overwrite possible
# 4. Entry point inside the expected LOAD segment?
readelf -h binary | grep 'Entry'
readelf -l binary | grep 'LOAD'
# Entry should be inside the R E LOAD segment (code), not the RW segment (data)
# Entry in a data segment → unpacker stub running from a data region → packedStripped vs Non-Stripped
# Non-stripped: rich symbol table — gdb shows function names
readelf -s binary | grep -c FUNC # dozens or hundreds of lines
nm binary | head -20
# Stripped: only dynamic symbols remain
file binary
# → stripped
nm binary
# → nm: binary: no symbols
# Dynamic symbols (imported functions) are never stripped — they are required at runtime
nm -D binary | grep ' U ' # undefined = imported from .so
# Example: printf@@GLIBC_2.2.5, socket@@GLIBC_2.2.5, connect@@GLIBC_2.2.5Key insightEven stripped binaries expose their imported function names via dynamic symbols. socket, connect, bind, recv, send → network capability. popen, execve, system → command execution. open, read, write, unlink → filesystem operations.
Sections That Should Not Exist
readelf -S binary | grep -E 'Name|Flags'
# Suspicious section names:
# .UPX0, .UPX1, .UPX2 → UPX packer
# .packed → custom packer
# .enigma, .vmp → VMProtect / Enigma
# No section names at all → heavily stripped or custom format
# Missing sections that should exist:
# No .text section → code is in an unnamed section (anti-analysis)
# No .dynamic section → statically linked (common in malware for portability)Pattern Extraction — Strings and IOCs
Network Indicators
# URLs and URIs
strings -n 8 binary | grep -oE 'https?://[a-zA-Z0-9./_?=&%-]+'
strings -n 8 binary | grep -oE 'ftp://[a-zA-Z0-9./_?=&%-]+'
# IP addresses (IPv4)
strings -n 7 binary | grep -oE '\b([0-9]{1,3}\.){3}[0-9]{1,3}\b' \
| grep -v '^127\.' | grep -v '^0\.'
# IPv6
strings binary | grep -oE '([0-9a-fA-F]{1,4}:){7}[0-9a-fA-F]{1,4}'
# Domain names (catch C2 domains)
strings -n 8 binary | grep -oE '[a-zA-Z0-9.-]+\.(com|net|org|io|ru|cn|cc|tk|top|xyz|info)'
# Port numbers in context (look for socket configuration)
strings binary | grep -oE ':[0-9]{2,5}\b'
# User-Agent strings (HTTP C2 often has a hardcoded UA)
strings binary | grep -i 'User-Agent\|Mozilla\|curl\|python-requests'Credentials and Secrets
# Hardcoded passwords and keys
strings -n 8 binary | grep -iE 'password|passwd|passw|secret|apikey|api_key|access_key|auth_token'
# Base64-encoded blobs (potential embedded keys or second-stage payloads)
strings binary | grep -oE '[A-Za-z0-9+/]{20,}={0,2}' | while read b64; do
echo "$b64" | base64 -d 2>/dev/null | strings -n 4
done
# SSH key headers
strings binary | grep -E 'BEGIN (RSA|EC|OPENSSH|PRIVATE|PUBLIC) KEY'
# AWS credentials pattern
strings binary | grep -oE 'AKIA[A-Z0-9]{16}'
strings binary | grep -oE '[a-zA-Z0-9/+]{40}'Filesystem Indicators
# Persistence paths
strings binary | grep -E '/etc/(cron|rc|init|systemd|profile|bash|passwd|shadow|ld\.so)'
strings binary | grep -E '(~|/home|/root)/\.(bashrc|profile|ssh|config|local)'
strings binary | grep -E '/var/(spool|run|tmp|log)'
strings binary | grep -E '\.service$|\.timer$|\.socket$' # systemd unit persistence
# Suspicious temp paths
strings binary | grep -E '/tmp/\.|/dev/shm/\.' # hidden files in temp directories
# /proc filesystem access (rootkit, credential theft)
strings binary | grep -E '/proc/[0-9]*/|/proc/self/'
strings binary | grep -E '/proc/net/|/proc/sys/'Command Execution Patterns
# Shell commands embedded in binary
strings binary | grep -E '^(bash|sh|zsh|python|perl|ruby|php|node) '
strings binary | grep -E 'execve|system\(|popen\(|fork\(\)'
strings binary | grep -E 'chmod [0-7]{3,4}|chown root|setuid|setgid'
# Reverse shell indicators
strings binary | grep -E '/dev/tcp/|/dev/udp/'
strings binary | grep -E 'bash -i|nc -e|ncat|socat'
strings binary | grep -iE 'reverse.?shell|bind.?shell'
# Privilege escalation
strings binary | grep -E 'sudo|su -|NOPASSWD|visudo'
strings binary | grep -E 'LD_PRELOAD|LD_LIBRARY_PATH'
# C2 command keywords (common RAT command strings)
strings binary | grep -iE '\b(shell|exec|upload|download|screenshot|keylog|persist)\b'Crypto and Ransom Indicators
# Ransomware patterns
strings binary | grep -iE 'encrypt|decrypt|AES|RSA|ransom|bitcoin|wallet|tor'
strings binary | grep -iE '\.encrypted$|\.locked$|\.crypto$'
strings binary | grep -iE 'README|HELP|HOW.TO.DECRYPT|RECOVERY'
# Cryptocurrency miner patterns
strings binary | grep -iE 'stratum\+tcp|mining.pool|xmrig|monero|hashrate'
strings binary | grep -iE 'nicehash|pool\.minexmr\.com|cryptonight'Obfuscation and Packing Detection
High-Entropy Detection — Per Section
python3 << 'EOF'
import pefile, math, sys
def section_entropy(data):
if not data: return 0
freq = {}
for b in data: freq[b] = freq.get(b, 0) + 1
return -sum((c/len(data)) * math.log2(c/len(data)) for c in freq.values())
try:
pe = pefile.PE(sys.argv[1])
print(f"{'Section':<20} {'Size':>10} {'Entropy':>8} Verdict")
print("-" * 55)
for s in pe.sections:
data = s.get_data()
ent = section_entropy(data)
verdict = "PACKED/ENCRYPTED" if ent > 7.0 else ("suspicious" if ent > 6.5 else "ok")
print(f"{s.Name.decode().strip():<20} {len(data):>10} {ent:>7.3f} {verdict}")
except Exception as e:
# For ELF, use a different approach
import struct
with open(sys.argv[1], 'rb') as f:
data = f.read()
chunk_size = 4096
print(f"{'Offset':<12} {'Entropy':>8} Verdict")
for i in range(0, len(data), chunk_size):
chunk = data[i:i+chunk_size]
ent = section_entropy(chunk)
if ent > 6.5:
print(f"0x{i:<10x} {ent:>7.3f} {'HIGH ENTROPY'}")
EOF suspicious_binaryUPX Detection and Unpacking
UPX is the most common open-source packer. Detecting and unpacking it is straightforward:
# Detection
strings binary | grep -iE 'UPX[0-9!]'
readelf -S binary | grep -E 'UPX[0-9]'
upx -t binary # test if valid UPX
# Unpack in place
upx -d -o binary_unpacked binary
# If UPX headers are stripped (common anti-analysis measure):
# The attacker removed "UPX0"/"UPX1" magic strings to prevent automatic unpacking
# Solution: run the binary in a debugger to the OEP, then dump memory
# x64dbg: run to entry, then Plugins → Scylla → Dump + Fix IATIdentifying the Real Entry Point (OEP)
When a packed binary runs, the stub decrypts the payload and jumps to the Original Entry Point (OEP). Finding it lets you dump the unpacked binary.
# Strategy 1: Look for a "tail jump" (jmp to OEP after unpacking loop)
objdump -d packed_binary | tail -30
# Common pattern: loop ending with jmp [eax] or jmp [rax] — that's the OEP jump
# Strategy 2: hardware breakpoint on VirtualAlloc return
# The stub allocates memory, writes the real binary, then jumps to it.
# Break when VirtualAlloc returns, note the allocation address.
# Set a hardware execute breakpoint at that address.
# Run → hit the breakpoint at the OEP.Dynamic Analysis — strace and ltrace
GoalObserve what the malware actually does when running — network connections, files created, processes spawned, credentials read. Dynamic analysis bypasses all static obfuscation: XOR-encrypted strings are decrypted before use and appear in function arguments you can observe; packed code is unpacked into memory and executes normally.
Why dynamic is essential alongside staticA binary with zero interesting strings in static analysis may have a full C2 URL and command list that appears only at runtime after decryption. strace captures every system call — it doesn't matter how obfuscated the code is, because system calls are the boundary between userspace and kernel and cannot be avoided.
The isolation requirementdynamic analysis means running the malware, so it must happen somewhere it can't hurt you:
Host-only or no networking, or a sinkholed network where every DNS name resolves to 127.0.0.1, so C2 calls go nowhere.
Take one before every run and revert afterwards. Never reuse a "dirty" VM.
Nothing mounted between the VM and the host — a shared folder is a path out of the sandbox.
Different MAC address and hostname each run: malware fingerprints analysis environments and goes quiet.
strace — System Call Tracer
# Trace all syscalls with timestamps
strace -tt ./binary
# Follow child processes (forked daemons, shell spawns)
strace -f ./binary 2>&1 | tee strace.log
# Trace a specific category of syscalls
strace -e trace=network ./binary # socket, connect, bind, recv, send
strace -e trace=file ./binary # open, read, write, stat, unlink
strace -e trace=process ./binary # execve, fork, clone, exit
strace -e trace=signal ./binary # signal, sigaction
strace -e trace=ipc ./binary # shared memory, semaphores
# Count syscalls (profiling)
strace -c ./binary
# Attach to a running process
strace -p <pid>
strace -fp <pid> # follow children tooKey patterns to look for in strace output:
# C2 connection attempt
connect(3, {sa_family=AF_INET, sin_port=htons(4444), sin_addr=inet_addr("1.2.3.4")}, 16)
# File creation (dropper)
openat(AT_FDCWD, "/tmp/.cache/.x", O_WRONLY|O_CREAT|O_TRUNC, 0755)
write(3, "\x7fELF\x02\x01\x01\x00", 64) # writing a new ELF to disk
# Privilege escalation attempt
setuid(0) → EPERM (failed — not running as root yet)
# or
openat(AT_FDCWD, "/etc/sudoers", O_RDONLY) = -1 EACCES
# Credential access
openat(AT_FDCWD, "/etc/shadow", O_RDONLY) = -1 EACCES
openat(AT_FDCWD, "/proc/1/environ", O_RDONLY) = -1 EPERM
# Shell spawn
execve("/bin/bash", ["/bin/bash", "-i"], 0x7fff... /* 23 vars */)
# Persistence (systemd service or cron)
openat(AT_FDCWD, "/etc/systemd/system/malware.service", O_WRONLY|O_CREAT)
# Network listen (bind shell or local port for lateral movement)
bind(3, {sa_family=AF_INET, sin_port=htons(1337), sin_addr=inet_addr("0.0.0.0")}, 16)
listen(3, 5)
# Anti-debug check
ptrace(PTRACE_TRACEME, 0, NULL, NULL) = -1 EPERM ← already traced → self-detects
openat(AT_FDCWD, "/proc/self/status", O_RDONLY) ← reading TracerPidltrace — Library Call Tracer
ltrace intercepts calls to shared libraries. Unlike strace (kernel boundary), ltrace works at the libc boundary — you see printf, malloc, strcmp, connect etc.
# Basic trace
ltrace ./binary
# Follow child processes
ltrace -f ./binary
# Filter for specific functions
ltrace -e 'strcmp+strcpy+strcat+malloc+free+connect' ./binary
# Show timestamps
ltrace -tt ./binary
# Attach to running process
ltrace -p <pid>Key patterns in ltrace output
# String comparison (often reveals decrypted passwords being checked)
strcmp("admin", "admin") = 0 # match
strcmp("correct_password", "wrongpass") = non-zero
# Memory operations (can reveal decrypted strings mid-execution)
malloc(256) = 0x55a1b2c3
# ... binary copies decrypted string here ...
puts("https://c2.evil.com/beacon") # now visible in plaintext
# File operations
fopen("/etc/passwd", "r") = 0x55a1...
fread(buf, 1, 4096, 0x55a1...) = 1024 # reading credentials
# Network operations (higher level than strace's connect)
getaddrinfo("c2.evil.com", "443", ...) # DNS resolution
connect(sockfd, {AF_INET, "1.2.3.4", 443}, ...)
# Crypto operations (if using OpenSSL / libcrypto)
EVP_EncryptInit_ex(ctx, EVP_aes_256_cbc(), ...) # AES encryption startingltrace vs strace trade-offstrace is always available and harder to evade (syscall level). ltrace requires debug symbol hooks and can be defeated by statically-linked binaries (no shared lib calls to intercept). Use both.
GDB — Debugging Basics
GoalStep through the binary's execution at instruction level, inspect and modify memory and registers in real time, extract decrypted data, and bypass anti-debug checks. GDB is what you reach for when strace/ltrace tells you what happened but not how — or when you need to intercept a decrypt function and read the plaintext before it gets used.
When to use GDB over strace
- You need to see inside a function call (strace only shows the syscall boundary)
- You want to extract a decrypted string or key from a register or memory region
- The binary is packed and you need to find the OEP (set a breakpoint on
mprotectormmap) - You need to bypass an anti-debug check (patch
ptracereturn value, skip a JNZ)
# Start debugging
gdb ./binary
gdb -q ./binary # quiet mode (no banner)
# Attach to running process
gdb -p <pid>
# Better experience: install pwndbg or gdb-peda
pip install pwndbg # then: echo "source /path/to/pwndbg/gdbinit.py" >> ~/.gdbinitEssential GDB Commands
# Execution control
run [args] → start the program
run < input_file → start with stdin redirected
start → run and break at main()
continue / c → resume execution
next / n → step over (don't enter function calls)
step / s → step into function calls
finish → run until current function returns
until <line> → run until source line
# Breakpoints
break main → break at function main
break *0x401234 → break at absolute address
break *main+0x20 → break 32 bytes into main
info breakpoints → list all breakpoints
delete 1 → delete breakpoint #1
disable 1 → disable without deleting
condition 1 rax==0 → conditional breakpoint
# Hardware breakpoints (no INT 3 — evades code-checksum anti-debug)
hbreak *0x401234 → hardware execute breakpoint
watch *0x603018 → hardware write watchpoint (fires when address written)
rwatch *0x603018 → hardware read watchpoint
# Examining memory
x/10i $rip → disassemble 10 instructions from current RIP
x/20x $rsp → dump 20 hex words from stack pointer
x/s 0x402000 → print string at address
x/b *0x401234 → read single byte
print $rax → print register value
info registers → show all registers
info registers rax rbx rsp rip → specific registers
# Memory search
find 0x400000, 0x500000, "password" → search for string in range
find /b 0x400000, 0x500000, 0x90, 0x90 → search for NOP NOP bytes
# Modify execution
set $rax = 0 → overwrite register
set *0x601020 = 0 → overwrite memory
set {int}0x601020 = 1337
# Shared libraries and symbols
info sharedlibrary → loaded .so files and their address ranges
info functions → list all known function symbols
info symbols 0x401234 → find nearest symbol to address
# Backtrace (call stack)
backtrace / bt → show call stack
frame 2 → switch to frame 2GDB Scripting for Anti-Debug Bypass
# gdb_bypass.py — auto-bypass ptrace PTRACE_TRACEME
# Run: gdb -x gdb_bypass.py ./binary
import gdb
class PtraceBypass(gdb.Breakpoint):
def stop(self):
# When ptrace is called, check if it's PTRACE_TRACEME (arg0 == 0)
ptrace_request = int(gdb.parse_and_eval("$rdi"))
if ptrace_request == 0: # PTRACE_TRACEME
gdb.execute("set $rax = 0") # fake success
gdb.execute("return") # skip the actual syscall
print("[*] ptrace(PTRACE_TRACEME) intercepted — returning 0")
return False # don't stop
PtraceBypass("ptrace", internal=False)
gdb.execute("run")Frida — Dynamic Instrumentation
Frida injects a JavaScript engine into any process without recompilation. Use it to trace function calls, hook specific functions, and extract data from memory at runtime.
# Install
pip install frida frida-tools
# List running processes
frida-ps
# Attach and run a script
frida -l script.js <process_name_or_pid>
# Spawn a new process under Frida
frida -l script.js -f ./binary --no-pause
# Trace all calls to a specific function
frida-trace -i "malloc" ./binary
frida-trace -i "connect" -i "send" -i "recv" ./binary
frida-trace -i "fopen" -i "fread" ./binary
# Trace all calls to any function matching a pattern
frida-trace -I "libc*" ./binary # trace everything in libcFrida Script Examples
// Hook connect() to log all network connections
Interceptor.attach(Module.getExportByName(null, "connect"), {
onEnter: function(args) {
// args[1] = sockaddr struct
var sa_family = Memory.readU16(args[1]);
if (sa_family == 2) { // AF_INET
var port = Memory.readU16(args[1].add(2));
port = ((port & 0xFF) << 8) | ((port >> 8) & 0xFF); // ntohs
var ip_bytes = Memory.readByteArray(args[1].add(4), 4);
var ip = Array.from(new Uint8Array(ip_bytes)).join(".");
console.log("[connect] " + ip + ":" + port);
}
}
});// Hook fopen to log all file opens
Interceptor.attach(Module.getExportByName(null, "fopen"), {
onEnter: function(args) {
var path = Memory.readUtf8String(args[0]);
var mode = Memory.readUtf8String(args[1]);
console.log("[fopen] " + path + " (" + mode + ")");
}
});// Intercept and log XOR decryption to extract decrypted strings
// Assumes decrypt_string(char* output, const char* input, size_t len, char key)
var decrypt_addr = ptr("0x401234"); // address from Ghidra/objdump
Interceptor.attach(decrypt_addr, {
onEnter: function(args) {
this.output_ptr = args[0];
this.len = parseInt(args[2]);
},
onLeave: function(retval) {
// After function returns, read the decrypted string
var decrypted = Memory.readUtf8String(this.output_ptr, this.len);
console.log("[decrypt_string] → " + decrypted);
}
});// Universal: dump all string arguments to any function (spray-and-pray)
["strcmp", "strcpy", "strcat", "puts", "printf"].forEach(function(fn) {
var addr = Module.getExportByName(null, fn);
if (addr) {
Interceptor.attach(addr, {
onEnter: function(args) {
try {
var s = Memory.readUtf8String(args[0]);
if (s && s.length > 3) console.log("[" + fn + "] " + s);
} catch(e) {}
}
});
}
});YARA — Writing Detection Rules
YARA is a pattern-matching engine for identifying malware. Rules describe byte patterns, string patterns, and structural conditions.
# Install
apt install yara
pip install yara-python
# Scan a file
yara rule.yar suspicious_binary
# Scan a directory recursively
yara -r rules/ /path/to/samples/
# Scan running processes
yara -p 50 rule.yar # uses /proc/<pid>/memWriting a Rule from Analysis
rule Suspicious_ELF_C2_Beacon {
meta:
description = "ELF binary with hardcoded C2 IP and XOR string obfuscation pattern"
author = "analyst"
severity = "high"
strings:
// Network indicators found via strings
$c2_ip = "192.168.10.100" ascii
$c2_path = "/beacon" ascii
$ua = "Mozilla/5.0 (compatible; malbot/1.0)" ascii
// XOR decryption loop pattern (x86-64 bytes)
// lea rsi, [rip+offset] ; xor [rsi+rcx], al ; loop
$xor_loop = { 48 8D 35 ?? ?? ?? ?? 30 04 0E 48 FF C1 }
// Anti-debug: ptrace PTRACE_TRACEME check
$ptrace_check = { B8 65 00 00 00 0F 05 48 85 C0 } // syscall 101 = ptrace
// Persistence path
$persist = "/etc/cron.d/updater" ascii
$persist2 = "/tmp/.x" ascii nocase
// Interesting imported symbols
$import_execve = "execve" ascii
$import_conn = "connect" ascii
condition:
// Must be an ELF binary
uint32(0) == 0x464C457F
// At least 2 network or execution indicators
and 2 of ($c2_ip, $c2_path, $ua, $import_execve, $import_conn)
// Anti-debug or xor loop
and ($ptrace_check or $xor_loop)
// And at least one persistence indicator
and 1 of ($persist, $persist2)
}Rule for Cryptominer Detection
rule Cryptominer_XMRig_Pattern {
meta:
description = "XMRig or XMRig-compatible miner"
strings:
$pool1 = "stratum+tcp://" ascii
$pool2 = "pool.minexmr.com" ascii
$pool3 = "mining.pool" ascii
$algo1 = "cryptonight" ascii nocase
$algo2 = "randomx" ascii nocase
$wallet = "4" ascii // Monero wallet starts with 4
$xmrig = "xmrig" ascii nocase
$hashrate = "H/s" ascii
$config = "--donate-level" ascii
condition:
uint32(0) == 0x464C457F
and (
2 of ($pool1, $pool2, $pool3)
or ($xmrig and 1 of ($algo1, $algo2))
or ($hashrate and $config)
)
}Rule from Byte Pattern (Shellcode)
rule Reverse_Shell_Shellcode {
meta:
description = "x86-64 reverse shell shellcode pattern (execve /bin/sh)"
strings:
// Classic x64 execve("/bin/sh", ...) syscall
// push 0x68 ; mov rbx, 0x732f2f2f6e69622f ; push rbx ; push rsp ; pop rdi
$execve_binsh = {
6A 68
48 B8 2F 62 69 6E 2F 73 68 00
50 54 5F
}
// syscall 59 (execve)
$syscall_execve = { B8 3B 00 00 00 0F 05 }
// /bin/sh string
$binsh = "/bin/sh" ascii
condition:
all of them
or ($execve_binsh and $syscall_execve)
}Decompilation Tools
| Tool | Type | Strengths | How to use |
|---|---|---|---|
| Ghidra | Decompiler + RE platform | Free, NSA-developed, excellent decompiler, scripting (Python/Java), collaborative | ghidraRun → import file → auto-analyse |
| IDA Pro | Disassembler + decompiler | Industry standard, best FLIRT signatures, Hex-Rays decompiler | Commercial; IDA Free available |
| Radare2 (r2) | Disassembler + analysis | CLI, scriptable, supports 100+ archs, carve from memory | r2 -A binary |
| Cutter | GUI for Radare2 | Same power as r2, visual call graph, Ghidra decompiler plugin | cutter binary |
| Binary Ninja | Disassembler + decompiler | Clean UI, HLIL/MLIL IR, Python scripting | Commercial; personal licence available |
| RetDec | Decompiler | Free online/CLI, outputs C pseudocode | retdec-decompiler binary |
| angr | Symbolic execution | Path exploration, vulnerability finding, constraint solving | Python library |
Ghidra Quick Reference
# Run headless analysis and export decompiled C
analyzeHeadless /path/to/project MyProject \
-import suspicious_binary \
-postScript DecompileHeadless.py \
-scriptPath /path/to/scripts/ \
-deleteProject
# Inside Ghidra GUI key shortcuts:
# G → go to address
# L → set label (rename symbol)
# F → open function at cursor
# Ctrl+F → search text in listing
# Ctrl+L → search labels/symbols
# X → cross references (who calls this function)
# P → disassemble at cursor
# Ctrl+Alt+K → define function at cursor
# Right-click → Rename Variable (in decompiler view)Radare2 Quick Reference
r2 -A binary # open and auto-analyse
r2 -d binary # open in debugger mode
# Inside r2:
aa # analyse all (equivalent to -A flag)
afl # list all functions
pdf @ main # print disassembly of function main
pdf @ sym.func_name
pdg @ main # print decompiled code (needs r2ghidra plugin)
iz # list strings in data sections
izz # list strings in entire binary
iE # list exports
ii # list imports
is # list symbols
axt @ sym.connect # cross-references to connect (who calls it)
/r 0x401234 # find cross-references to address
s sym.main; pdf # seek to main, disassemble
VV # visual call graph (ASCII art)
q # quitMalware Family Patterns
Cryptominer Indicators
# String patterns
strings binary | grep -iE 'stratum\+tcp|xmrig|minerd|cryptonight|randomx'
strings binary | grep -iE 'hashrate|difficulty|pool_address|wallet_address'
# Behavioral: CPU usage skyrockets
# strace shows network connection to pool, followed by continuous send/recv
strace -e trace=network ./miner 2>&1 | grep -E 'connect|send|recv'
# Process name masquerading as kernel worker
strings binary | grep -E 'kworker|kdevtmpfs|migration|ksoftirqd'Botnet / RAT Indicators
# C2 beaconing pattern in strace
# Regular connect() + send() + sleep() loop
grep -E '(connect|send|recv|sleep)' strace.log | head -30
# Command handling keywords
strings binary | grep -iE '\b(shell|exec|cmd|upload|download|screenshot)\b'
strings binary | grep -iE 'reverse.?shell|bind.?shell|pty|tty'
# Persistence mechanisms
strings binary | grep -E '(crontab|rc\.local|\.bashrc|\.profile|systemd|service)'
strings binary | grep -E 'chattr\s*\+i' # make self immutable
# Process duplication / daemonization
# strace shows: fork() → setsid() → fork() again (double-fork daemon pattern)
grep -E 'fork|setsid|chdir.*root' strace.logRootkit Indicators
# LD_PRELOAD-based rootkit
strings binary | grep 'LD_PRELOAD\|/etc/ld.so.preload'
strings binary | grep -E 'readdir|getdents|opendir' # hiding files by intercepting these
# Hooking patterns
strings binary | grep -E 'dlsym|dlopen|RTLD_NEXT' # RTLD_NEXT = classic LD_PRELOAD hook
# RTLD_NEXT tells dlsym to find the NEXT occurrence of the function (the real libc one)
# Self-hiding
strings binary | grep -E 'list_del|kobject_del|THIS_MODULE' # LKM self-hiding
# /proc manipulation
strings binary | grep '/proc/net/tcp\|/proc/pid\|/proc/self'Full Workflow Example
# --- PHASE 1: TRIAGE ---
file sample.elf
# → ELF 64-bit LSB executable, x86-64, dynamically linked, stripped
sha256sum sample.elf # check VT
xxd sample.elf | head -2
# → 7f 45 4c 46 02 01 01 00 (ELF)
python3 entropy.py sample.elf
# → Overall entropy: 6.82 (moderately high — possible packing)
# --- PHASE 2: STATIC ---
strings -n 8 sample.elf | grep -iE 'http|tcp|connect|exec|/tmp|passwd'
# → stratum+tcp://pool.minexmr.com:4444
# → /tmp/.x
# → kworker/0:1H
readelf -d sample.elf | grep NEEDED
# → libpthread, libc — standard; no libssl (may use embedded crypto)
nm -D sample.elf | grep ' U '
# → connect, socket, send, recv, fork, setsid, execve, prctl
checksec --file sample.elf
# → NX: enabled, PIE: disabled, No Canary, stripped
# Verdict: likely a cryptominer masquerading as a kernel worker thread
# using stratum protocol. Daemonises itself (fork+setsid). Writes to /tmp.
# --- PHASE 3: DYNAMIC (isolated VM) ---
strace -f ./sample.elf 2>&1 | tee strace.log &
sleep 5; kill %1
grep 'connect\|execve\|openat.*tmp\|prctl' strace.log
# → connect(3, {AF_INET, 195.148.127.141, 4444}, 16) = 0
# → prctl(PR_SET_NAME, "kworker/0:1H")
# → openat("/tmp/.x", O_WRONLY|O_CREAT|O_TRUNC, 0755)
# Confirmed: network connection to miner pool, process rename, drop to /tmp
# --- PHASE 4: YARA ---
cat > miner.yar << 'EOF'
rule XMRig_Miner {
strings:
$pool = "stratum+tcp://" ascii
$proc = "kworker" ascii
$tmp = "/tmp/." ascii
condition:
uint32(0) == 0x464C457F and all of them
}
EOF
yara miner.yar sample.elf
# → XMRig_Miner sample.elf
# --- PHASE 5: DECOMPILATION (Ghidra) ---
# Import → auto-analyse → search for "connect" in symbol tree
# Trace callers → find beacon loop → rename functions → export IOCsAnalyst OPSEC Failures
Everything below is something an analyst does during normal investigation that inadvertently signals to the attacker that their malware has been found. Sophisticated threat actors — APT groups, ransomware operators, some crimeware authors — actively monitor for these signals and use them to burn infrastructure, update tooling, or escalate before containment is complete.
The central principle: every action that reaches attacker-controlled infrastructure or a public platform the attacker can monitor is a potential burn.
Threat Intelligence Lookups
VirusTotal hash queries
VT logs every API and web query. Threat actors with VirusTotal Intelligence (a paid tier) can set up YARA or hash-based Hunting notifications — they receive an email when someone searches for or submits their sample. A hash lookup tells the attacker:
- Their implant was discovered and is under investigation
- Approximate time of discovery (they can correlate with C2 beaconing logs to identify the victim)
- If you use a corporate VT account: your organisation
OPSEC failure: curl -s "https://www.virustotal.com/api/v3/files/<sha256>" -H "x-apikey: $VT_KEY"
Safer: Query private feeds first — CrowdStrike Intelligence, Recorded Future, Mandiant,
internal MISP. Only query VT after containment is complete and you accept the burn.File uploads to public sandboxes
Uploading the binary (not just the hash) to Any.run, Hybrid-Analysis, or Joe Sandbox with public visibility is worse than a hash query. The full analysis report becomes searchable by anyone. Attackers can search their own malware name or hash and see your report — which may include:
- The original filename (may contain the victim's hostname or path)
- Your submitter country
- The exact time you ran it
- Network indicators you triggered, revealing how far you traced the campaign
OPSEC failure: Drag malware onto hybrid-analysis.com → submit as "Public"
Safer: Use private submission (Any.run has private mode, Joe Sandbox Enterprise does too).
Better: self-hosted Cuckoo/CAPE sandbox — nothing leaves your network.WHOIS and certificate lookups via public resolvers
WHOIS queries themselves are not dangerous (registrars don't notify registrants). However, certificate transparency log searches (crt.sh, Censys, Shodan certificate search) can reveal that a specific cert was queried — and Shodan/Censys notify some customers of searches against their infrastructure.
DNS and Network Exposure
Resolving C2 domains from your analysis network
When malware runs in your sandbox and makes a DNS request for c2.example[.]com, that query leaves your network via your DNS resolver. The authoritative DNS server (which the attacker controls or can read logs from) sees:
- The query
- Your resolver's IP (often the corporate NAT or cloud egress IP)
- Timestamp
Even if you only look up the domain manually (dig c2.example.com) from your analyst machine, that query is visible.
OPSEC failure: dig malware-c2.attacker-domain.com
curl -I http://185.220.101.12/beacon # "just checking if it's up"
Safer: Use passive DNS databases — Farsight DNSDB, Cisco Umbrella Investigate,
VirusTotal passive DNS (check TTL/resolution history without making live queries).
Run dynamic analysis in a network-isolated sandbox that sinkhole-resolves DNS.Directly connecting to C2 infrastructure to verify it's alive
Connecting to the C2 server from an analyst machine — even a single TCP SYN — gives the attacker:
- Your IP address (confirms investigation)
- User-Agent if HTTP (analyst tools have recognisable UAs)
- Timing
C2 frameworks like Cobalt Strike, Sliver, and Mythic log all inbound connections. Some operators have tripwires: if a non-beacon IP connects to a listener port, they get an alert.
OPSEC failure: nmap -sV 185.220.101.12 # port scan from corp IP
curl http://185.220.101.12/ # verify HTTP server
Safer: Check from passive sources: Shodan/Censys historical data, Censys IPv4 banners,
Greynoise context. Never touch attacker infra from a corp or identifiable IP.
If you must probe: use a disposable VPS in a neutral cloud region with no links to your org.Scanning attacker infrastructure from a corporate IP
Even if you use nmap rather than curl, a scan from a corporate IP range against a threat actor's server tells them exactly what they need to know. Ransomware groups have posted in forums about detecting investigative scans within hours of deployment.
Sandbox and Analysis Environment Fingerprinting
Re-using analysis VMs with consistent fingerprints
Many malware families perform sandbox detection by checking for consistent environment characteristics:
- Hostname seen in a previous run (
DESKTOP-ANALYSIS01,ANY.RUN,CUCKOO) - Same MAC address prefix (VMware:
00:0C:29, VirtualBox:08:00:27) - Specific usernames (
analyst,malware,sandbox,virus) - Absence of real user artefacts (no browser history, no documents, no real installed apps, only fresh install)
- Screen resolution exactly 1024×768 or 800×600 (common sandbox defaults)
- CPU core count of 1 or 2 (analysts economise on VM resources; real machines have 8–16)
- Uptime less than 5 minutes (sandbox spins up, runs sample, tears down)
When malware detects these patterns it runs a clean code path — your analysis captures nothing meaningful, and you've spent time on a decoy.
OPSEC failure: Run every sample in the same "MalwareAnalysis-Win10" snapshot
Safer: Rotate VM names, MACs, usernames. Use 8 vCPUs. Pre-populate with fake documents,
browser cookies, and plausible user history. Use long-running snapshots (uptime > 7 days).
Tools: FlareVM deployed fresh per engagement, VMCloak for automated VM customisation.Running malware on an analyst machine outside a sandbox
Accidentally double-clicking a sample outside an isolated VM results in a live C2 beacon from your real machine. The attacker now has your actual IP, username, hostname, and OS version. This happens during triage ("let me just check what this does quickly") more often than people admit.
Timing and Coordination Failures
Querying VT or sandbox at the same time the C2 beacon goes silent
If you find the implant, quarantine the host, and immediately query VT — all in the same 5-minute window — a competent attacker correlates:
- Beacon from victim IP stops → host was quarantined
- VT query for their hash arrives 3 minutes later → analyst found the file
They now know which victim, which host, and the investigation timeline. They may have other implants on the same network that they haven't burned yet.
Safer: Complete network-level containment, scope the full compromise (lateral movement,
persistence mechanisms, all beaconing hosts) BEFORE doing any public lookups.
VT queries should happen after you've isolated everything and accepted the attacker
knows.Blocking C2 at the firewall before mapping full scope
Blocking the known C2 IP at the perimeter is satisfying but premature. The moment the beacon dies, the attacker knows:
- Which victim network lost connectivity to the C2
- That detection has occurred
- They may pivot to secondary implants, lateral movement targets, or exfil that was queued
OPSEC failure: Add C2 IP to deny list in the firewall immediately upon discovery
Safer: Monitor first. Let the beacon run (in isolation if possible) while you:
1. Map all hosts beaconing to the same C2
2. Identify any lateral movement or secondary persistence
3. Check for data staging / exfil activity
Block everything simultaneously once you have a complete picture.Publishing IOCs before containment
Sharing indicators (IP, domain, hash, YARA rule) to a public MISP instance, Twitter, a threat intel ISAC, or a blog post while the attacker still has active access:
- They see their IOCs published and immediately know the campaign is burned
- They rotate infrastructure within hours
- They may accelerate exfil or destruction before you can evict them
OPSEC failure: Post "new Lazarus IOCs" to Twitter while still in active IR
Safer: Share IOCs via TLP:RED or TLP:AMBER in closed-circle ISACs only.
Public disclosure only after full eviction or infrastructure has already decayed.Phishing and Email Investigation
Forwarding phishing emails to analysis tools without stripping headers
Email headers contain a full routing trace from the victim's mail server — including your mail infrastructure's IPs and server names. Forwarding a phishing email to an external analysis service exposes this metadata to the service (and potentially to the attacker if the service is compromised or the attacker operates it).
Opening phishing emails or HTML attachments without isolation
HTML emails and attachments commonly contain tracking pixels — 1×1 transparent images at a unique URL. When your email client (or browser) loads the image, the attacker's web server logs:
- Your IP address
- User-Agent (reveals email client, browser, OS version)
- Timestamp of when you opened it
OPSEC failure: Open phishing.html in a browser to "see what it looks like"
Forward original .eml to an external email security vendor for analysis
Safer: Open HTML files in a text editor or use `cat`/`strings` first.
For full rendering: use an isolated VM with no real network access (or sinkholed DNS).
Strip headers before forwarding: use `munpack`, Sublime Text, or mail header parsers
that don't make external network calls.OSINT and Attribution Research
Performing OSINT on attacker personas from a corporate IP
Looking up attacker handles, email addresses, or usernames from LinkedIn, GitHub, Telegram, or forums from a corporate IP or on a logged-in account can:
- Alert the attacker (honeytokens — some actors set canary links in their profiles)
- Reveal your organisation's interest to the platform, which may notify the account
- Tie your investigation to your identity if the attacker reviews their profile visitors
OPSEC failure: Google "APT29 alias" from a corp laptop while logged into Google
Search an attacker's GitHub username from your work machine
Safer: Use a sterile browser profile with no signed-in accounts.
Use a VPN or Tor exit node with no link to your organisation.
Use dedicated OSINT VMs (e.g., OSINT Kombine or a clean Tails session).Attributing publicly before eviction
Publishing a full attribution report ("this is Fancy Bear") while the threat actor still has access to your environment or partner networks is a strategic failure. The actor reads the report, validates what you know and don't know, and adapts. Nation-state groups have cited published IR reports in their own internal retooling decisions.
Communication Channel Compromise
If the attacker has access to your email, Slack, Teams, or ticketing system (all common in persistent access scenarios), then any IR communication on those channels is visible to them in real time. They can read your containment plan, know which hosts are being reimaged, and move laterally to hosts not yet on your radar.
Safer: Establish an out-of-band communication channel for IR:
- Separate Signal group with only IR responders (not on corp infrastructure)
- New Slack workspace outside the company org
- Phone calls / in-person for critical steps
Assume all internal systems are compromised until proven otherwise.Summary: What to Do Before Any Public Lookup
1. Scope first — identify ALL beaconing hosts, lateral movement, persistence mechanisms
2. Contain simultaneously — isolate everything at once, not incrementally
3. Out-of-band comms — use a channel the attacker cannot read
4. Private intel only — internal MISP, closed-feed threat intel, self-hosted sandbox
5. Accept the burn — only then query VT, publish IOCs, notify ISACInterview Questions and Answers
file → sha256sum → entropy check → strings -n 8 | head -40 → readelf -d | grep NEEDED → nm -D | grep ' U ' → checksec. First establish the file type and hash for threat intel, then check entropy (>7.0 = likely packed, deal with unpacking before static analysis is useful). Quick string preview reveals obvious IOCs even without unpacking.
file command tell you, and what does "stripped" mean for analysis?file reads magic bytes and the ELF header to identify architecture (x86-64, ARM, MIPS), linking mode (dynamically/statically linked), and whether the binary is stripped. "Stripped" means the .symtab symbol table was removed — GDB and objdump see only hex addresses instead of function names, making analysis significantly harder. Dynamic symbols (.dynsym) are never stripped.
strings to extract IOCs? What are its limitations?strings -n 8 binary | grep -iE 'https?://|[0-9]{1,3}\.[0-9]{1,3}|/etc/|/tmp/|password' extracts printable ASCII sequences that are at least 8 chars long. Critical limitation: any XOR-encrypted, RC4-encrypted, or otherwise obfuscated strings produce no useful output. Use dynamic analysis (strace/ltrace/Frida) to catch those at runtime when they're decrypted into memory before use.
readelf -S and readelf -l?-S shows sections — the linker's logical view (.text, .data, .rodata, etc.) used by analysis tools. -l shows segments (program headers) — the kernel loader's view of how the file maps into memory at runtime, with permissions (R, W, X flags). Sections don't exist at runtime; segments do. A packed binary's -l may show an RWE LOAD segment — writable and executable — which is a major red flag.
nm -D binary | grep ' U ' tell you, and why is it useful even on stripped binaries?It shows undefined (imported) dynamic symbols — functions the binary calls from shared libraries. These cannot be stripped because the dynamic linker needs them at runtime. Even a fully stripped binary with no function names exposes its capabilities here: connect + send + recv = network capable; execve = can run programs; ptrace = anti-debug or tracing; openat /etc/shadow pattern would appear in strace.
ldd on an untrusted binary? What should you use instead?ldd works by setting LD_TRACE_LOADED_OBJECTS=1 and actually executing the binary. A malicious binary can detect this environment variable and run harmful code — C2 callbacks, persistence installation, or data destruction. Use readelf -d binary | grep NEEDED instead — it reads the .dynamic section directly without executing anything.
Four signals: (1) high entropy — overall entropy >7.0 or .text section >7.2 (normal code is 4.5–6.0); (2) minimal import table — only VirtualAlloc, ExitProcess, LoadLibrary, GetProcAddress; (3) UPX magic strings — strings | grep UPX or readelf -S | grep UPX0; (4) tiny .text section relative to a large high-entropy .data or unnamed section — the stub is in .text, real code is compressed in .data.
strace and ltrace. When would you use each?strace intercepts at the kernel syscall boundary — open, connect, execve. Always works, unavoidable, works on statically linked binaries. ltrace intercepts at the libc function boundary — strcmp, malloc, puts, fopen. More readable (shows function names) but defeated by statically linked binaries (no shared lib calls to intercept). Use strace first for network/file/process activity; ltrace for string comparisons and config parsing to catch decrypted values.
Reverse shellconnect() to remote IP → dup2(fd, 0) / dup2(fd, 1) / dup2(fd, 2) (redirect stdin/stdout/stderr to socket) → execve("/bin/sh", ...). Miner: connect() to mining pool IP → tight alternating send()/recv() loop. Credential stealer: openat("/etc/shadow", O_RDONLY) or /proc/*/environ reads, or PAM library calls, followed by a network connect() to exfil.
Find the decrypt function address in Ghidra (look for XOR loop pattern). Then: Interceptor.attach(ptr("0xADDRESS"), { onEnter: function(args) { this.out = args[0]; this.len = parseInt(args[2]); }, onLeave: function() { console.log(Memory.readUtf8String(this.out, this.len)); } }). By the time onLeave fires, the output buffer contains plaintext.
/bin/sh.rule ELF_Reverse_Shell { strings: $elf = { 7F 45 4C 46 } $ip = "1.2.3.4" ascii $sh = "/bin/sh" ascii $exec = { B8 3B 00 00 00 0F 05 } // execve syscall condition: $elf at 0 and $ip and $sh and $exec }Key: uint32(0) == 0x464C457F or $elf at 0 anchors to the ELF magic. Real rule would combine port 4444 context with connect import.
.text section tell you? What is the threshold?High entropy in .text means the "code" is actually compressed or encrypted data — you're looking at a packer stub, not real assembly. Normal x86-64 code entropy is 4.5–6.0 because instructions use a non-uniform byte distribution. Threshold: >7.0 is suspicious, >7.2 in .text is essentially certain to be packed. The real code lives elsewhere (usually .data) in a compressed blob.
(1) Break on mprotect or VirtualAlloc — the stub allocates executable memory for the real binary. Note the returned address. (2) Set a hardware execute breakpoint at that address. (3) Resume — the stub decompresses the real binary into the allocation, then jumps to it. (4) Hardware breakpoint fires at the OEP. (5) Dump memory with Scylla (x64dbg) or gdb's dump binary memory out.bin START END. Try upx -d first — only falls back to debugger if UPX headers were stripped.
Ghidrafree, NSA-developed, excellent C decompiler (often better than IDA's), GUI with call graph, best for complex stripped binaries and collaborative analysis. radare2: CLI-first, fast startup, supports 100+ architectures, best for automation/scripting, embedded firmware (MIPS/ARM IoT), quick terminal-based analysis. For RE work on a campaign, Ghidra. For a quick "what does this do?" on an ARM router binary from a terminal, radare2.
Look for: (1) RTLD_NEXT in strings output — the classic pattern for hooking libc functions (get the real function pointer, then wrap it); (2) dlsym/dlopen imports; (3) hooks on readdir, getdents64, opendir (for file hiding) or getpwnam, pam_authenticate (for credential theft); (4) /etc/ld.so.preload path in strings (persistence mechanism). nm -D rootkit.so | grep 'readdir\|getdents\|fopen' reveals which libc functions are being replaced.