tcpdump, BPF Filters, eBPF & TCP Stack Deep Dive
Visual referenceTCPDUMP_OFFSETS_HTML_v2.2_bpf.svg — byte-offset diagram for writing BPF filters against TCP/IP headers. Open in a browser for the interactive version.
Part 1: TCP/IP Packet Header Anatomy
IPv4 Header (20 bytes minimum, IHL × 4 bytes total)
Byte 0: Version (4 bits) | IHL (4 bits)
Byte 1: TOS = DSCP (6 bits) | ECN (2 bits)
Bytes 2-3: Total Length
Bytes 4-5: Identification
Byte 6: IP Flags (3 bits) | Fragment Offset (13 bits)
Byte 7: (fragment offset continued — use 6:2 mask 0x1FFF)
Byte 8: TTL
Byte 9: Protocol (6=TCP, 17=UDP, 1=ICMP, 47=GRE)
Bytes 10-11: Header Checksum
Bytes 12-15: Source Address
Bytes 16-19: Destination Address
Bytes 20+: Options (if IHL > 5)TOS byte breakdown (byte 1):
bits 7-2 = DSCP (6 bits)
bits 1-0 = ECN (2 bits)
00 = Not ECN-Capable Transport (Non-ECT)
01 = ECN Capable Transport (ECT(1))
10 = ECN Capable Transport (ECT(0))
11 = Congestion Experienced (CE)IP Flags (byte 6 upper 3 bits):
Bit 7 (0x80) = Reserved (must be 0)
Bit 6 (0x40) = DF (Don't Fragment)
Bit 5 (0x20) = MF (More Fragments)IPv6 Header (40 bytes fixed)
Bytes 0-3: Version (4b) | Traffic Class (8b) | Flow Label (20b)
Bytes 4-5: Payload Length
Byte 6: Next Header (same protocol numbers as IPv4 Protocol field)
Byte 7: Hop Limit
Bytes 8-23: Source Address (128-bit)
Bytes 24-39: Destination Address (128-bit)Traffic Class = DSCP (6 bits) + ECN (2 bits) — same as IPv4 TOS.
TCP Header (20 bytes minimum)
Bytes 0-1: Source Port
Bytes 2-3: Destination Port
Bytes 4-7: Sequence Number
Bytes 8-11: Acknowledgment Number
Byte 12: Data Offset (4b) | Reserved (3b) | NS flag (1b)
Byte 13: TCP Flags: CWR | ECE | URG | ACK | PSH | RST | SYN | FIN
Bytes 14-15: Window Size
Bytes 16-17: Checksum
Bytes 18-19: Urgent Pointer
Bytes 20+: OptionsTCP Flag byte 13 bit mapping
Bit 7 (0x80) = CWR (Congestion Window Reduced)
Bit 6 (0x40) = ECE (ECN-Echo)
Bit 5 (0x20) = URG
Bit 4 (0x10) = ACK
Bit 3 (0x08) = PSH
Bit 2 (0x04) = RST
Bit 1 (0x02) = SYN
Bit 0 (0x01) = FINCommon flag values for tcp[13]:
| Value | Decimal | Flags |
|---|---|---|
0x02 | 2 | SYN |
0x12 | 18 | SYN+ACK |
0x10 | 16 | ACK |
0x18 | 24 | PSH+ACK |
0x04 | 4 | RST |
0x01 | 1 | FIN |
0xC2 | 194 | SYN+ECE+CWR |
UDP Header (8 bytes)
Bytes 0-1: Source Port
Bytes 2-3: Destination Port
Bytes 4-5: Length
Bytes 6-7: ChecksumData begins at byte 8 of the UDP header.
ICMP Header (8 bytes minimum)
Byte 0: Type
Byte 1: Code
Bytes 2-3: Checksum
Bytes 4-7: Type-specific (e.g., ID + Sequence for Echo)Common ICMP types:
| Type | Code | Meaning |
|---|---|---|
| 0 | 0 | Echo Reply |
| 3 | 0-15 | Destination Unreachable |
| 8 | 0 | Echo Request |
| 11 | 0 | TTL Exceeded in Transit |
| 11 | 1 | Fragment Reassembly Timeout |
Ethernet Frame
Bytes 0-5: Destination MAC
Bytes 6-11: Source MAC
Bytes 12-13: EtherType (0x0800=IPv4, 0x0806=ARP, 0x86DD=IPv6, 0x8100=802.1Q)
Bytes 14+: Payload802.1Q (VLAN tag) inserts 4 bytes after the Source MAC:
Bytes 12-13: 0x8100 (TPID)
Bytes 14-15: PCP (3b) | DEI (1b) | VLAN ID (12b)
Bytes 16-17: EtherType (0x0800 for IPv4)
Bytes 18+: Payload (so IP header starts at ether[18])Part 2: tcpdump BPF Filter Reference
BPF syntax: proto[offset:size] operator value
proto= ip / ip6 / tcp / udp / icmp / etheroffset= byte offset from start of that headersize= 1, 2, or 4 bytes (default 1)- Masks: use
& 0xNNfor bit masking
IPv4 Filters
# Source/destination IP
src host 192.168.1.1
dst host 192.168.2.2
host 1.1.1.1 # either direction
# Protocol field (byte 9)
ip[9]=6 # TCP
ip[9]=17 # UDP
ip[9]=1 # ICMP
ip[9]=47 # GRE
# TTL (byte 8)
ip[8]=255 # TTL=255 (often Cisco routers)
ip[8]=64 # TTL=64 (Linux default)
# DF flag (Don't Fragment) — byte 6 upper bits
ip[6] & 0x40 = 0x40 # DF set
ip[6] & 0x20 = 0x20 # MF set (fragmented packet, not last)
ip[6] & 0x80 = 0x00 # Reserved bit clear (should always be)
# Fragment offset — bytes 6-7, lower 13 bits
ip[6:2] & 0x1FFF != 0 # non-zero fragment offset (not first fragment)
# Total length (bytes 2-3)
ip[2:2] = 123 # specific total length
# Identification field (bytes 4-5)
ip[4:2] = 0x3805 # specific identification value
# Header checksum (bytes 10-11)
ip[10:2] = 0x4051
# DSCP — byte 1 upper 6 bits (& 0xFC)
ip[1] & 0xfc = 0x00 # CS0 (default, best effort)
ip[1] & 0xfc = 0x28 # CS2 (0x10 << 2 = 0x28... actually DSCP 10 = 0x28)
ip[1] & 0xfc = 0xb8 # EF (Expedited Forwarding, decimal 46 << 2 = 0xb8)
# ECN — byte 1 lower 2 bits
ip[1] & 0x03 = 0x01 # ECT(1)
ip[1] & 0x03 = 0x02 # ECT(0)
ip[1] & 0x03 = 0x03 # CE (Congestion Experienced)
ip[1] & 0x03 = 0x00 # Non-ECTTCP Filters
# Port filters (preferred high-level syntax)
tcp port 443
tcp dst port 22
tcp src port 1024
# SYN packets only
tcp[13] = 2 # SYN (no ACK)
tcp[13] & 0x02 = 0x02 # SYN bit set (matches SYN and SYN+ACK)
# SYN+ACK
tcp[13] = 18 # 0x12 = SYN+ACK
# RST
tcp[13] & 0x04 = 0x04 # RST bit set
# FIN
tcp[13] & 0x01 = 0x01 # FIN bit set
# PSH+ACK (data packets)
tcp[13] = 24 # 0x18 = PSH+ACK
# ECN flags in TCP
tcp[13] & 0x40 = 0x40 # ECE (ECN-Echo) set
tcp[13] & 0x80 = 0x80 # CWR (Congestion Window Reduced) set
# SYN with ECN negotiation (SYN+ECE+CWR = 0xC2)
tcp[13] = 0xC2
# SYN+ACK with ECE (server agrees to ECN = 0x52)
tcp[13] = 0x52
# Sequence number (bytes 4-7)
tcp[4:4] = 0 # ISN=0 (might indicate spoofing or test traffic)
# ACK number (bytes 8-11)
tcp[8:4] = 0 # ACK=0 (should be 0 on SYN only)
# Urgent pointer (bytes 18-19)
tcp[18:2] = 0 # Urgent not active
# Data offset = header length in 32-bit words (upper 4 bits of byte 12)
tcp[12] & 0xF0 = 0x50 # 20-byte header (standard, no options)
tcp[12] & 0xF0 = 0x60 # 24-byte header
tcp[12] & 0xF0 = 0x80 # 32-byte header (with full timestamp options)
# Detect port scans: SYN with no other flags
'tcp[13] = 2 and not src net 10.0.0.0/8'
# Detect NULL scan (no flags)
tcp[13] = 0
# Detect XMAS scan (FIN+PSH+URG)
tcp[13] = 0x29UDP Filters
# Source/destination port
udp port 53 # DNS
udp dst port 514 # Syslog
udp src port 123 # NTP
# Raw offset filters (8-byte header, data at offset 8)
udp[0:2] = 123 # source port = 123
udp[2:2] = 443 # destination port = 443
udp[4:2] = 22 # length field = 22
# DNS query (in UDP payload, data at udp[8])
# QR bit = 0 means query, bit 7 of byte 2 of DNS header = udp[10] bit 7
udp[10] & 0x80 = 0x00 # DNS query (QR=0)
udp[10] & 0x80 = 0x80 # DNS response (QR=1)
# DNS opcode (bits 3-6 of byte 2)
udp[10] & 0x78 = 0x00 # standard query
udp[10] & 0x78 = 0x08 # inverse query
udp[10] & 0x78 = 0x10 # server status
# DNS authoritative answer
udp[10] & 0x04 = 0x00 # non-authoritative
udp[11] & 0x80 = 0x80 # recursion availableICMP Filters
# Type
icmp[0] = 8 # Echo Request (ping)
icmp[0] = 0 # Echo Reply
icmp[0] = 3 # Destination Unreachable
icmp[0] = 11 # Time Exceeded (TTL expired)
# Type + Code
icmp[0] = 11 and icmp[1] = 0 # TTL exceeded in transit
icmp[0] = 11 and icmp[1] = 1 # Fragment reassembly timeout
icmp[0] = 3 and icmp[1] = 3 # Port Unreachable
# ICMP Echo ID + Sequence
icmp[4:2] = 0x0840 # specific ICMP ID
icmp[6:2] = 0x0002 # sequence number = 2Ethernet Filters
# EtherType
ether[12:2] = 0x0800 # IPv4
ether[12:2] = 0x0806 # ARP
ether[12:2] = 0x86DD # IPv6
ether[12:2] = 0x8100 # 802.1Q (VLAN tagged)
# 802.1Q VLAN ID (lower 12 bits of bytes 14-15)
ether[14:2] & 0x0FFF = 101 # VLAN 101
# 802.1Q: PCP priority (upper 3 bits of byte 14)
ether[14] & 0xE0 = 0x00 # priority 0 (best effort)
# Source MAC prefix (first 3 bytes = OUI)
ether[6:2] = 0x10f3 and ether[8] = 0x11 # specific OUI
# Destination MAC prefix
ether[0:2] = 0x10f3 and ether[2] = 0x11GRE Tunnel Filters (IP in IP)
When traffic is GRE-encapsulated, the inner packet starts after:
- Outer Ethernet (14 bytes)
- Outer IPv4 (20 bytes)
- GRE header (4-16 bytes, typically 4 for plain GRE) = inner IP starts at byte 38 (for standard GRE)
# Inner TCP in GRE (plain GRE, no key/seq)
# Outer IP protocol = GRE (47)
ip[9] = 47
# Inner IPv4 header starts at ip[24] (after 20B outer IP + 4B GRE header)
ip[24] & 0xF0 = 0x40 # inner packet is IPv4
# Inner destination port (inner IP + 20B + TCP port at offset 2)
# inner TCP dst port at ip[46:2]
ip[46:2] = 443 # HTTPS inside GRE
ip[46:2] = 22 # SSH inside GRE
# Inner ICMP type inside GRE
ip[44] = 11 # TTL exceeded inside GRE
ip[45] = 1 # code 1 (frag reassembly)IPv6 Filters
# Payload length (bytes 4-5)
ip6[4:2] = 0x0020 # payload = 32 bytes
# Next header (byte 6)
ip6[6] = 0x06 # TCP
ip6[6] = 0x11 # UDP
ip6[6] = 0x3a # ICMPv6
# Hop limit (byte 7)
ip6[7] = 0x31 # hop limit = 49
# ECN in IPv6 Traffic Class (byte 1, bits 1-0)
ip6[1] & 0x03 = 0x03 # CE (Congestion Experienced)
ip6[1] & 0x03 = 0x00 # Non-ECT
# Inner TCP destination port in IPv6 (TCP header starts at byte 40)
ip6[42:2] = 443 # dst port 443
ip6[44:4] = 0x8b164221 # sequence numberUseful tcpdump Filter Combinations
# Catch all SYN packets (new connections) except from our monitoring host
tcpdump -n 'tcp[13] = 2 and not src host 10.0.0.1'
# Catch RST packets (connection resets — useful for detecting port scans)
tcpdump -n 'tcp[13] & 0x04 = 0x04'
# DNS exfiltration: long DNS queries (> 50 bytes of name data)
tcpdump -n 'udp port 53 and udp[4:2] > 60'
# ECN-capable connections
tcpdump -n 'tcp[13] = 0xC2' # client offering ECN
# Fragmented packets (potential evasion)
tcpdump -n '(ip[6] & 0x20 = 0x20) or (ip[6:2] & 0x1FFF != 0)'
# ICMP tunneling detection (large ICMP packets)
tcpdump -n 'icmp[0] = 8 and ip[2:2] > 100'
# Null/XMAS scan detection
tcpdump -n 'tcp[13] = 0 or tcp[13] = 0x29'
# SYN flood: many SYN, no ACK back
tcpdump -n -c 1000 'tcp[13] = 2' | awk '{print $3}' | sort | uniq -c | sort -rn | headPart 3: DSCP & QoS Reference
DSCP — Differentiated Services Code Point (6 bits in TOS/Traffic Class)
The TOS byte (IPv4 byte 1) is split: upper 6 bits = DSCP, lower 2 bits = ECN.
Class Selector Values (backward-compatible with IP Precedence)
| DSCP Name | Binary | Hex | Decimal | Typical Use |
|---|---|---|---|---|
| CS0 (Default) | 000 000 | 0x00 | 0 | Best effort |
| CS1 | 001 000 | 0x08 | 8 | Scavenger (YouTube, Gaming, P2P) |
| CS2 | 010 000 | 0x10 | 16 | OAM (SNMP, SSH, Syslog) |
| CS3 | 011 000 | 0x18 | 24 | Signaling (SCCP, SIP, H.323) |
| CS4 | 100 000 | 0x20 | 32 | Realtime (TelePresence) |
| CS5 | 101 000 | 0x28 | 40 | Broadcast video (Cisco IPVS) |
| CS6 | 110 000 | 0x30 | 48 | Network control (EIGRP, OSPF, HSRP, IKE) |
| CS7 | 111 000 | 0x38 | 56 | Reserved |
Assured Forwarding (AF) Values — Per-Hop Behavior
Format: AFXY where X = class (1-4), Y = drop probability (1=low, 2=medium, 3=high)
| DSCP | Binary | Hex | Decimal | Drop probability |
|---|---|---|---|---|
| AF11 | 001 010 | 0x0a | 10 | Low |
| AF12 | 001 100 | 0x0c | 12 | Medium |
| AF13 | 001 110 | 0x0e | 14 | High |
| AF21 | 010 010 | 0x12 | 18 | Low |
| AF22 | 010 100 | 0x14 | 20 | Medium |
| AF23 | 010 110 | 0x16 | 22 | High |
| AF31 | 011 010 | 0x1a | 26 | Low |
| AF32 | 011 100 | 0x1c | 28 | Medium |
| AF33 | 011 110 | 0x1e | 30 | High |
| AF41 | 100 010 | 0x22 | 34 | Low |
| AF42 | 100 100 | 0x24 | 36 | Medium |
| AF43 | 100 110 | 0x26 | 38 | High |
| EF | 101 110 | 0x2e | 46 | N/A (Expedited Forwarding — VoIP) |
BPF filter to match DSCP value
# DSCP = EF (46 decimal = 0x2e, but shifted left 2 = 0xb8 as full byte value)
ip[1] & 0xfc = 0xb8
# Match any AF4x class (bits 7-4 = 0100 = 0x40, masked & 0xE0)
ip[1] & 0xe0 = 0x80Part 4: ECN (Explicit Congestion Notification)
RFC 3168 — ECN is a mechanism that allows routers to signal congestion to endpoints before packet drops occur, enabling earlier, smoother congestion response.
ECN Bits (TOS byte 1, lower 2 bits)
| Bits | Name | Meaning |
|---|---|---|
00 | Non-ECT | Not ECN-Capable Transport |
01 | ECT(1) | ECN-Capable Transport (endpoint supports ECN) |
10 | ECT(0) | ECN-Capable Transport (endpoint supports ECN) |
11 | CE | Congestion Experienced (set by congested router) |
ECN TCP Flags (byte 13)
| Flag | Bit | Meaning |
|---|---|---|
| CWR | bit 7 (0x80) | Congestion Window Reduced — sender acknowledges it received CE |
| ECE | bit 6 (0x40) | ECN-Echo — receiver echoes CE back to sender |
ECN Three-Way Handshake (RFC 3168)
Client → Server: SYN, ECE, CWR (flags 0xC2: I support ECN)
Server → Client: SYN, ACK, ECE (flags 0x52: I also support ECN, agreed)
Client → Server: ACK (standard ACK, ECN negotiation complete)tcpdump filter
# Capture ECN handshake
tcpdump 'tcp[13] = 0xC2 or tcp[13] = 0x52'ECN Congestion Signaling (during data transfer)
1. ECN-capable endpoints set ECT(0) or ECT(1) on IP packets
2. A congested router marks packets with CE (0x11 in ECN bits) instead of dropping
3. The receiver echoes CE back to the sender via ECE flag in TCP ACK
4. The sender reduces its congestion window and sets CWR in the next TCP segment
5. Receiver clears ECE once CWR is seenSender Congested Router Receiver │ │ │ ├──── data, ECT(0) ─────────────────►│ │ │ ├──── data, CE ───────────►│ │ │ │ (sets ECE) │◄──────────────────────────────── ACK + ECE ──────────────────┤ │ (reduces cwnd, sets CWR) │ │ ├──── data, CWR+ECT(0) ─────────────► │
ECN Lab Setup (from presentation)
Topology: Ubuntu_14.04_1 (192.168.2.2) → R1 → R2 → Ubuntu_14.04_2 (192.168.1.2)
- Both endpoints ECN-capable
- File transfer used to generate traffic
- Captured and analysed in Wireshark
Linux ECN Kernel Settings
# Check current ECN settings
sysctl -a | grep ecn
# Enable ECN
sysctl net.ipv4.tcp_ecn=1 # Enable ECN (request ECN in new connections)
sysctl net.ipv4.tcp_ecn=2 # Only respond to ECN requests (more conservative)
sysctl net.ipv4.tcp_ecn=0 # Disable ECN
# ECN fallback (if ECN negotiation fails, fall back to non-ECN)
sysctl net.ipv4.tcp_ecn_fallback=1ECN values
tcp_ecn | Behaviour |
|---|---|
0 | Disabled |
1 | Enabled — actively request ECN in SYN |
2 | Passive — only agree to ECN if peer requests |
# Verify via sysctl
sysctl net.ipv4.tcp_ecn # should return 1 if enabled
sysctl net.ipv4.tcp_ecn_fallback # should return 1Scenario Analysis (from presentation)
Scenario 1: ECN capable, CWR not sent back
- SYN has ECE+CWR (0xC2) — client offers ECN
- SYN+ACK has ECE (0x52) — server agrees
- ACK sent (ECN negotiated)
- Data flows: ECT(0) marked in IP TOS
- CE returned by router → receiver sends ECE in ACK
- Bug scenario: sender doesn't send CWR — this is a protocol violation; receiver keeps sending ECE indefinitely
Scenario 2: ECN full handshake, real file transfer
- Full ECN negotiation visible in Wireshark
- Data segments marked ECT(0):
ip[1] & 0x03 = 0x02 - When congestion occurs: CE marking visible on packets
- CWR flag observed in subsequent sender segments
Scenario 3: Non-ECN setup (control/baseline)
- Standard SYN without ECN flags
- No ECN marking in data flow
- Packet drops instead of CE marking = worse performance
tcpdump Filters for ECN Analysis
# Complete ECN negotiation capture
tcpdump -n 'tcp[13] & 0xC0 != 0' # any ECN TCP flags (ECE or CWR)
# ECN-capable packets (ECT in IP)
tcpdump -n 'ip[1] & 0x03 != 0' # ECT(0), ECT(1), or CE
# Congestion events (CE marked by router)
tcpdump -n 'ip[1] & 0x03 = 0x03' # CE codepoint
# ECN echo (receiver signaling congestion back to sender)
tcpdump -n 'tcp[13] & 0x40 = 0x40' # ECE bit set
# Congestion window reduced (sender acknowledging)
tcpdump -n 'tcp[13] & 0x80 = 0x80' # CWR bit set
# Full ECN handshake SYN
tcpdump -n 'tcp[13] = 0xC2'
# Full ECN handshake SYN+ACK
tcpdump -n 'tcp[13] = 0x52'Part 5: eBPF Deep Dive
What is eBPF?
eBPF (extended Berkeley Packet Filter) is a technology that allows safely running sandboxed programs in the Linux kernel without changing kernel source code or loading kernel modules.
User space application
│
│ bpf() syscall
▼
BPF Verifier ─── static analysis, safety checks
│ (no loops that won't terminate,
│ no invalid memory access,
│ privilege check)
▼
JIT Compiler → native machine code
│
▼
Hook Point (where eBPF attaches)
┌──────────────────────────────────────────┐
│ kprobes/kretprobes — kernel functions │
│ uprobes/uretprobes — user-space │
│ tracepoints — stable kernel events │
│ XDP — earliest network packet hook │
│ TC (traffic control) — network hooks │
│ Socket filters — per-socket packet filter│
│ LSM hooks — security policy enforcement │
│ perf events — performance monitoring │
└──────────────────────────────────────────┘cBPF vs eBPF
| Feature | cBPF (classic) | eBPF |
|---|---|---|
| Registers | 2 (A, X) | 11 (R0-R10) |
| Register width | 32-bit | 64-bit |
| Maps | None | Yes (hash, array, LRU, ring buffer, etc.) |
| Helper functions | None | Many (200+) |
| Program types | socket filter | 30+ types |
| Verification | Basic | Full static verification |
| JIT | Limited | Full on all arches |
| Use cases | tcpdump, iptables | Networking, security, tracing, profiling |
tcpdump still uses cBPF under the hood — the filter expression is compiled to cBPF bytecode and passed to the kernel socket.
eBPF Program Types (Security-Relevant)
| Type | Attach Point | Security Use Case |
|---|---|---|
BPF_PROG_TYPE_SOCKET_FILTER | Socket | Deep packet inspection, filtering |
BPF_PROG_TYPE_KPROBE | Kernel function entry/exit | Syscall monitoring, kernel tracing |
BPF_PROG_TYPE_TRACEPOINT | Kernel tracepoints | System call auditing |
BPF_PROG_TYPE_XDP | Driver RX before network stack | DDoS mitigation, fast packet drop |
BPF_PROG_TYPE_TC | Traffic Control ingress/egress | Network policy enforcement |
BPF_PROG_TYPE_LSM | Linux Security Module hooks | MAC policy (BPF-LSM) |
BPF_PROG_TYPE_CGROUP_SKB | cgroup network | Container network policy |
BPF_PROG_TYPE_PERF_EVENT | perf subsystem | Performance monitoring, profiling |
eBPF Maps — Data Structures
// Hash map for tracking connections
struct {
__uint(type, BPF_MAP_TYPE_HASH);
__type(key, struct ip_port);
__type(value, u64);
__uint(max_entries, 65536);
} connection_count SEC(".maps");
// Ring buffer for userspace communication
struct {
__uint(type, BPF_MAP_TYPE_RINGBUF);
__uint(max_entries, 4096 * 1024);
} events SEC(".maps");Map types:
| Type | Description | Use Case |
|---|---|---|
BPF_MAP_TYPE_HASH | Key-value hash table | Connection tracking |
BPF_MAP_TYPE_ARRAY | Fixed-size indexed array | Per-CPU counters |
BPF_MAP_TYPE_LRU_HASH | Hash with LRU eviction | Connection state (bounded) |
BPF_MAP_TYPE_RINGBUF | Ring buffer for events | Efficient kernel→user events |
BPF_MAP_TYPE_PERCPU_HASH | Per-CPU hash map | Lock-free counters |
BPF_MAP_TYPE_PERF_EVENT_ARRAY | Perf events | Older event mechanism |
XDP for Network Security
XDP (eXpress Data Path) runs eBPF programs at the NIC driver level, before the kernel network stack. Can process millions of packets per second.
SEC("xdp")
int xdp_drop_syn_flood(struct xdp_md *ctx) {
void *data = (void *)(long)ctx->data;
void *data_end = (void *)(long)ctx->data_end;
struct ethhdr *eth = data;
if ((void *)(eth + 1) > data_end) return XDP_PASS;
struct iphdr *ip = (void *)(eth + 1);
if ((void *)(ip + 1) > data_end) return XDP_PASS;
if (ip->protocol != IPPROTO_TCP) return XDP_PASS;
struct tcphdr *tcp = (void *)ip + (ip->ihl * 4);
if ((void *)(tcp + 1) > data_end) return XDP_PASS;
// Drop SYN packets from blocked IPs
if (tcp->syn && !tcp->ack) {
if (is_blocked(ip->saddr))
return XDP_DROP;
}
return XDP_PASS;
}XDP return codes:
| Code | Action |
|---|---|
XDP_PASS | Pass to kernel network stack |
XDP_DROP | Drop packet (fast, no allocation) |
XDP_TX | Redirect back out same interface |
XDP_REDIRECT | Redirect to another interface/socket |
XDP_ABORTED | Drop + trace (error path) |
eBPF for Security Observability (Falco)
Falco uses eBPF (or kernel module) to hook into kernel syscalls and detect anomalous behaviour in containers.
User space: Falco daemon
│
│ reads events from ring buffer
▼
eBPF ring buffer (kernel space)
│
│ written by eBPF programs at:
▼
tracepoints:
sys_enter_execve — process execution
sys_enter_connect — outbound connection
sys_enter_openat — file open
sys_enter_write — file write
sys_exit_clone — process creationeBPF Security Vulnerabilities (Interview Must-Know)
Kernel eBPF verifier bugs are a serious attack surface for local privilege escalation.
| CVE | Type | Details |
|---|---|---|
| CVE-2021-3490 | OOB r/w in ALU32 verifier | Bitwise op on 32-bit sub-registers; wrong bounds tracking |
| CVE-2021-31440 | OOB write in verifier | ALU32 ops not properly bounds-checked |
| CVE-2021-34866 | Type confusion in verifier | Pointers aliased through map of maps |
| CVE-2022-23222 | OOB r/w via map_push_elem | Wrong pointer range in verifier |
| CVE-2021-4001 | Race condition in eBPF maps | Concurrent access to shared map |
Attack path
- User triggers verifier bug via
bpf()syscall with crafted program - Verifier incorrectly believes memory access is safe
- JIT-compiled program executes arbitrary kernel read/write
- Use kernel r/w to escalate to root (e.g., overwrite cred structure)
Prerequisites
unprivileged_bpf_disabled = 0(allows unprivileged users to load BPF programs)- Or user namespaces enabled (gives
CAP_BPFin namespace)
Mitigations
sysctl kernel.unprivileged_bpf_disabled=1 # disable unprivileged eBPF
sysctl kernel.unprivileged_userns_clone=0 # disable user namespaces (Debian/Ubuntu)Part 6: Practical tcpdump Recipes
# Live capture on interface, write to pcap
tcpdump -i eth0 -w /tmp/capture.pcap
# Read pcap file, verbose output
tcpdump -r capture.pcap -v
# No DNS resolution, show port numbers
tcpdump -n -nn
# Capture with snaplen 1500 (full packet)
tcpdump -s 1500
# Capture to rotating files (100MB each, keep 10)
tcpdump -C 100 -W 10 -w /tmp/capture.pcap
# Filter by host and port
tcpdump 'host 192.168.1.1 and port 443'
# Exclude SSH (port 22) to avoid capturing own session
tcpdump 'not port 22'
# Show hex and ASCII
tcpdump -X -s 0 'port 80'
# Capture HTTPS SNI (TLS Client Hello)
tcpdump -n 'tcp port 443 and tcp[((tcp[12:1] & 0xf0) >> 2):4] = 0x16030100'
# Detect SYN flood
tcpdump -n -c 10000 'tcp[13] = 2' 2>/dev/null | awk '{print $3}' | cut -d. -f1-4 | sort | uniq -c | sort -rn | head -20
# Capture DNS traffic
tcpdump -n 'udp port 53 or tcp port 53'
# Capture ECN-related packets
tcpdump -n '(tcp[13] & 0xC0 != 0) or (ip[1] & 0x03 != 0)'Memory hookeBPF is "JavaScript for the kernel": safe, sandboxed, event-driven code you load at runtime. Just as a browser runs a web page's code in a sandbox without crashing the browser, eBPF runs small programs attached to kernel events — every syscall, packet, or function entry — without writing a kernel module or rebooting. The magic and the safety come from the verifier: before loading, the kernel statically proves the program can't loop forever, can't read arbitrary memory, and will terminate — so near-untrusted code runs inside the kernel safely. That's why modern security tooling (Falco, Cilium, Tetragon) is eBPF-based: kernel-level visibility into every process and packet, low overhead, no kernel patching. Mnemonic: eBPF = sandboxed kernel programs gatekept by the verifier — the engine behind modern runtime detection.
Part 7: Interview Questions
tcpdump & BPF
tcp[tcpflags] == tcp-syn — equality matches segments where the flags byte is exactly SYN and nothing else, so SYN+ACK (both bits set) is excluded. The distinction matters: tcp[tcpflags] & tcp-syn != 0 matches any segment with SYN set, including SYN+ACK. Exact equality isolates connection-initiation attempts — the client's first packet — which is what you want for detecting SYN scans or counting inbound connection attempts per source.
ICMP tunneling smuggles data in echo request/reply payloads, which are normally tiny and uniform. I'd capture with tcpdump -i eth0 -X icmp and inspect the payload: legitimate pings carry a small predictable pattern, whereas tunneled ICMP shows unusually large payloads, high-entropy or base64-looking content, asymmetric request/reply sizes, and high volume to a single host. The -X hex/ASCII dump reveals the content, and statistically a flood of large echo packets to one destination is the tell. The same logic generalizes — covert channels appear as a boring protocol carrying atypical volume and entropy.
tcp port 80 and tcp[2:2] = 80?tcp port 80 is a high-level primitive matching port 80 as either source or destination, with BPF handling header offsets. tcp[2:2] = 80 reads 2 raw bytes at offset 2 of the TCP header — the destination port specifically — and compares to 80. So the raw form is narrower (destination only) and more brittle, but it's how you express conditions the named primitives don't cover. It shows tcpdump filters can use convenient keywords or drop to raw header arithmetic when you need precision.
GRE wraps the inner packet, so a plain port 443 misses it because the outer packet is IP protocol 47, not TCP. Capture the tunnel with proto gre or ip proto 47, then decode the inner payload — modern tcpdump dissects GRE, or carve the inner TLS in Wireshark. For fragmented IPv4, the info is in the IP flags/offset field: tcpdump 'ip[6:2] & 0x3fff != 0' matches any fragment — nonzero when the More Fragments bit is set or the fragment offset is nonzero. Fragmentation filters matter because fragmentation is a classic IDS-evasion and DoS technique.
ECN & TCP Stack
Explicit Congestion Notification lets routers signal congestion by marking a bit in the IP header instead of dropping a packet. Without ECN the only congestion signal is a drop, which the sender detects as loss and reacts to — at the cost of a retransmission and added latency. With ECN, a congested router sets the CE codepoint, the receiver echoes it back via the TCP ECE flag, and the sender reduces its congestion window without anything being lost or retransmitted. So you get the congestion-control benefit without the loss and latency spike, which especially helps latency-sensitive and high-throughput flows. It needs support at both endpoints and the routers between.
ECN is negotiated in the handshake: the client's SYN sets both ECE and CWR ("I support ECN"), the server's SYN-ACK sets ECE alone to confirm. During data flow the signal path is: a congested router sets the CE codepoint in the IP header's ECN field; the receiver sees CE and sets the ECE flag on its next ACK; the sender, seeing ECE, reduces its congestion window as if it detected loss and sets CWR on its next data segment to acknowledge it reacted; the receiver then stops setting ECE. So CWR — Congestion Window Reduced — means "I got your congestion echo and shrank my window." On Linux, tcp_ecn=1 enables ECN inbound and outbound, while tcp_ecn=2, the common default, accepts ECN when asked but doesn't initiate it.
eBPF
cBPF, classic BPF, is the original limited VM for packet filtering — what tcpdump compiles filters to, with a tiny instruction set, two registers, and no persistent state. eBPF, extended BPF, generalizes it into a full in-kernel VM: more registers, a richer instruction set, persistent state via maps, helper-function calls, and the ability to attach not just to packets but to syscalls, tracepoints, and kprobes. So cBPF filters packets; eBPF runs sandboxed programs across the whole kernel for networking, observability, and security. eBPF underpins modern tools like Cilium and Falco, while cBPF's legacy survives in tcpdump filter syntax.
Before loading any eBPF program, the verifier statically analyzes it to prove it's safe in kernel context. It checks the program terminates — historically by forbidding loops, now allowing bounded ones — so it can't hang the kernel; that every memory access is within checked bounds, so it can't touch arbitrary kernel memory; that registers are initialized and types tracked; and that it only calls permitted helpers. It walks all possible paths and rejects anything it can't prove safe, which is what lets the kernel run essentially untrusted internal programs. The catch, and a favorite interview point, is that the verifier itself is extremely complex, so bugs in it have been the root cause of privilege-escalation CVEs.
XDP, eXpress Data Path, is an eBPF hook running at the earliest point in the network stack — in the driver, before the kernel allocates a socket buffer. Because it processes packets before almost any kernel overhead, it makes drop-or-pass decisions at extremely high rates with minimal CPU. That's ideal for volumetric DDoS mitigation: an XDP program inspects each incoming packet and drops attack traffic — bad source ranges, malformed packets, flood patterns — at line rate before it costs the stack anything, while passing legitimate traffic up. Cloudflare and others use XDP to absorb massive floods on commodity hardware. It's the fast path; stateful or L7 logic falls back to higher hooks.
Falco is a runtime security tool whose eBPF program taps kernel events — mainly syscalls — from every process and container, streaming them to a userspace engine that evaluates rules like "a shell spawned in a container" or "a sensitive file was opened." eBPF gives it deep, low-overhead visibility without a custom kernel module, which is why it's the common Kubernetes runtime-detection choice. On verifier CVEs, the recurring root cause is the verifier mis-analyzing certain instruction sequences — CVE-2021-3490 was an ALU bounds-tracking flaw where it mis-tracked 32-bit bitwise ops, allowing out-of-bounds access and root escalation; others stem from speculative-execution mishandling or pointer-arithmetic tracking bugs. The theme is the verifier's complexity makes it a high-value attack surface, which is why hardening disables unprivileged eBPF via kernel.unprivileged_bpf_disabled=1.
eBPF maps are key-value structures giving programs persistent state and a channel to share data with userspace, since the programs themselves are short-lived per event. For security tooling the most useful are hash maps for per-process or per-connection state, per-CPU arrays for low-contention counters, ring buffers (or older perf event arrays) for efficiently streaming events up to a userspace agent like Falco, and LRU maps for bounded caches of recent activity. The pattern is that the in-kernel program observes events and updates a map, and the userspace component reads it for alerting, correlation, or enforcement. Maps are what turn stateless per-event hooks into stateful detection.