TCP/IP Deep Dive
Breadth layernotes-security-core-knowledge.md Also see: tcpdump & eBPF · TLS & Cryptography
IP Fundamentals
IP (Internet Protocol) provides connectionless, best-effort packet delivery across networks. It does not guarantee delivery, ordering, or error recovery — those are left to TCP.
IPv4 Header (key fields)
0 1 2 3 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 ├─────────────────────────────────────────────────────────────────┤ │ Version(4) │ IHL(4) │ DSCP/ECN(8) │ Total Length(16) │ ├─────────────────────────────────────────────────────────────────┤ │ Identification(16) │ Flags(3) │ Fragment Off(13) │ ├─────────────────────────────────────────────────────────────────┤ │ TTL(8) │ Protocol(8) │ Header Checksum(16) │ ├─────────────────────────────────────────────────────────────────┤ │ Source IP Address(32) │ ├─────────────────────────────────────────────────────────────────┤ │ Destination IP Address(32) │ └─────────────────────────────────────────────────────────────────┘
decremented by 1 at each hop; packet dropped when TTL reaches 0. Prevents routing loops. Typical starting values: Linux = 64, Windows = 128, Cisco = 255.
6 = TCP, 17 = UDP, 1 = ICMP, 50 = ESP (IPsec), 47 = GRE.
Don't Fragment (DF) — used for Path MTU Discovery; More Fragments (MF) — set on all fragments except the last.
Subnets and CIDR
192.168.1.0/24
↑
prefix length — how many bits are the network portion
192.168.1.0 = network address (all host bits 0)
192.168.1.255= broadcast address (all host bits 1)
192.168.1.1–192.168.1.254 = usable hosts (254 hosts)
/24 → 256 addresses, 254 usable
/25 → 128 addresses, 126 usable (half of /24)
/30 → 4 addresses, 2 usable (point-to-point links)
/32 → single host (loopback, static routes to one IP)
# Quick math: 2^(32-prefix) = total addresses
# /24 → 2^8 = 256 /16 → 2^16 = 65,536 /8 → 2^24 = 16,777,216TCP — The Three-Way Handshake
TCP is connection-oriented — before any data is exchanged, a connection must be established. The three-way handshake synchronises sequence numbers and confirms both endpoints can send and receive.
Client Server │ │ │─── SYN (seq=x) ──────────────►│ [1] Client picks random ISN x; sets SYN flag │ │ Server allocates connection state │◄── SYN-ACK (seq=y, ack=x+1) ─│ [2] Server picks random ISN y; ACK=x+1 (expecting x+1 next) │ │ │─── ACK (seq=x+1, ack=y+1) ──►│ [3] Client acknowledges server's ISN; connection ESTABLISHED │ │ │════ DATA EXCHANGE ════════════│ Both sides can now send data
Why three steps (not two)?Two steps would let the client know the server received its SYN, but the server would have no confirmation the client received the SYN-ACK. Three steps confirm both directions are working.
ISN (Initial Sequence Number)chosen randomly at connection start. Prevents sequence number prediction attacks (TCP hijacking, blind RST injection). Modern stacks use a CSPRNG for ISN selection.
SYN cookies (SYN flood protection, RFC 4987): instead of allocating state at step 1, the server encodes connection parameters into the ISN (SYN-ACK sequence number). State is only created when the client's ACK arrives with the matching sequence number. Defeats SYN floods because no memory is allocated until the three-way handshake completes.
TCP — Connection Establishment Variants
The standard three-way handshake is not the only way TCP connections form. RFC 793 itself defined several variants, and later RFCs added more.
1. Standard Three-Way Handshake (RFC 793)
The normal case — covered above. Client initiates; server responds.
Client: SYN(x) →
Server: ← SYN-ACK(y, ack=x+1)
Client: ACK(y+1) →2. Simultaneous Open (RFC 793, §3.4)
Both endpoints send SYN at the same time (each in SYN_SENT state). Neither is the "server." Results in four segments instead of three:
Host A Host B │ │ │── SYN(x) ───────────────►│ Both enter SYN_SENT │◄─────────────── SYN(y) ──│ │ │ │── SYN-ACK(x, ack=y+1) ──►│ Both transitions to SYN_RECEIVED │◄── SYN-ACK(y, ack=x+1) ─│ then ESTABLISHED │ │
Both sides end up sending a SYN-ACK in response to the other's SYN. Connection succeeds with 4 segments. This requires both endpoints to know each other's address and port in advance — rare in practice but valid per spec. You can observe it in P2P hole-punching (STUN/ICE).
3. Half-Open Connection
One side believes the connection is established (ESTABLISHED state); the other has crashed, rebooted, or has no record of the connection.
Server: ESTABLISHED (thinks connection is alive)
Client: rebooted, connection state lost
Client receives data from Server → sends RST (unknown connection)
Server receives RST → connection torn downDetectionss -tanp shows ESTABLISHED connections with no traffic for a long time. TCP keepalive (SO_KEEPALIVE) probes the peer periodically with empty ACK segments; if no response after N probes, the OS sends RST and closes the connection.
4. Half-Close
After sending FIN, a side can no longer send data but can still receive. The remote end may continue sending until it also sends FIN. Used in HTTP/1.0:
Client: FIN → (done sending request)
Server: ← data, data... (still sending response)
Server: ← FIN (done sending response)
Client: ACK →5. SYN Cookies (Stateless Handshake, RFC 4987)
A defence against SYN floods: the server doesn't allocate connection state until the handshake completes.
Client: SYN(x) →
Server: SYN-ACK(cookie, ack=x+1) ← cookie encodes: MSS, timestamp, HMAC(src/dst/port/isn)
NO state allocated yet
Client: ACK(cookie+1) →
Server: validates ACK sequence number matches cookie → creates socket, ESTABLISHEDThe cookie encodes MSS, window scale, and a timestamp HMAC. If the ACK never arrives (spoofed SYN), nothing was wasted.
Trade-offSYN options (SACK, window scaling, timestamps) can't be stored in the cookie — some features degrade gracefully under SYN cookies.
6. TCP Fast Open — TFO (RFC 7413, 2014)
TFO reduces latency for repeated connections by allowing data to be sent in the SYN packet itself on the second and subsequent connections to the same server.
First connection (TFO cookie request):
Client: SYN + TFO request option →
Server: ← SYN-ACK + TFO cookie (opaque server token)
Client: ACK →
(data exchange proceeds normally)
Subsequent connection (TFO in use):
Client: SYN + TFO cookie + HTTP GET → ← data piggybacked on the SYN!
Server: SYN-ACK + response data ← ← can reply before ACK arrives
Client: ACK →Benefitsaves one full RTT for short connections (e.g., HTTP/1.0, DNS over TCP, RPC). The server validates the cookie (HMAC of client IP, server key) before accepting the early data.
Security noteTFO data is sent before the three-way handshake completes, which has replay implications — replayed SYNs with the same cookie can re-deliver the same data on retransmit. RFC 7413 restricts early data to idempotent requests. Linux enables it via net.ipv4.tcp_fastopen.
7. TIME_WAIT Assassination and Avoidance
After closing, the active closer enters TIME_WAIT (2×MSL ≈ 60s) to absorb delayed packets. High-throughput servers accumulate many TIME_WAIT sockets.
Avoidance mechanisms
net.ipv4.tcp_tw_reuse=1— reuse TIME_WAIT sockets for new outgoing connections (if timestamps are enabled)SO_REUSEPORT— multiple sockets can bind the same port; kernel load-balances- RST-based early close: send RST instead of FIN+ACK — terminates immediately, no TIME_WAIT. Some load balancers use this.
TIME_WAIT assassination (RFC 1337)an attacker or delayed packet with an in-window sequence number can inject a RST into a TIME_WAIT connection, causing it to close early and potentially allowing a new connection to reuse the same 4-tuple — risking packet confusion. RFC 1337 describes this and recommends ignoring RSTs during TIME_WAIT.
TCP — Connection Termination (Four-Way)
TCP connections are full-duplex — each direction is closed independently. A full close requires four segments (two FINs + two ACKs):
Active Closer (A) Passive Closer (B) │ │ │─── FIN (seq=u) ────────────────────►│ [1] A: "I'm done sending" │ │ B enters CLOSE_WAIT (can still send data) │◄── ACK (ack=u+1) ────────────────── │ [2] B acknowledges A's FIN │ │ │ ← B may send more data here → │ B flushes remaining data │ │ │◄── FIN (seq=v) ──────────────────── │ [3] B: "I'm also done sending" │ │ │─── ACK (ack=v+1) ──────────────────►│ [4] A acknowledges B's FIN │ │ │ A enters TIME_WAIT (2×MSL = ~60s) │ B: CLOSED │ A: CLOSED after TIME_WAIT │
TIME_WAIT statethe active closer waits 2×MSL (Maximum Segment Lifetime, typically 60s) before fully closing. This ensures:
- The final ACK (step 4) reaches B — if it's lost, B retransmits FIN, A re-sends ACK
- Any delayed packets from the old connection expire before a new connection reuses the same 4-tuple
Half-closeafter A sends FIN (step 1), A can no longer send data, but B can still send data to A. This is used by protocols like HTTP/1.0 to signal end of request while waiting for the response.
RST (Reset)abrupt connection teardown — no graceful four-way. Used for error conditions (port not listening, firewall drop, connection timeout). No TIME_WAIT.
TCP Flags
The TCP header carries a 9-bit flags field — grown incrementally across three RFCs. Originally 6 flags (RFC 793, 1981), expanded to 8 for ECN (RFC 3168, 2001), then to 9 for the experimental Nonce Sum (RFC 3540, 2003). The bits that were once marked "reserved, must be zero" are now in active use — see the Shadow Bits section below.
| Flag | Bit | RFC | Meaning |
|---|---|---|---|
| NS | 0x100 | RFC 3540 | Nonce Sum — ECN-nonce extension; detects receivers that conceal ECN marks (experimental) |
| CWR | 0x080 | RFC 3168 | Congestion Window Reduced — confirms sender received ECE and reduced cwnd |
| ECE | 0x040 | RFC 3168 | ECN Echo — receiver signals network congestion (CE bit seen in IP header) |
| URG | 0x020 | RFC 793 | Urgent pointer field is significant — out-of-band data (rarely used) |
| ACK | 0x010 | RFC 793 | Acknowledgement field is valid — set on all segments after the first SYN |
| PSH | 0x008 | RFC 793 | Push — deliver buffered data to application immediately |
| RST | 0x004 | RFC 793 | Reset connection — abort immediately, no graceful close |
| SYN | 0x002 | RFC 793 | Synchronise sequence numbers — used in connection establishment |
| FIN | 0x001 | RFC 793 | No more data from sender — initiates graceful close |
Flag combinations in normal traffic
SYN — connection initiation (step 1)
SYN+ACK — connection response (step 2)
ACK — normal data and acknowledgement
PSH+ACK — data transfer (PSH tells receiver not to buffer)
FIN+ACK — graceful close
RST — reset/abort
RST+ACK — reset with acknowledgementScanning techniques (Nmap uses these)
SYN scan (-sS): send SYN; SYN-ACK = open, RST = closed, no-reply = filtered
NULL scan (-sN): no flags; RST = closed, no-reply = open|filtered (evades stateless firewalls)
FIN scan (-sF): only FIN; RST = closed, no-reply = open|filtered
XMAS scan (-sX): FIN+PSH+URG; RST = closed, no-reply = open|filtered
ACK scan (-sA): probe firewall rules; RST = unfiltered, no-reply = filteredTCP Shadow Bits — Flag Field RFC History
"Shadow bits" refers to the TCP header bits that were originally reserved and later repurposed. Older devices that enforce "reserved bits must be zero" (RFC 793) will reject or misinterpret modern ECN/NS traffic. IDS rules and firewalls from pre-2001 deployments may flag ECE/CWR/NS as anomalous.
Where the bits live in the wire format
The TCP header is 20 bytes minimum. Byte 12 (high nibble = data offset, low nibble = reserved nibble) and byte 13 (the flags byte) together form the flag field:
Byte 12: [ Data Offset (4 bits) ][ Reserved (3 bits) ][ NS ]
Byte 13: [ CWR ][ ECE ][ URG ][ ACK ][ PSH ][ RST ][ SYN ][ FIN ]
┌────────────────────────────────────────────────────────────┐
│ Bit: 15 14 13 12 │ 11 10 9 │ 8 │ 7 6 5 4 3 2 1 0 │
│ ← Data Off → │ ← Rsvd → │ NS │ CWR ECE URG ACK PSH RST SYN FIN│
└────────────────────────────────────────────────────────────┘
(nibble) (zeros) ^RFC3540^ ^RFC3168^ ^RFC 793^RFC Evolution Table
| RFC | Year | Change | Bits Used |
|---|---|---|---|
| RFC 793 | 1981 | TCP defined; upper bits "reserved, must be zero" | URG, ACK, PSH, RST, SYN, FIN |
| RFC 3168 | 2001 | ECN — Explicit Congestion Notification | + ECE (bit 6), CWR (bit 7) |
| RFC 3540 | 2003 | ECN-nonce — Nonce Sum (experimental) | + NS (bit 8, in reserved nibble) |
ECN — How CWR and ECE Work
ECN lets routers signal congestion before dropping packets:
[Negotiation — SYN handshake]
Client SYN: ECE + CWR set → "I want to use ECN"
Server SYN-ACK: ECE set only → "I also support ECN"
[Data Transfer]
Congested router sets CE bit in IP ECN field (not TCP)
└→ Receiver sees CE mark → sets ECE flag in next ACK
└→ Sender sees ECE → reduces cwnd → sets CWR to confirm
└→ Receiver clears ECE after receiving CWR-flagged segmentThis avoids a full packet drop cycle: cwnd reduction happens from the CE signal, not a retransmit timeout.
NS (Nonce Sum) — RFC 3540
The NS bit (bit 8, in the reserved nibble of byte 12) was designed to catch a cheating receiver:
- A well-behaved receiver that gets CE-marked packets signals them back via ECE
- A misbehaving receiver could silently drop CE marks to prevent the sender from reducing cwnd (gaining more throughput unfairly)
- RFC 3540 adds a one-bit running XOR checksum of random nonce bits embedded in each CE mark
- If the receiver is suppressing CE marks, the nonce sum will be wrong → sender detects cheating
Practical statusRFC 3540 is experimental and essentially unused in production. In practice you'll see NS only in:
- Academic research on ECN deployment
- Nmap
-OOS fingerprinting (custom stack behaviour) - Wireshark flag dissectors (it displays it)
- Custom packet crafters (Scapy, hping3): a TCP segment with NS set is unusual and may trigger IDS alerts
TCP State Machine
TCP endpoints move through states tracked in the kernel. ss -tanp or netstat -tanp shows current states.
CLOSED
│ SYN sent
▼
SYN_SENT ─────── SYN+ACK received ──► ESTABLISHED ─── FIN sent ──► FIN_WAIT_1
│ │
FIN received ACK received
│ │
CLOSE_WAIT FIN_WAIT_2
│ │
FIN sent FIN received
│ │
LAST_ACK TIME_WAIT
│ │ (2×MSL)
ACK recv CLOSED
│
CLOSEDKey states explained
LISTEN— server waiting for incoming connections (afterbind()+listen())SYN_RECEIVED— server received SYN, sent SYN-ACK, waiting for ACKESTABLISHED— data transfer phaseTIME_WAIT— waiting for all delayed packets to expire; shows up as many entries on busy servers
# View TCP state counts
ss -s
# or breakdown by state
ss -tan | awk '{print $1}' | sort | uniq -c | sort -rn
# TIME_WAIT accumulation — common on high-traffic servers
# Tune: net.ipv4.tcp_tw_reuse=1 (reuse TIME_WAIT sockets for new outgoing connections)
sysctl net.ipv4.tcp_tw_reuseTCP Reliability Mechanisms
Sequence Numbers and Acknowledgements
Every byte of data has a sequence number. The receiver ACKs the next byte it expects:
Sender: SEQ=1000, data=100 bytes → receiver expects ACK=1100
Sender: SEQ=1100, data=200 bytes → receiver expects ACK=1300
Receiver: ACK=1300 → "send me byte 1300 next"Sliding Window (Flow Control)
The receiver advertises a window size — how many bytes it can buffer. The sender can transmit up to window bytes without waiting for ACK. This enables pipelining.
Window size = 64KB → sender can have 64KB of unacknowledged data in flight
When window = 0, sender must pause (receiver's buffer is full)Congestion Control — Foundations
All TCP congestion control algorithms share the same two core mechanisms. Everything else is a variation on how aggressively to grow, and what to do when things go wrong.
Congestion window (cwnd)how many bytes the sender can have in-flight. The effective send rate = min(cwnd, receiver window) / RTT.
Two universal phases
cwnd starts at 1 MSS (Maximum Segment Size ≈ 1460 bytes for Ethernet), doubles each RTT until it hits the slow-start threshold (ssthresh) or packet loss occurs. Despite the name, exponential growth is fast — it reaches 64KB in ~6 RTTs.
after reaching ssthresh, grow cwnd linearly (+1 MSS per RTT). This is the steady-state phase.
Loss signals
nothing received in time → severe; the path may be broken
receiver got a gap but keeps ACKing the last good byte → a single packet is likely lost, not the whole path
TCP Tahoe (1988 — Van Jacobson, Lawrence Berkeley Lab)
The first modern congestion control algorithm. Before Tahoe, a TCP "congestion collapse" in 1986 reduced ARPANET throughput by a factor of 1000.
Loss event (ANY — timeout or triple-dup-ACK):
ssthresh = cwnd / 2
cwnd = 1 MSS ← back to slow start, hard reset
→ slow start phase again until ssthresh, then linearKey insightloss always means cwnd=1. No distinction between "one packet dropped" and "path is broken." This is conservative but correct.
TCP Reno (RFC 5681, 1990s — became the standard)
Reno adds Fast Recovery to avoid slamming cwnd back to 1 for a single loss.
Triple duplicate ACK (partial loss):
ssthresh = cwnd / 2
cwnd = ssthresh ← cut in half, NOT back to 1
→ enter Fast Recovery: send missing segment, inflate cwnd for each dup-ACK
→ when new ACK arrives (hole filled), exit Fast Recovery, continue from ssthresh
Timeout (severe loss):
ssthresh = cwnd / 2
cwnd = 1 MSS ← same as Tahoe for timeouts
→ slow startReno vs Tahoefor a single dropped packet detected by 3 dup-ACKs, Reno halves cwnd but stays in flight; Tahoe resets to 1. Reno is far more efficient on links with occasional random loss.
TCP New Reno (RFC 6582, 1999)
Reno has a bug: if multiple packets are lost in a single window, each loss is only discovered one at a time (each "partial ACK" advances the window by one segment). New Reno fixes this:
During Fast Recovery, if a partial ACK arrives (not a full new ACK):
→ retransmit the next unacknowledged segment immediately
→ deflate cwnd by 1 MSS (don't exit Fast Recovery yet)
→ continue until a full new ACK fills the whole holeThis means New Reno can recover from multiple losses in one window without triggering a timeout.
TCP SACK — Selective Acknowledgement (RFC 2018, 1996)
SACK is an extension to the ACK mechanism, not a standalone algorithm. It works alongside Reno or New Reno.
Without SACK, the receiver can only say "I've received everything up to byte N." With SACK:
Receiver sends: ACK=1000, SACK: [2000-3000, 4000-5000]
Meaning: "I'm missing 1000-2000 and 3000-4000; I have 2000-3000 and 4000-5000 already"
Sender retransmits only the missing blocks — not the full windowNegotiated in SYN/SYN-ACK via the SACK-permitted option. Nearly universally supported today. Dramatically improves recovery from bursty loss (e.g., buffer overflow dropping multiple packets at once).
TCP Vegas (1994 — University of Arizona)
Vegas is the first delay-based algorithm. Instead of waiting for packet loss, it uses RTT increase as the congestion signal.
Baseline RTT = minimum observed RTT (the "uncongested" RTT)
Expected throughput = cwnd / BaseRTT
Actual throughput = cwnd / current_RTT
Diff = Expected - Actual
if Diff < α: increase cwnd (we have headroom)
if Diff > β: decrease cwnd (queues are building)Advantageproactive — reduces cwnd before packets are dropped, keeping queues shorter (lower latency). Disadvantage: unfair against loss-based algorithms (Reno). Vegas backs off when it senses queue growth; Reno keeps pushing until loss; Reno gets more bandwidth.
TCP CUBIC (RFC 8312, 2008 — Linux default since kernel 2.6.19)
CUBIC replaces the linear growth of Reno with a cubic function of time since the last loss. This allows fast recovery of bandwidth after a congestion event while remaining stable near the saturation point.
cwnd(t) = C × (t - K)³ + W_max
where:
t = time since last congestion event
W_max = cwnd just before the last loss (the "maximum")
K = time it will take to reach W_max again (K = ∛(W_max × β / C))
C = scaling constant (default 0.4)
β = multiplicative decrease factor (default 0.7 — reduces cwnd to 70% on loss)Shape of the curve
cwnd │ . ← W_max (congestion point) │ . . │ . . │ . . │ . (concave: fast initial recovery) │ . └──────────────────────── time K (time to reach W_max)
- Concave phase (before W_max): aggressive recovery, catches up quickly
- Inflection at W_max: slows down to probe carefully at the previous congestion point
- Convex phase (beyond W_max): probes for more bandwidth slowly
Why it's better for high-bandwidth, high-latency links (e.g., 10Gbps transoceanic): Reno's linear growth (+1 MSS/RTT) takes forever to fill a large pipe after a loss. CUBIC's cubic growth fills it much faster.
TCP BBR — Bottleneck Bandwidth and RTT (Google, 2016)
BBR is a fundamentally different design philosophy. All prior algorithms react to loss or delay. BBR models the network directly by estimating two parameters:
bottleneck bandwidth (max observed delivery rate)
round-trip propagation delay (minimum observed RTT, the "empty pipe" delay)
Optimal operating point: cwnd = BtlBw × RTprop (fill the pipe exactly, no queueing)BBR's probing cycle
1. STARTUP: exponential growth until bandwidth estimate stops increasing (same as slow start)
2. DRAIN: quickly drain queue built during STARTUP
3. PROBE_BW: cycle through 8 RTTs at various rates to continuously re-estimate BtlBw
4. PROBE_RTT: every 10 seconds, reduce cwnd to 4 MSS to get a clean RTprop measurementWhy BBR outperforms on high-BDP (Bandwidth-Delay Product) links:
- Loss-based algorithms fill the queue → high latency. BBR targets the empty-pipe point.
- Lossy links (satellite, mobile): Reno/CUBIC interpret every loss as congestion; BBR uses bandwidth as the signal and ignores sporadic loss.
Why BBR can be unfair in mixed deployments:
- BBR v1 was sometimes more aggressive than CUBIC in multi-tenant environments
- BBR v2 (2019) adds a loss signal component to improve fairness
# Check current congestion control on Linux
sysctl net.ipv4.tcp_congestion_control
# Set to BBR
sysctl -w net.ipv4.tcp_congestion_control=bbr
sysctl -w net.core.default_qdisc=fq # BBR works best with fair queueing
# List available algorithms
sysctl net.ipv4.tcp_available_congestion_controlCongestion Control Algorithm Comparison
| Algorithm | Loss Signal | Delay Signal | Default On | Best For |
|---|---|---|---|---|
| Tahoe | Loss → cwnd=1 | No | Legacy only | — |
| Reno (RFC 5681) | Loss → cwnd/2 | No | Widely supported | General baseline |
| New Reno (RFC 6582) | Multi-loss recovery | No | Many OS | General use |
| SACK (RFC 2018) | Selective retransmit | No | Extension to Reno/NR | High-loss links |
| Vegas | RTT increase | Yes | Rarely | Low-latency LANs |
| CUBIC (RFC 8312) | Loss → cubic | No | Linux default | High-BDP paths |
| BBR v1/v2 | Bandwidth model | Yes (RTprop) | Google/options | Lossy/high-BDP |
UDP — When Not to Use TCP
UDP is connectionless — no handshake, no ordering, no guaranteed delivery, no congestion control.
| Property | TCP | UDP |
|---|---|---|
| Connection setup | 3-way handshake | None |
| Ordering | Guaranteed | Not guaranteed |
| Reliability | Retransmission on loss | No; application must handle |
| Congestion control | Yes (reduces send rate) | No |
| Overhead | ~20 byte header | 8 byte header |
| Latency | Higher (RTT for handshake) | Lower |
| Use cases | HTTP, SSH, SMTP, file transfer | DNS, DHCP, streaming, gaming, VoIP |
When to choose UDP
- Latency is critical and some loss is tolerable (gaming, VoIP, live streaming)
- Application has its own reliability (QUIC re-implements reliability over UDP for HTTP/3)
- Request-response that fits in one datagram and the client retries on timeout (DNS, DHCP, NTP, SNMP)
ICMP
ICMP (Internet Control Message Protocol) is used for network diagnostics and error reporting. Not a transport protocol — it rides directly on IP (protocol 1).
| Type | Code | Name | Use |
|---|---|---|---|
| 0 | 0 | Echo Reply | ping response |
| 3 | 0 | Dest Unreachable: Net | routing failure |
| 3 | 1 | Dest Unreachable: Host | host unreachable |
| 3 | 3 | Dest Unreachable: Port | port closed (TCP RST equivalent for UDP) |
| 3 | 4 | Fragmentation Needed | Path MTU Discovery (DF bit set but packet too large) |
| 8 | 0 | Echo Request | ping |
| 11 | 0 | Time Exceeded | TTL expired (traceroute uses this) |
# ping — sends ICMP Echo Request, measures RTT
ping -c 4 8.8.8.8
# traceroute — sends probes with increasing TTL; each hop returns ICMP Time Exceeded
traceroute google.com # Linux (UDP by default)
traceroute -I google.com # Linux (ICMP, like Windows)
tracert google.com # Windows
# How traceroute works:
# TTL=1 → first router decrements to 0 → returns ICMP Time Exceeded → reveals first hop's IP
# TTL=2 → second router → reveals second hop → etc.ICMP Redirecta router sends ICMP Redirect (type 5) to tell a host to use a better route. Historically used in on-path attacks to redirect traffic.
Common Port Numbers
# Well-known ports (0-1023) — require root/admin to bind
20, 21 FTP (data, control) 22 SSH
23 Telnet 25 SMTP
53 DNS 67, 68 DHCP (server, client)
80 HTTP 110 POP3
119 NNTP 123 NTP
143 IMAP 161, 162 SNMP (agent, trap)
179 BGP 389 LDAP
443 HTTPS 445 SMB
465 SMTPS 514 Syslog (UDP)
587 SMTP submission 636 LDAPS
993 IMAPS 995 POP3S
# Registered ports (1024-49151)
1433 MSSQL 1521 Oracle
3306 MySQL 3389 RDP
5432 PostgreSQL 5900 VNC
6379 Redis 8080 HTTP alternate
8443 HTTPS alternate 9200 Elasticsearch
27017 MongoDB
# Dynamic/ephemeral (49152-65535) — client source portsNetwork Address Translation (NAT)
NAT allows many private IP addresses to share one public IP. The NAT device (router/firewall) rewrites source/destination IP and port on the fly.
Internal: 192.168.1.10:54321 → 8.8.8.8:53
NAT table entry: (192.168.1.10:54321) ↔ (203.0.113.1:40001) [public IP:mapped port]
Packet on wire: 203.0.113.1:40001 → 8.8.8.8:53
Return packet: 8.8.8.8:53 → 203.0.113.1:40001
NAT reverse-translates: → 192.168.1.10:54321NAT types (relevant for P2P/VoIP/gaming)
any external host can send to the mapped port
only hosts the internal client has contacted can reply
only the specific IP:port combination that was contacted can reply
different external port for each destination — hardest for P2P; requires STUN/TURN
Security implicationNAT provides implicit inbound filtering (no unsolicited inbound connections reach internal hosts). Not a firewall — it does not filter packets or detect attacks.
TCP/IP Security Attacks
SYN Flood
Attacker sends many SYN packets with spoofed source IPs. Server allocates state for each and sends SYN-ACK into the void, exhausting the backlog queue.
# Each SYN consumes a slot in the SYN backlog (net.ipv4.tcp_max_syn_backlog)
# Default: 1024-2048; large floods exhaust this in seconds
Mitigations:
- SYN cookies (stateless; no memory allocated until ACK)
- net.ipv4.tcp_syncookies=1 (Linux sysctl)
- Rate limiting SYN packets (iptables -m limit, cloud WAF)
- Anycast-based DDoS scrubbing (Cloudflare, Akamai)TCP Session Hijacking
If an attacker can predict sequence numbers (pre-CSPRNG ISNs) or observe traffic, they can inject segments into an existing TCP session.
Attacker observes: SEQ=1000, ACK=5000 between A and B
Attacker crafts: src=A, seq=1000, ACK=5000, data=malicious command
Server accepts it as coming from A → command executedModern mitigations: randomised ISNs; TLS (TCP hijacking is meaningless against encrypted application data).
RST Injection
Send a forged TCP RST with a valid sequence number → immediately terminates the connection. Used by the Great Firewall of China to terminate connections to blocked sites.
Blind RST requires guessing the sequence number — made harder by randomised ISNs and PAWS (Protection Against Wrapped Sequences).
IP Spoofing
Forge the source IP in IP packets. Easy for UDP (stateless); harder for TCP (attacker can't receive SYN-ACK responses so can't complete the handshake with a spoofed source unless on the same L2 network or using reflected/amplified attacks).
# Amplification DDoS example: spoof victim's IP as source
# Send small request to NTP server → large response → directed at victim
# Amplification factor: ~556x for NTP monlist
hping3 -S --flood -V --spoof <victim_IP> -p 123 <ntp_server>Packet Analysis with tcpdump
Sequence Number Display — A Common Gotcha
By default, tcpdump shows relative sequence numbers — it subtracts the ISN so the first byte always appears as sequence 0. This makes output readable but diverges from what's actually on the wire.
Default output (relative):
10:00:00.000000 IP client.54321 > server.80: Flags [S], seq 0, win 65535
10:00:00.001000 IP server.80 > client.54321: Flags [S.], seq 0, ack 1, win 65535
← Both sides appear to start at seq=0 — NOT what's on the wire
Actual wire (absolute):
client ISN = 1234567890
server ISN = 9876543210Use -S to see absolute sequence numbers — essential when:
- Correlating tcpdump output with Wireshark captures or kernel logs
- Debugging sequence number prediction attacks or RST injection
- Comparing what two different hosts see on the same connection
- Writing tcpdump BPF filters that match specific sequence ranges
tcpdump -i eth0 -S port 80 # -S = absolute (real) sequence numbers
# Now you'll see seq=1234567890 instead of seq=0Flag display also has quirksBy default, tcpdump abbreviates flags in the Flags [...] field:
[S] = SYN
[S.] = SYN-ACK (dot = ACK bit)
[.] = ACK only
[P.] = PSH+ACK
[F.] = FIN+ACK
[R] = RST
[R.] = RST+ACKECN flags (ECE, CWR) and NS are not shown in the default abbreviated format. To see them, use -v (verbose) which switches to the numeric flags representation, or use Wireshark/tshark for full flag visibility.
# Verbose mode — shows all flags including ECN
tcpdump -i eth0 -v port 80
# Example verbose output for a SYN with ECN negotiation:
# Flags [SEC], cksum 0x1234 (correct), seq 0, win 65535, options [...]
# SEC = SYN + ECE + CWR (client requesting ECN)Common tcpdump Commands
# Capture all traffic on eth0
tcpdump -i eth0
# Capture with hex + ASCII output
tcpdump -i eth0 -XX
# Write to file, read back
tcpdump -i eth0 -w capture.pcap
tcpdump -r capture.pcap
# Filter by host and port
tcpdump -i eth0 host 10.0.0.1 and port 80
# TCP flags filters
tcpdump -i eth0 'tcp[tcpflags] & tcp-syn != 0' # SYN packets (connection attempts)
tcpdump -i eth0 'tcp[tcpflags] & tcp-rst != 0' # RST packets (resets)
tcpdump -i eth0 'tcp[tcpflags] == tcp-syn' # only SYN (not SYN-ACK)
tcpdump -i eth0 'tcp[tcpflags] == (tcp-syn|tcp-ack)'# SYN-ACK
# Show full TCP handshakes (SYN, SYN-ACK, ACK)
tcpdump -i eth0 -n 'tcp[tcpflags] & (tcp-syn|tcp-fin|tcp-rst) != 0'
# ICMP only
tcpdump -i eth0 icmp
# DNS queries
tcpdump -i eth0 -n 'udp port 53'
# Show absolute sequence numbers (not relative) — ALWAYS use this when debugging
tcpdump -i eth0 -S port 80
# Show all flags including ECN/NS
tcpdump -i eth0 -v port 80
# Count packets per source IP (SYN flood detection)
tcpdump -i eth0 -n 'tcp[tcpflags] == tcp-syn' | awk '{print $3}' | cut -d. -f1-4 | sort | uniq -c | sort -rnLong-Distance and Space Communication Protocols
TCP/IP was designed for terrestrial networks where RTT is measured in milliseconds and packet loss is rare. These assumptions collapse completely in satellite and deep-space contexts. This section covers why TCP fails at distance and what protocols replaced it.
Why TCP Breaks at Long Distance
The core problem: propagation delay
Speed of light in vacuum: ~300,000 km/s Speed of light in fibre: ~200,000 km/s (refractive index ≈ 1.5) Distance One-way delay RTT ────────────────────────────────────────── New York → London ~28ms ~56ms (TCP works fine) Geostationary orbit ~240ms ~600ms (TCP struggles) Moon ~1.3s ~2.6s (TCP barely works) Mars (closest) ~3 min ~6 min (TCP useless) Mars (farthest) ~22 min ~44 min (TCP completely unusable)
TCP's cwnd problem at high latency
Even with RFC 7323 window scaling (up to ~1GB window), TCP's congestion control defeats itself:
Bandwidth-Delay Product (BDP) = Bandwidth × RTT
10 Mbps × 44 min RTT = 3.3 GB BDP — must keep 3.3 GB in flight to saturate the link
After a single packet loss on Mars link:
CUBIC: cwnd → 70% of 3.3GB → sender immediately cuts throughput
Reno: cwnd → 50% of 3.3GB → much worse
Then: 44 minutes before sender learns the retransmit succeededAdditional problems in space
| Problem | Why it matters for TCP |
|---|---|
| High Bit Error Rate (BER) | Space radiation, antenna misalignment → packet loss that is NOT congestion; TCP wrongly cuts cwnd |
| Intermittent connectivity | Spacecraft orbits behind a planet; ground station windows limited to hours/day; TCP connection drops |
| Asymmetric bandwidth | High-rate downlink (telemetry), low-rate uplink (commands); TCP ACKs compete with uplink data |
| Single-copy links | One RF link to spacecraft; no alternate path; no fast retransmit because there is no competing traffic |
Satellite Communication (Near-Earth)
| Orbit | Examples | Altitude | Round trip | What it means for TCP |
|---|---|---|---|---|
| LEO (Low Earth) | Starlink, OneWeb | ~550 km | ~20–40 ms | Like a cross-continent fibre link. Works well; prefer BBR or CUBIC over Reno so radio bit errors aren't mistaken for congestion (net.ipv4.tcp_congestion_control=bbr) |
| MEO (Medium Earth) | O3b mPOWER; GPS | ~2,000–35,000 km (GPS at 20,200 km) | ~100–150 ms for data constellations | Moderate delay. GPS carries no user data at all — it only broadcasts navigation signals |
| GEO (Geostationary) | HughesNet, Viasat | 35,786 km | ~600 ms (≈240 ms each way + processing) | Slow start needs many round trips to fill the pipe; throughput collapses without help |
TipHow GEO providers cope: Performance Enhancing Proxies (PEPs). The provider splits each TCP connection at the ground station and the satellite gateway, and runs its own protocol (often UDP-based, with huge windows) across the satellite hop. Your machine sees normal TCP; the long-delay segment never runs TCP at all. Security side effect: a PEP must see TCP headers, which is one reason satellite links and end-to-end encrypted transports like QUIC don't always get along.
CCSDS — The Space Data Standard
The Consultative Committee for Space Data Systems (CCSDS) — established 1982, consortium of NASA, ESA, JAXA, and 28 other agencies — defines the standard protocols for spacecraft communications. TCP/IP is not used between ground and spacecraft for commanding and telemetry.
CCSDS stack (simplified)
Application Layer
│
CCSDS Space Packet Protocol ← analogous to IP; fixed-length primary header
│
CCSDS Transfer Frame ← synchronisation, error detection (Reed-Solomon / LDPC)
│
RF Link (X-band, Ka-band, optical)Key CCSDS protocols
| Protocol | Direction | Use |
|---|---|---|
| Telecommand (TC) | Ground → Spacecraft | Sending commands; low data rate (~2 kbps); protected by BCH error-correcting codes |
| Telemetry (TM) | Spacecraft → Ground | Sensor data, housekeeping, science; 100 bps to 100 Mbps depending on mission |
| CFDP — File Delivery Protocol | Both | Reliable file transfer over intermittent links; supports store-and-forward; used for software uplinks to spacecraft |
| Proximity-1 | Short range | Rover ↔ lander ↔ orbiter relay (Mars rovers use this to relay via MRO orbiter) |
Deep Space Network (DSN)
NASA's Deep Space Network is the ground infrastructure for communicating with spacecraft beyond ~2 million km. Three sites 120° apart for continuous coverage:
Goldstone, California (USA) — 34m + 70m dish antennas
Madrid, Spain — 34m + 70m antennas
Canberra, Australia — 34m + 70m antennasFrequency bands used
legacy, lower data rate, penetrates ionosphere better
primary deep space band; ~6–10 Mbps downlink for near-Earth; ~1–4 kbps for outer planets
high bandwidth; atmosphere attenuates more; used for high-rate science data (Cassini, Mars Reconnaissance Orbiter)
next-generation; Lunar Laser Communication Demo (2013) achieved 622 Mbps from lunar orbit; TBIRD (2022) demonstrated 200 Gbps from LEO
DTN — Delay/Disruption Tolerant Networking
DTN is the protocol architecture designed specifically for deep space and other challenged networks. The key insight: store-and-forward, treating the network as a series of custody transfers rather than an end-to-end connection.
Bundle Protocol (RFC 5050, 2007; RFC 9171, 2022):
Traditional TCP model:
Source → Router1 → Router2 → Destination (all links must be up simultaneously)
DTN Bundle Protocol model:
Source → [store at Node A until contact window opens]
→ [transfer to Node B, which stores until its contact window]
→ [transfer to Destination]
Each hop transfers custody; next hop doesn't need to exist yetThink of it like store-and-forward email, but for arbitrary data at the transport layer.
Bundle Protocol header fieldscreation timestamp (for ordering), lifetime (TTL for the bundle), custody transfer flags, source/destination EIDs (Endpoint IDs, URI-like: dtn://voyager1/telemetry).
LTP — Licklider Transmission Protocol (RFC 5326):
LTP provides reliable data transfer for a single deep-space radio link. Unlike TCP, it is designed around long delays:
- Uses negative ACKs (NAKs) for specific lost segments rather than streaming ACKs
- No connection setup — LTP sessions are unidirectional
- Timers are set to multiples of the one-way light-time (not RTT as in TCP)
- Used as the convergence layer beneath Bundle Protocol for deep-space links
DeploymentDTN/Bundle Protocol is deployed on the International Space Station (experimental, since ~2008), and is the planned networking architecture for NASA's LunaNet (lunar internet infrastructure for Artemis program).
SCPS-TP — Space Communications Protocol Specifications (Transport Protocol)
SCPS-TP is a modified TCP designed for space. Developed by the CCSDS for use cases where IP-based networking is desirable but standard TCP performs poorly.
Key changes from TCP
receiver explicitly reports missing segments (similar to SACK but also for NAKs)
uses RTT increase rather than loss as the congestion signal — avoids penalising for BER-induced loss
fewer ACKs on asymmetric links (conserves uplink bandwidth)
SCPS-TP uses header flags to distinguish "packet lost to corruption" from "packet lost to congestion" — prevents cwnd reduction on corrupted links
Used in some US military satellite systems. Not widely deployed commercially.
Summary: Which Protocol to Use
| Scenario | Protocol | Why |
|---|---|---|
| LEO satellite (Starlink) | TCP/IP + BBR | RTT is low enough; BBR handles BER loss without congestion backoff |
| GEO satellite (HughesNet) | TCP at endpoints, PEP at satellite gateway | 600ms RTT kills TCP performance; PEP hides the satellite hop |
| Ground to spacecraft | CCSDS (TM/TC/CFDP) | Designed for intermittent contact, asymmetric rates, error correction |
| Deep space file transfer | CFDP (over DTN) | Store-and-forward, custody transfer, works with contact gaps |
| Rover to orbiter to Earth | Proximity-1 + Bundle Protocol | Multi-hop relay, contact windows, disruption tolerance |
| ISS / LunaNet | Bundle Protocol (RFC 9171) | DTN architecture for challenged networks |
Interview Questions
Connection Establishment and State
Three steps are needed to confirm both directions work. Step 1 (SYN): client proves it can send. Step 2 (SYN-ACK): server proves it can receive and send, and ACKs the client's ISN. Step 3 (ACK): client proves it can receive, and ACKs the server's ISN. Two steps would leave the server not knowing if its SYN-ACK was received — meaning the server would commit resources (socket, buffers) with no confirmation the client is actually there.
Both endpoints send SYN at the same time, resulting in a 4-segment exchange (SYN, SYN, SYN-ACK, SYN-ACK) rather than 3. Both start in SYN_SENT and transition to SYN_RECEIVED then ESTABLISHED. It's rare in practice but valid per RFC 793 §3.4 — occurs in P2P hole-punching (STUN/ICE) where both peers initiate simultaneously.
TCP is full-duplex, so each direction closes independently — 4 segments: FIN, ACK, FIN, ACK. TIME_WAIT (2×MSL ≈ 60s) serves two purposes: (1) ensures the final ACK reaches the other side — if it's lost, the other side retransmits FIN and we re-send ACK; (2) ensures any delayed packets from the old connection expire before a new connection reuses the same 4-tuple, preventing old data from confusing a new session.
TFO (RFC 7413) eliminates the RTT cost of the handshake for repeat connections. On first connect, the server issues a cookie (HMAC of the client IP). On subsequent connects, the client includes the cookie + data in the SYN. The server validates the cookie and can begin processing the request before the handshake completes — saving one full RTT. Useful for short connections (DNS-over-TCP, RPC). Risk: TFO data on SYN can be replayed by retransmits, so it should only carry idempotent requests.
Attacker sends SYNs with spoofed source IPs. The server allocates a connection slot (SYN backlog entry) for each and sends SYN-ACK into the void, exhausting net.ipv4.tcp_max_syn_backlog. SYN cookies (RFC 4987) fix this by not allocating state: the server encodes MSS + timestamp + HMAC into the ISN of the SYN-ACK. If the ACK arrives with the right value, state is created then. Spoofed SYNs never produce a valid ACK, so nothing is wasted.
CLOSE_WAIT means the remote end sent FIN (peer is done sending) and the local application acknowledged it — but the local application has not yet called close(). The application is still holding the socket open. This is a bug: a connection leak. Common cause: a thread waiting on the socket without checking for EOF, or a resource management bug in the application. Fix: find the process with ss -tanp | grep CLOSE_WAIT and investigate why it's not closing.
Attacker forges a TCP RST segment with a valid sequence number in the current window. The receiver closes the connection immediately — no graceful teardown, no TIME_WAIT. Used by the Great Firewall of China to terminate connections to blocked sites. Also used in BGP session disruption attacks (target: TCP port 179). Defence: RFC 5961 "blind in-window attacks" requires the RST to match the next expected sequence number exactly (not just be in-window); this is now the default on modern Linux/Windows.
TCP Flags and Shadow Bits
Both were added in RFC 3168 (2001) for ECN — Explicit Congestion Notification. ECE (ECN Echo, bit 6): the receiver sets this when it has seen a CE (Congestion Experienced) mark in the IP header, telling the sender the network is congested. CWR (Congestion Window Reduced, bit 7): the sender sets this to acknowledge it received ECE and has already reduced cwnd — prevents the receiver from continuing to set ECE for the same event. They allow congestion signalling without dropping packets.
NS (Nonce Sum) was added in RFC 3540 (2003) as an experimental extension to ECN. It lives in the previously-reserved nibble of byte 12, making the effective flag field 9 bits. Its purpose: detect a misbehaving receiver that suppresses ECN marks to avoid triggering cwnd reduction (gaining unfair bandwidth). The sender embeds random nonces in CE marks; the receiver's NS bit is a running XOR of those nonces. If marks are being hidden, the XOR value will be wrong. RFC 3540 is experimental and essentially unused in production.
RFC 793 (1981) marked all bits above the 6 original flags as "reserved, must be zero." Firewall rules and IDS signatures written before RFC 3168 (2001) were written to alert on non-zero reserved bits because that was a reliable indicator of a custom/malformed packet. ECN flipped those bits into active use, so old rules generate false positives on all ECN-capable connections. The fix is to update rules to specifically allow ECE/CWR/NS and only flag unexpected combinations.
Congestion Control
Both detect congestion via triple duplicate ACKs or timeout. The difference is what they do when triple-dup-ACK triggers fast retransmit. Tahoe: always sets cwnd=1 MSS (returns to slow start) — treats triple-dup-ACK the same as timeout. Reno (RFC 5681): adds fast recovery — on triple-dup-ACK, sets cwnd = ssthresh = old_cwnd/2 and enters fast recovery, staying near the throughput curve. Only on timeout does Reno also reset to cwnd=1. Reno is far more efficient for single random packet loss.
Reno has a partial ACK problem: if multiple packets are lost in one window, fast recovery exits on the first "new ACK" (which may only advance the window past the first lost packet). Reno then enters congestion avoidance with the remaining losses unresolved, eventually timing out. New Reno (RFC 6582) stays in fast recovery after a partial ACK — it retransmits the next unacknowledged segment and continues until a full new ACK fills the entire hole. This recovers from multiple per-window losses without triggering RTO.
Standard ACK is cumulative — it only says "I've received everything up to byte N." With SACK (RFC 2018), the receiver can report non-contiguous received blocks: "I'm missing 1000-2000, but I have 2000-3000 and 4000-5000." The sender retransmits only the missing blocks. This is critical when multiple packets are dropped in a burst — without SACK, the sender must infer losses one at a time (one per RTT). SACK is negotiated in the SYN/SYN-ACK options and is nearly universal today.
CUBIC (RFC 8312, Linux default since 2.6.19) replaces Reno's linear cwnd growth with a cubic function of time since the last congestion event. On loss, it reduces cwnd to 70% (β=0.7 vs Reno's 50%). After recovery, it grows quickly (concave phase), then slows down near the previous congestion point, then explores above it slowly (convex phase). On a 10Gbps transatlantic link with RTT=100ms, Reno needs thousands of RTTs to fill the pipe after a loss; CUBIC's cubic growth fills it much faster. The cubic function is also RTT-independent, so CUBIC is fair between long and short RTT flows.
BBR (Google, 2016) doesn't react to loss or delay — it models the network. It estimates two values: BtlBw (bottleneck bandwidth — max sustained delivery rate) and RTprop (minimum observed RTT — the "empty pipe" delay). The target cwnd is BtlBw × RTprop — exactly enough to fill the pipe without building a queue. BBR cycles through probing phases to keep these estimates fresh. On lossy links (satellite, LTE), Reno/CUBIC interpret every drop as congestion; BBR ignores sporadic loss and uses bandwidth as the signal. BBR v1 was unfair in some multi-tenant setups; BBR v2 (2019) adds a loss signal for fairness.
General
TCP: connection-oriented (3-way handshake), ordered, reliable (retransmit on loss), congestion-controlled, higher overhead (~20 byte header + RTT setup). UDP: connectionless, unordered, unreliable, no congestion control, 8 byte header. Choose UDP when: (1) latency beats reliability — gaming, VoIP, live streaming tolerate loss but not delay; (2) app has its own reliability — QUIC implements reliability over UDP for HTTP/3; (3) request/response fits in one datagram and the client retries on timeout — DNS, DHCP, NTP, SNMP.
Every byte of data has a sequence number. The ISN is the starting number for a new connection. Randomisation (CSPRNG) prevents two attacks: (1) TCP session hijacking — a predictable ISN lets an off-path attacker forge segments with the right sequence number; (2) blind RST injection — guessing an in-window sequence number to terminate connections. Modern stacks also use ISN clocks with per-connection secrets (RFC 6528) rather than purely random ISNs, to prevent sequence number reuse within TIME_WAIT.
Traceroute discovers the path to a destination by exploiting TTL (Time To Live). It sends probes with TTL=1, TTL=2, TTL=3, … Each router that receives a packet with TTL=0 drops it and returns an ICMP Time Exceeded (type 11) message — revealing its IP address. Linux default: UDP probes to high ports. Linux -I: ICMP Echo. Windows tracert: ICMP Echo. The path may differ per probe (ECMP routing), and some hops suppress ICMP Time Exceeded (show as * * *).
NAT rewrites source/destination IP and port on packets to allow many private addresses to share one public IP. The NAT table maps internal (IP:port) ↔ external (IP:mapped-port). It provides implicit inbound filtering — unsolicited inbound connections can't reach internal hosts because there's no NAT table entry. But NAT is not a firewall: it doesn't inspect packet content, doesn't detect attacks, doesn't enforce policies. A firewall behind NAT is still needed for: blocking outbound to malicious IPs, detecting port scans, layer-7 inspection, and explicit inbound allow-rules.
tcpdump shows relative sequence numbers by default — it subtracts the ISN so the first byte appears as 0. This makes output more readable but hides the actual wire values. Use -S (absolute/full sequence numbers) to see the real ISNs. This matters for: correlating with Wireshark or kernel logs, debugging sequence number attacks, comparing captures from two hosts. Also note: ECN flags (ECE, CWR) and NS are not shown in the default abbreviated Flags [...] format — use -v (verbose) to see all 9 flag bits.