Security Notes
System Design

Design a URL Shortener (Bitly)

4 min read 6 sections

DifficultyEasy | HelloInterview: problem breakdown


Problem Statement

Build a service that takes a long URL and returns a short alias (e.g., https://bit.ly/3xYz9a). When users visit the short URL, they are redirected to the original URL.

๐Ÿ“–

Real-world & security angle: URL shorteners are a deceptively rich security topic. Short codes must be unguessable โ€” early shorteners used sequential IDs, so you could enumerate /1, /2, โ€ฆ and harvest everyone's "private" links (people shortened Google Docs, Dropbox shares, even password-reset URLs). Researchers in 2016 enumerated Microsoft and Google short links and found live OneDrive/Maps data. Two more security gotchas interviewers love: shorteners are a favorite open-redirect / phishing laundering tool (the short domain looks trusted, the destination is malicious), and the 301 vs 302 choice has a security/ops consequence โ€” a 301 is cached by the browser forever, so you can't change or revoke the destination later, and you lose click analytics. The lesson: use 302 + random (not sequential) codes, and scan destinations for malware.


Requirements

Functional

  • Shorten a long URL โ†’ unique short code
  • Redirect short URL โ†’ original URL (HTTP 301 or 302)
  • Custom aliases (optional)
  • Expiry for short URLs (optional)
  • Analytics: click counts per short URL (optional)

Non-Functional

  • Reads heavily outweigh writes (100:1 read/write ratio)
  • Low redirect latency (< 10ms P99) โ€” this is in the critical path of every page load
  • 100M URLs shortened total; 10B redirects/day
  • High availability: 99.99% uptime for redirects

Scale Estimation

  • Write: 100M URLs / (365 ร— 86400) โ‰ˆ 3 writes/s
  • Read: 10B / 86400 โ‰ˆ 115,000 redirects/s
  • Storage: 100M URLs ร— 500 bytes avg โ‰ˆ 50 GB (easily fits in one DB)

Core Design

Client โ†’ CDN/Edge Cache โ†’ Load Balancer โ†’ Redirect Service โ†’ Cache (Redis) โ†’ DB
                                         โ†“
                                     Write Service โ†’ DB

Key Components

Short Code Generation โ€” two approaches:

  1. Hash-basedmd5(long_url) โ†’ take first 7 chars. Problem: collisions. Fix: append incrementing counter on collision. Problem: deterministic โ€” same URL always maps to same code.

  2. Counter + Base62 encoding (recommended):

    • Global counter (or distributed snowflake ID)
    • base62(counter) โ†’ 7 chars supports 62^7 โ‰ˆ 3.5 trillion unique codes
    • Atomic counter: use Redis INCR or a DB sequence
    • Pre-generate batches of IDs per server to avoid central bottleneck

Redirect Service

  • Check Redis cache first (hot short codes)
  • On miss: query DB, populate cache with TTL
  • Return HTTP 301 (permanent, browser caches = fewer hits) vs 302 (temporary, each visit hits server = better analytics)

Database

urls(
  code        VARCHAR(10) PRIMARY KEY,
  long_url    TEXT NOT NULL,
  user_id     BIGINT,
  created_at  TIMESTAMP,
  expires_at  TIMESTAMP,
  click_count BIGINT DEFAULT 0
)

Key Design Decisions

301 vs 302 Redirect

  • 301: browser caches the redirect โ†’ future visits bypass your server โ†’ no analytics, but lower load
  • 302: browser always asks your server โ†’ full analytics visibility, higher QPS
  • Decision: use 302 if analytics matter; 301 for pure performance

Cache Strategy

  • Cache: code โ†’ long_url in Redis, TTL = 24h
  • 80/20 rule: 20% of short codes drive 80% of traffic โ€” small hot cache covers most requests
  • With 115k redirects/s and ~1ฮผs Redis get, Redis easily handles the load

ID Generation at Scale

  • Single counter = single point of failure
  • Solution: range allocation โ€” each app server pre-fetches a range of 10,000 IDs from a central counter. Exhausts the range before fetching more. Central counter is hit rarely.
  • Alternative: Twitter Snowflake (timestamp + datacenter ID + sequence) for globally unique IDs without coordination

Custom Aliases

  • Allow user-specified codes: check availability in DB, insert with their code as primary key
  • Separate namespace from auto-generated codes (e.g., auto codes are numeric base62; custom are alpha-prefixed)

Security Considerations

ThreatMitigation
Phishing via short URLsScan long_url against SafeBrowsing API / malicious URL DB before shortening
Spam / abuseRate limit per user/IP on writes; require account for high-volume usage
Open redirect abuseValidate that long_url is a real HTTP/HTTPS URL; block localhost and private IP ranges
Enumeration of short codesRandom/unpredictable code generation; don't use sequential codes that are trivially enumerable
Click hijacking via analyticsIf exposing analytics API, require authentication and validate ownership
SSRF via long_urlWhen previewing/validating URLs server-side, use a safe SSRF-blocking HTTP client

Interview Tips

Start with scale estimation

writing 3/s, reading 115k/s tells you immediately: this is a read-heavy system, caching is the primary scaling lever.

Mention 301 vs 302 trade-off

interviewers love this because it shows you understand that requirements (analytics vs performance) drive technical decisions.

Don't skip the ID generation problem

it's the most interesting technical challenge here. Single counter โ†’ range allocation โ†’ Snowflake ID evolution is a good story.

Security add-on

mention SafeBrowsing API integration to prevent phishing โ€” quick win that shows security awareness.