Design YouTube (Video Platform)
DifficultyHard | HelloInterview: problem breakdown
Problem Statement
Design a video streaming platform where users can upload videos, and viewers can watch them with adaptive bitrate streaming. Handle the full pipeline: upload → processing → storage → streaming.
📖Real-world: A fun fact that captures the scale: "Gangnam Style" (2012) broke YouTube's view counter — it was a signed 32-bit integer, capped at ~2.1 billion, and the video blew past it, forcing Google to move to 64-bit counts. That's the whole "approximate counting at scale" theme in one anecdote. The core of this design is the transcoding pipeline — one uploaded video is fanned out into many resolutions/codecs (the classic async queue + worker-per-format pattern), chunked for adaptive bitrate streaming (HLS/DASH), and pushed to a global CDN where the real bytes are served. Security/trust dimensions to raise: Content ID (automated copyright matching against a reference fingerprint database) and upload moderation (CSAM/abuse scanning) are first-class subsystems — at YouTube's "500 hours uploaded per minute," moderation has to be automated pipeline infrastructure, not human review alone.
Requirements
Functional
- Upload video (any format, up to 100GB)
- Transcode video to multiple resolutions (360p, 720p, 1080p, 4K)
- Stream video to users (adaptive bitrate)
- Video search, recommendations
- Comments, likes, view counts
Non-Functional
- 2B users, 500M DAU
- 500 hours of video uploaded per minute
- 1B hours watched per day → ~700M simultaneous streams
- Upload success even with poor network (resumable)
- Low buffering: < 200ms startup latency, < 1% rebuffer rate
- Global availability
Core Design
Upload Pipeline
Client → Upload Service → Object Store (raw video)
↓
Message Queue (Kafka)
↓
Transcoding Workers (FFmpeg)
↓
┌───────────────────────────┐
│ 360p.mp4 │
│ 720p.mp4 │→ CDN Origin Store (S3)
│ 1080p.mp4 │
│ thumbnail.jpg │
└───────────────────────────┘
↓
Update DB: video status = READY
↓
Notify uploader (email/push)Resumable Uploads (critical for large files):
- Client uses chunked upload (e.g., 5MB chunks)
- Server issues an upload session ID
- Client uploads chunks; server tracks which chunks received
- On failure, client resumes from last successful chunk
- Implementation: AWS Multipart Upload, Google Resumable Upload, or TUS protocol
Why async transcodingvideo processing is CPU-intensive, takes minutes. Don't block the HTTP upload response — queue the job and process asynchronously.
Adaptive Bitrate Streaming (ABR)
Video is divided into small segments (2-10 seconds each) at multiple bitrates. A manifest file (.m3u8 for HLS, .mpd for DASH) lists all segments and bitrates.
video.m3u8:
#EXT-X-STREAM-INF:BANDWIDTH=800000,RESOLUTION=640x360
360p/segment_%d.ts
#EXT-X-STREAM-INF:BANDWIDTH=2800000,RESOLUTION=1280x720
720p/segment_%d.ts
#EXT-X-STREAM-INF:BANDWIDTH=5000000,RESOLUTION=1920x1080
1080p/segment_%d.tsClient player monitors download speed → switches to higher/lower bitrate segment URLs dynamically → smooth playback under variable network conditions.
Architecture
Upload Service
(chunked, resumable)
|
Raw Video Store (S3)
|
Message Queue (Kafka)
|
┌──────────┴──────────────┐
Transcoding Workers Thumbnail Workers
(auto-scaling, GPU-enabled)
|
Processed Video Store (S3)
|
CDN Network (CloudFront, Akamai)
|
User's Video Player
Metadata DB (PostgreSQL/Spanner)
- Video metadata, status
- User profiles, subscriptions
Search (Elasticsearch)
- Video title, description, tags
Recommendations (ML pipeline - Kafka → Flink → serving)Transcoding at Scale
- Videos are independent; parallelizable: spin up many workers
- Each worker picks one job from queue (Kafka consumer group)
- Worker downloads raw video from S3, runs
ffmpeg, uploads segments to S3 - Worker marks job complete in DB; updates video status
For very long videos: segment-level parallelism — split video into 10-minute chunks, transcode in parallel, merge.
CDN Strategy
Videos are served via CDN. CDN caches video segments at edge nodes close to users.
- Popular videos: hot in CDN cache → very low origin load
- Long-tail videos: cache miss → CDN fetches from S3 → caches for next request
- Pre-warming: when a video goes viral, proactively push to CDN edge nodes
- Regional distribution: store popular video copies in S3 per region (to reduce cross-region CDN cost)
Key Design Decisions
| Decision | Choice | Reason |
|---|---|---|
| Upload | Chunked / resumable | Large files, unreliable networks |
| Transcoding | Async queue + workers | CPU-intensive; don't block upload response |
| Streaming | HLS/DASH with ABR | Industry standard; works on all devices |
| CDN | Multi-CDN (CloudFront + Akamai) | Resilience; optimize per region |
| Video storage | S3-compatible object store | Cheap, durable, globally replicated |
| Metadata | PostgreSQL | Rich queries; video/user/comment data |
| View count | Redis (approx) + batch to DB | High write rate; approximate count acceptable |
Security Considerations
| Threat | Mitigation |
|---|---|
| Malicious video upload (malware in container) | Virus scan raw uploads before transcoding; sandbox transcoding workers |
| Copyright infringement | Content ID fingerprinting (audio/video hash) on upload; compare against known database |
| CSAM / illegal content | Hash matching (PhotoDNA); AI classifier; human review queue |
| DRM / piracy prevention | Widevine/FairPlay DRM; signed URLs with TTL for premium content |
| Upload abuse (bot uploads) | Account verification; CAPTCHA; upload quota; ML detection |
| Signed URLs for private videos | Pre-signed S3 URLs with short expiry; validate token on each segment request |
| Transcoding resource abuse | Resource quotas per user; queue priority; kill runaway jobs |
| View count manipulation | Deduplicate views (IP + user agent + time window); ML anomaly detection on view spikes |
Interview Tips
async pipeline is the right design and the interviewer wants to hear it.
"we split the video into segments at multiple bitrates, client picks based on bandwidth" — shows you know how streaming actually works.
here — 700M simultaneous streams cannot be served from a single origin.
mention Redis INCR per video, batch-sync to DB periodically — a classic approximation trade-off.
mention these as security/legal requirements. Interviewers at Google/YouTube will care.