Design File Storage (Dropbox/Google Drive)
DifficultyEasy | HelloInterview: problem breakdown
Problem Statement
Design a cloud file storage and sync service where users can upload files, access them from any device, and share them with others. Files should sync automatically across devices.
๐Real-world: Dropbox's core trick is content-addressed chunking with deduplication โ files are split into ~4MB blocks, each hashed, and only changed blocks are uploaded (delta sync). The famous payoff: if a million users store the same popular file, it's stored once. But that same dedup created a notorious security subtlety โ if dedup is global across users and the system tells you "already uploaded" for a block you didn't have, an attacker can confirm whether a specific file exists in someone else's storage (a side channel), which is why providers dedup carefully and per-account where privacy matters. Also famous: Dropbox's 2011 four-hour auth bug where any password would unlock any account, and its later move to client-side encryption debates (if the server holds the keys, it โ and anyone who breaches it โ can read your files). The security framing interviewers reward: who holds the encryption keys, and what does dedup leak?
Requirements
Functional
- Upload and download files
- Folder hierarchy
- File sync across devices (bidirectional)
- File sharing (link sharing + permission control)
- File versioning / history
- Mobile and desktop clients
Non-Functional
- 500M users; 1B files
- Average file size: 1MB; P99: 100MB
- Files should sync within 30 seconds of change
- Storage optimized: deduplication across users (same file content = stored once)
Core Design
Chunked Upload
Large file uploads over a single HTTP connection fail on network interruption. Solution: chunk the file.
Client splits file into 4MB chunks
Each chunk: hash(chunk_content) โ content-addressed ID
Upload: PUT /chunks/{hash} for each chunk
Commit: POST /files { chunk_hashes: [...], path: "/Documents/report.pdf" }Benefits
- Resumable: on failure, resume from last missing chunk
- Deduplication: if chunk hash already exists in storage, skip upload (content-addressed)
- Delta sync: on file modification, only upload changed chunks (typically 1-2 chunks out of many)
Deduplication
Two users upload the same file (e.g., Ubuntu ISO):
- Content hash of file matches a stored file โ store metadata pointing to existing content
- No duplicate storage
chunks table:
hash (PK), size, s3_key
-- stored once even if 10,000 users have the same file
file_chunks table:
file_id, chunk_sequence, chunk_hash
-- mapping: file โ ordered list of chunksSync Protocol
Client change โ Local event queue (file watcher: inotify / FSEvents)
โ Compute changed chunks
โ Upload new/modified chunks
โ POST file metadata to Sync Service
โ Sync Service publishes change event
โ Other clients for same user poll / receive WebSocket notification
โ Client downloads changed chunks onlyArchitecture
Client (Desktop/Mobile)
- File watcher (inotify/FSEvents)
- Local chunk cache
โ
Upload Service โ S3 (chunk storage, content-addressed by hash)
Metadata Service โ PostgreSQL (files, folders, permissions, versions)
Sync Service โ WebSocket / SSE (notify clients of changes)
Download Service โ CDN-fronted S3 (fast downloads)
Sharing Service โ PostgreSQL (share links, permissions)File Versioning
file_versions(
file_id UUID,
version INT,
created_at TIMESTAMP,
created_by UUID,
chunk_hashes TEXT[], -- ordered array of chunk hashes
size BIGINT
)Restoring a version: read chunk_hashes array, reassemble from S3. Since chunks are content-addressed, no additional storage if a restored version shares chunks with the current version.
Key Design Decisions
| Decision | Choice | Reason |
|---|---|---|
| Chunk size | 4-8 MB | Balance between number of round trips and granularity for delta sync |
| Addressing | Content hash | Deduplication; immutable (same hash = same content) |
| Storage | S3 | Cheap, durable, globally replicated |
| Metadata | PostgreSQL | Tree hierarchy (folder structure); ACID transactions |
| Sync | WebSocket + push | < 30s sync requirement; avoid polling |
Security Considerations
| Threat | Mitigation |
|---|---|
| Unauthorized access | Per-file ACL; presigned S3 URLs with short TTL |
| Malware upload | Hash-based malware scanning (known bad hashes); sandbox execution for suspicious files |
| Ransomware via sync | File versioning (recover previous versions); anomaly detection on mass deletion/encryption |
| Link sharing abuse | Signed URLs with expiry; password-protect share links; revocation capability |
| Chunk enumeration | Chunk hashes are SHA-256: not guessable; but add auth check on chunk download |
| Client-side encryption | Offer E2EE option where chunks are encrypted before upload (key stays on client) |
Interview Tips
is the key design โ explain the chunk split/hash/upload/commit flow.
is a quick win โ same hash = same content = store once = huge storage savings.
is the real differentiator for user experience โ only upload changed chunks, not the whole file.
is worth mentioning โ it's a standard feature and shows you think about data safety.