Security Notes
System Design

Design File Storage (Dropbox/Google Drive)

4 min read 7 sections

DifficultyEasy | HelloInterview: problem breakdown


Problem Statement

Design a cloud file storage and sync service where users can upload files, access them from any device, and share them with others. Files should sync automatically across devices.

๐Ÿ“–

Real-world: Dropbox's core trick is content-addressed chunking with deduplication โ€” files are split into ~4MB blocks, each hashed, and only changed blocks are uploaded (delta sync). The famous payoff: if a million users store the same popular file, it's stored once. But that same dedup created a notorious security subtlety โ€” if dedup is global across users and the system tells you "already uploaded" for a block you didn't have, an attacker can confirm whether a specific file exists in someone else's storage (a side channel), which is why providers dedup carefully and per-account where privacy matters. Also famous: Dropbox's 2011 four-hour auth bug where any password would unlock any account, and its later move to client-side encryption debates (if the server holds the keys, it โ€” and anyone who breaches it โ€” can read your files). The security framing interviewers reward: who holds the encryption keys, and what does dedup leak?


Requirements

Functional

  • Upload and download files
  • Folder hierarchy
  • File sync across devices (bidirectional)
  • File sharing (link sharing + permission control)
  • File versioning / history
  • Mobile and desktop clients

Non-Functional

  • 500M users; 1B files
  • Average file size: 1MB; P99: 100MB
  • Files should sync within 30 seconds of change
  • Storage optimized: deduplication across users (same file content = stored once)

Core Design

Chunked Upload

Large file uploads over a single HTTP connection fail on network interruption. Solution: chunk the file.

Client splits file into 4MB chunks
Each chunk: hash(chunk_content) โ†’ content-addressed ID
Upload: PUT /chunks/{hash} for each chunk
Commit: POST /files { chunk_hashes: [...], path: "/Documents/report.pdf" }

Benefits

  • Resumable: on failure, resume from last missing chunk
  • Deduplication: if chunk hash already exists in storage, skip upload (content-addressed)
  • Delta sync: on file modification, only upload changed chunks (typically 1-2 chunks out of many)

Deduplication

Two users upload the same file (e.g., Ubuntu ISO):

  • Content hash of file matches a stored file โ†’ store metadata pointing to existing content
  • No duplicate storage
chunks table:
  hash (PK), size, s3_key
  -- stored once even if 10,000 users have the same file

file_chunks table:
  file_id, chunk_sequence, chunk_hash
  -- mapping: file โ†’ ordered list of chunks

Sync Protocol

Client change โ†’ Local event queue (file watcher: inotify / FSEvents)
              โ†’ Compute changed chunks
              โ†’ Upload new/modified chunks
              โ†’ POST file metadata to Sync Service
              โ†’ Sync Service publishes change event
              โ†’ Other clients for same user poll / receive WebSocket notification
              โ†’ Client downloads changed chunks only

Architecture

Client (Desktop/Mobile)
  - File watcher (inotify/FSEvents)
  - Local chunk cache
          โ†“
Upload Service โ†’ S3 (chunk storage, content-addressed by hash)
Metadata Service โ†’ PostgreSQL (files, folders, permissions, versions)
Sync Service โ†’ WebSocket / SSE (notify clients of changes)
Download Service โ†’ CDN-fronted S3 (fast downloads)
Sharing Service โ†’ PostgreSQL (share links, permissions)

File Versioning

sql
file_versions(
  file_id       UUID,
  version       INT,
  created_at    TIMESTAMP,
  created_by    UUID,
  chunk_hashes  TEXT[],   -- ordered array of chunk hashes
  size          BIGINT
)

Restoring a version: read chunk_hashes array, reassemble from S3. Since chunks are content-addressed, no additional storage if a restored version shares chunks with the current version.


Key Design Decisions

DecisionChoiceReason
Chunk size4-8 MBBalance between number of round trips and granularity for delta sync
AddressingContent hashDeduplication; immutable (same hash = same content)
StorageS3Cheap, durable, globally replicated
MetadataPostgreSQLTree hierarchy (folder structure); ACID transactions
SyncWebSocket + push< 30s sync requirement; avoid polling

Security Considerations

ThreatMitigation
Unauthorized accessPer-file ACL; presigned S3 URLs with short TTL
Malware uploadHash-based malware scanning (known bad hashes); sandbox execution for suspicious files
Ransomware via syncFile versioning (recover previous versions); anomaly detection on mass deletion/encryption
Link sharing abuseSigned URLs with expiry; password-protect share links; revocation capability
Chunk enumerationChunk hashes are SHA-256: not guessable; but add auth check on chunk download
Client-side encryptionOffer E2EE option where chunks are encrypted before upload (key stays on client)

Interview Tips

Chunking

is the key design โ€” explain the chunk split/hash/upload/commit flow.

Deduplication

is a quick win โ€” same hash = same content = store once = huge storage savings.

Delta sync

is the real differentiator for user experience โ€” only upload changed chunks, not the whole file.

File versioning

is worth mentioning โ€” it's a standard feature and shows you think about data safety.