Design Facebook Post Search
DifficultyMedium | HelloInterview: problem breakdown
Problem Statement
Design a search system that lets users search across their own posts and their friends' public/shared posts by keyword.
๐Real-world: The hidden hard part is that search and access control collide. A naive inverted index returns "all posts containing the word," but here the results must be filtered by who's allowed to see them โ your friends' friends-only posts, privacy settings, blocks โ and doing that filtering after ranking breaks pagination and leaks existence. Facebook's actual system (Unicorn, their graph-aware search engine) had to make the index itself privacy-aware. This is the core security lesson for the audience: search is a notorious authorization-bypass surface โ if you forget to apply the viewer's permissions to every result, you've built a data-leak engine, and "search returned something I shouldn't see" is a classic real bug class (IDOR-via-search). Technically it's an inverted index (Elasticsearch) + async indexer (Kafka) re-ranked by recency and social signals, but the headline move is to bake per-viewer access control into retrieval, not bolt it on after.
Requirements
Functional
- Full-text search across posts
- Results filtered by: my posts, friends' posts, public posts
- Ranked by: relevance + recency
- Support for typo tolerance
- Privacy-aware: only show posts the searcher has permission to see
Non-Functional
- Billions of posts indexed
- Search latency: < 500ms
- Index freshness: new posts searchable within 1 minute
- Privacy enforcement: strict โ never return unauthorized posts
Core Design
Search Index with Privacy
The hard part isn't the search โ it's enforcing privacy at search time without re-querying the DB for every result.
Approach 1: Pre-filter index โ each document in the search index includes a visibility list (which user IDs can see it). Query filters by visibility.
- Problem: visibility list can be huge (public posts visible to all). Storing a list of 2B users in an index document doesn't work.
Approach 2: Search then post-filter โ search returns candidates, then DB lookup enforces privacy.
- Problem: if top 20 results are all private, you get 0 results after filtering. Need to fetch more candidates.
Approach 3: Visibility categories (recommended):
Post visibility types:
- SELF: only author sees it
- FRIENDS: author + friends list
- PUBLIC: everyone
Index documents with visibility=PUBLIC are searchable by all.
Index documents with visibility=FRIENDS: index separately; only search this sub-index when the searcher is a friend of the author.At search time: query the public index + the friends-only indexes of all the searcher's friends (or a subset if too many friends).
Write Path (Index Update)
New post created โ Post Service โ Kafka โ Search Indexer
โ
Elasticsearch (index document)
Document contains: post_id, author_id, content, created_at, visibilityRead Path
Search query โ Search Service
โ
1. Build permissioned query:
visibility=PUBLIC
OR (visibility=FRIENDS AND author_id IN [user's friend IDs])
2. Execute against Elasticsearch
3. Rank: BM25 relevance ร recency boost
4. Return with post metadata (from Post Service cache)Friend IDs at Search Time
"User's friend IDs" โ up to 5000 friends. Passing 5000 IDs to Elasticsearch filter is doable. For users with many friends, pre-compute a "friends set" in Redis per user.
Architecture
Post Write โ Kafka โ Search Indexer โ Elasticsearch
โ
Search Query โ Search Service (add friend IDs from Social Graph)
โ
Elasticsearch (permissioned query)
โ
Hydrate post details โ ResponseSecurity Considerations
| Threat | Mitigation |
|---|---|
| Privacy bypass | Visibility category in index; friends-only posts only returned when friendship verified |
| Content injection via search | Sanitize search query before Elasticsearch query construction (prevent injection) |
| PII in post content | Content moderation; GDPR right to erasure must propagate to search index |
| Enumeration attack (search to find private posts) | Consistent 0-result response for unauthorized posts; no timing difference |
Interview Tips
spend the most time on visibility categories and how you enforce them at index time vs query time.
standard choice; mention BM25 scoring and fuzzy matching for typo tolerance.
near-real-time via Kafka is better than polling DB.