Design a News Aggregator
DifficultyEasy | HelloInterview: problem breakdown
Problem Statement
Design a service that aggregates news from multiple sources, deduplicates similar stories, and presents a ranked feed to users. Think Google News, Feedly, or Techmeme.
๐Real-world: The deceptively hard part isn't fetching feeds โ it's near-duplicate detection: 50 outlets publish the same wire story with different headlines, and you must cluster them into one entry. Exact hashing fails (the text differs slightly), so the real tool is similarity hashing like SimHash/MinHash (Google used SimHash to dedup the web), which produces fingerprints where similar documents land close together โ the same locality-sensitive-hashing idea behind plagiarism detection and the vector/embedding similarity in RAG. The pipeline is crawler/RSS pull โ dedup/cluster โ rank by recency + source authority + engagement โ serve. Modern relevance for a security person: this is also where misinformation and source-trust become design concerns โ ranking has to weigh source credibility, and the same clustering tech is used to detect coordinated inauthentic content (the same false story seeded across many fake outlets).
Requirements
Functional
- Ingest news from multiple sources (RSS feeds, web scraping, news APIs)
- Deduplicate / cluster similar articles (same story from multiple outlets)
- Rank articles by relevance, recency, engagement
- Personalized feed (based on user interests / reading history)
- Category browsing (Tech, Politics, Sports...)
Non-Functional
- Ingest ~10k articles/hour across 10k sources
- Feed load: 50M DAU โ 1M feed requests/hour
- Article freshness: new articles appear in feed within 5 minutes
Core Design
Ingestion Pipeline
Scheduler (cron per source) โ Fetcher โ Raw Article Store (S3)
โ
Content Parser (extract title, body, pub_date)
โ
NLP Processor:
- Entity extraction (people, orgs, locations)
- Keyword / topic classification
- Sentence embedding (for deduplication)
โ
Article DB (Elasticsearch)Politenessrespect each source's rate limit and robots.txt. Cache robots.txt per domain.
Deduplication / Story Clustering
Multiple outlets publish the same event. Group them into a "story":
of article body: articles with near-identical content (syndicated) โ same SimHash bucket โ same story
compute embedding (BERT/sentence-transformers); cosine similarity > threshold โ same story
same entities (person + event + date) in articles published within 2h โ likely same story
One story = canonical cluster. Show the most authoritative source first.
Ranking
score = freshness_score ร source_authority ร engagement_score ร personalization_boost
freshness_score: exp(-ฮป ร age_in_hours) โ exponential decay
source_authority: domain PageRank-equivalent score
engagement_score: click-through rate + reading time signals
personalization: dot product of article embedding ร user interest vectorPersonalization
User interest profile built from reading history:
- Each article has a topic vector
- User interest vector = weighted average of read article vectors (recent reads weighted more)
- Similarity between user interest vector and article embedding โ personalization boost
Architecture
Source Scheduler โ Fetcher Workers โ S3 (raw HTML)
โ
Content Processor (NLP, embed) โ Elasticsearch
โ
Story Clusterer (SimHash + similarity) โ Story DB (Postgres)
โ
Feed Service:
User interest profile (Redis) + Story DB + Elasticsearch
โ Rank โ Paginate โ ResponseSecurity Considerations
| Threat | Mitigation |
|---|---|
| Malicious source injection | URL allowlist / source vetting; don't execute JavaScript in fetched content |
| SSRF via feed URLs | Validate feed URLs; block RFC 1918 IPs; use sandboxed HTTP client |
| Misinformation spread | Source credibility score; user reporting; ML fact-check signals |
| Copyright | Attribution always shown; don't store full article text; store excerpt + link to source |
Interview Tips
is the interesting problem โ SimHash for exact near-duplicates, semantic similarity for same story different angle.
the user never waits for fresh articles; background workers keep the index current.
balance: a 10-minute-old important article should rank above a 30-second-old minor one.