Security Notes
System Design

Design a News Aggregator

3 min read 6 sections

DifficultyEasy | HelloInterview: problem breakdown


Problem Statement

Design a service that aggregates news from multiple sources, deduplicates similar stories, and presents a ranked feed to users. Think Google News, Feedly, or Techmeme.

๐Ÿ“–

Real-world: The deceptively hard part isn't fetching feeds โ€” it's near-duplicate detection: 50 outlets publish the same wire story with different headlines, and you must cluster them into one entry. Exact hashing fails (the text differs slightly), so the real tool is similarity hashing like SimHash/MinHash (Google used SimHash to dedup the web), which produces fingerprints where similar documents land close together โ€” the same locality-sensitive-hashing idea behind plagiarism detection and the vector/embedding similarity in RAG. The pipeline is crawler/RSS pull โ†’ dedup/cluster โ†’ rank by recency + source authority + engagement โ†’ serve. Modern relevance for a security person: this is also where misinformation and source-trust become design concerns โ€” ranking has to weigh source credibility, and the same clustering tech is used to detect coordinated inauthentic content (the same false story seeded across many fake outlets).


Requirements

Functional

  • Ingest news from multiple sources (RSS feeds, web scraping, news APIs)
  • Deduplicate / cluster similar articles (same story from multiple outlets)
  • Rank articles by relevance, recency, engagement
  • Personalized feed (based on user interests / reading history)
  • Category browsing (Tech, Politics, Sports...)

Non-Functional

  • Ingest ~10k articles/hour across 10k sources
  • Feed load: 50M DAU โ†’ 1M feed requests/hour
  • Article freshness: new articles appear in feed within 5 minutes

Core Design

Ingestion Pipeline

Scheduler (cron per source) โ†’ Fetcher โ†’ Raw Article Store (S3)
                                            โ†“
                                     Content Parser (extract title, body, pub_date)
                                            โ†“
                                     NLP Processor:
                                       - Entity extraction (people, orgs, locations)
                                       - Keyword / topic classification
                                       - Sentence embedding (for deduplication)
                                            โ†“
                                     Article DB (Elasticsearch)

Politenessrespect each source's rate limit and robots.txt. Cache robots.txt per domain.

Deduplication / Story Clustering

Multiple outlets publish the same event. Group them into a "story":

SimHash

of article body: articles with near-identical content (syndicated) โ†’ same SimHash bucket โ†’ same story

Semantic similarity

compute embedding (BERT/sentence-transformers); cosine similarity > threshold โ†’ same story

Entity + time window

same entities (person + event + date) in articles published within 2h โ†’ likely same story

One story = canonical cluster. Show the most authoritative source first.

Ranking

score = freshness_score ร— source_authority ร— engagement_score ร— personalization_boost

freshness_score: exp(-ฮป ร— age_in_hours) โ€” exponential decay
source_authority: domain PageRank-equivalent score
engagement_score: click-through rate + reading time signals
personalization: dot product of article embedding ร— user interest vector

Personalization

User interest profile built from reading history:

  • Each article has a topic vector
  • User interest vector = weighted average of read article vectors (recent reads weighted more)
  • Similarity between user interest vector and article embedding โ†’ personalization boost

Architecture

Source Scheduler โ†’ Fetcher Workers โ†’ S3 (raw HTML)
                        โ†“
               Content Processor (NLP, embed) โ†’ Elasticsearch
                        โ†“
               Story Clusterer (SimHash + similarity) โ†’ Story DB (Postgres)
                        โ†“
               Feed Service:
                   User interest profile (Redis) + Story DB + Elasticsearch
                   โ†’ Rank โ†’ Paginate โ†’ Response

Security Considerations

ThreatMitigation
Malicious source injectionURL allowlist / source vetting; don't execute JavaScript in fetched content
SSRF via feed URLsValidate feed URLs; block RFC 1918 IPs; use sandboxed HTTP client
Misinformation spreadSource credibility score; user reporting; ML fact-check signals
CopyrightAttribution always shown; don't store full article text; store excerpt + link to source

Interview Tips

Story clustering

is the interesting problem โ€” SimHash for exact near-duplicates, semantic similarity for same story different angle.

Ingestion is async

the user never waits for fresh articles; background workers keep the index current.

Freshness vs ranking

balance: a 10-minute-old important article should rank above a 30-second-old minor one.