Security Notes
System Design

Design a Price Tracking Service

3 min read 6 sections

DifficultyMedium | HelloInterview: problem breakdown


Problem Statement

Design a service that tracks prices of products across e-commerce sites and alerts users when a product's price drops below their target. Think CamelCamelCamel for Amazon or Honey.

๐Ÿ“–

Real-world: Price tracking is a web-scraping problem with an adversary: the e-commerce sites you're scraping actively don't want to be scraped โ€” they deploy bot detection, rate limits, CAPTCHAs, and dynamic/obfuscated HTML, and prices are increasingly rendered by JavaScript, so you need headless browsers and rotating proxies. That cat-and-mouse is the real engineering, plus change detection (poll efficiently โ€” don't re-fetch a million products every minute; prioritize by volatility and user demand) and alerting on threshold crossings. The cautionary tale for this audience is Honey (2020+ controversy): the price/coupon tool was accused of hijacking affiliate-commission cookies (replacing the referrer's tracking code with its own at checkout) โ€” a vivid reminder that the client-side extension form of these tools has serious supply-chain and trust implications (a browser extension with access to every page you visit is a huge attack surface). Design-wise it's scheduler + scraper pool โ†’ price-history TSDB โ†’ alert queue, but the scraping arms race and the extension's privilege are the parts worth discussing.


Requirements

Functional

  • Track price of a product URL
  • Set a price alert (notify me when price โ‰ค target)
  • View price history (time-series chart)
  • Scheduled price checks (every N hours)
  • Support multiple retailers (Amazon, Walmart, Best Buy, etc.)

Non-Functional

  • 10M tracked products, 50M alerts
  • Price check frequency: hourly for popular products, daily for less-popular
  • Alert delivery: < 5 minutes after price drop detected
  • Scraping: must handle bot detection, CAPTCHAs, rate limiting by retailers

Core Design

Price Scraping Architecture

Scheduler (priority queue by next_check_time)
           โ†“
   Fetch Queue (Kafka, partitioned by retailer_domain)
           โ†“
   Scraper Workers (per domain โ€” respect rate limits)
           โ†“
   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
   โ”‚ Parse price from HTML/JSON API โ”‚
   โ”‚ Extract: price, currency, in-stock โ”‚
   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
           โ†“
   Price Event (product_id, price, timestamp) โ†’ Kafka
           โ†“
   Price DB (TimescaleDB / Cassandra: time-series)
           โ†“
   Alert Checker โ†’ Notification Service

Scraping Challenges

Retailers actively block scrapers. Mitigation strategies:

Rotating proxies/IPs

residential proxy pools; rotate per request

User-agent rotation

mimic real browsers

Headless browser

Puppeteer/Playwright for JavaScript-rendered pages

Rate limiting

respect crawl delays; back off on 429

API first

prefer official product APIs (Amazon PA-API) where available โ€” legitimate and reliable

Alert Evaluation

On new price event (product_id, new_price):
  Query alerts WHERE product_id = ? AND target_price >= new_price AND status = 'ACTIVE'
  For each matching alert:
    if last_notified < (now - 24h):  -- don't spam if price keeps dropping
      send notification
      update last_notified

At scale (50M alerts, 10M products): alert lookup is frequent. Index (product_id, target_price). Pre-compute per-product: "what's the lowest alert target?" โ€” only query alerts if new price โ‰ค that floor.

Price History (Time-Series)

TimescaleDB (or Cassandra):
  price_history(
    product_id  UUID,
    checked_at  TIMESTAMP,
    price       DECIMAL(10,2),
    currency    CHAR(3),
    in_stock    BOOLEAN
    PRIMARY KEY (product_id, checked_at DESC)
  )

Hypertable with time partitioning. Query last 90 days fast; older data compressed/archived.


Architecture

Scheduler โ†’ Fetch Queue (Kafka) โ†’ Scraper Workers
                                         โ†“
                                  Price DB (TimescaleDB) โ† Price History API
                                         โ†“
                                  Alert Checker (stream processing)
                                         โ†“
                                  Notification Service โ†’ Email / Push / SMS

Security Considerations

ThreatMitigation
Scraping ToS violationPrefer official APIs where available; follow robots.txt
SSRF via user-submitted URLsValidate URLs are allowed retailers; block internal IPs
Notification spamRate limit notifications per user (max 1/product/day); unsubscribe link in emails
Price manipulation alertsMark alerts triggered by flash sales differently from genuine price drops
User email phishing riskSend from authenticated domain (SPF/DKIM/DMARC); avoid clicking links in notification emails

Interview Tips

Scraping is the interesting infrastructure problem

discuss rotating proxies, headless browsers, and API-first preference.

Alert evaluation at scale

the floor optimization (only run alert query if new price is below lowest target for that product) reduces query volume dramatically.

Time-series for price history

TimescaleDB or Cassandra is the right choice; don't store in a flat SQL table without partitioning.