Design a Price Tracking Service
DifficultyMedium | HelloInterview: problem breakdown
Problem Statement
Design a service that tracks prices of products across e-commerce sites and alerts users when a product's price drops below their target. Think CamelCamelCamel for Amazon or Honey.
๐Real-world: Price tracking is a web-scraping problem with an adversary: the e-commerce sites you're scraping actively don't want to be scraped โ they deploy bot detection, rate limits, CAPTCHAs, and dynamic/obfuscated HTML, and prices are increasingly rendered by JavaScript, so you need headless browsers and rotating proxies. That cat-and-mouse is the real engineering, plus change detection (poll efficiently โ don't re-fetch a million products every minute; prioritize by volatility and user demand) and alerting on threshold crossings. The cautionary tale for this audience is Honey (2020+ controversy): the price/coupon tool was accused of hijacking affiliate-commission cookies (replacing the referrer's tracking code with its own at checkout) โ a vivid reminder that the client-side extension form of these tools has serious supply-chain and trust implications (a browser extension with access to every page you visit is a huge attack surface). Design-wise it's scheduler + scraper pool โ price-history TSDB โ alert queue, but the scraping arms race and the extension's privilege are the parts worth discussing.
Requirements
Functional
- Track price of a product URL
- Set a price alert (notify me when price โค target)
- View price history (time-series chart)
- Scheduled price checks (every N hours)
- Support multiple retailers (Amazon, Walmart, Best Buy, etc.)
Non-Functional
- 10M tracked products, 50M alerts
- Price check frequency: hourly for popular products, daily for less-popular
- Alert delivery: < 5 minutes after price drop detected
- Scraping: must handle bot detection, CAPTCHAs, rate limiting by retailers
Core Design
Price Scraping Architecture
Scheduler (priority queue by next_check_time)
โ
Fetch Queue (Kafka, partitioned by retailer_domain)
โ
Scraper Workers (per domain โ respect rate limits)
โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Parse price from HTML/JSON API โ
โ Extract: price, currency, in-stock โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
Price Event (product_id, price, timestamp) โ Kafka
โ
Price DB (TimescaleDB / Cassandra: time-series)
โ
Alert Checker โ Notification ServiceScraping Challenges
Retailers actively block scrapers. Mitigation strategies:
residential proxy pools; rotate per request
mimic real browsers
Puppeteer/Playwright for JavaScript-rendered pages
respect crawl delays; back off on 429
prefer official product APIs (Amazon PA-API) where available โ legitimate and reliable
Alert Evaluation
On new price event (product_id, new_price):
Query alerts WHERE product_id = ? AND target_price >= new_price AND status = 'ACTIVE'
For each matching alert:
if last_notified < (now - 24h): -- don't spam if price keeps dropping
send notification
update last_notifiedAt scale (50M alerts, 10M products): alert lookup is frequent. Index (product_id, target_price). Pre-compute per-product: "what's the lowest alert target?" โ only query alerts if new price โค that floor.
Price History (Time-Series)
TimescaleDB (or Cassandra):
price_history(
product_id UUID,
checked_at TIMESTAMP,
price DECIMAL(10,2),
currency CHAR(3),
in_stock BOOLEAN
PRIMARY KEY (product_id, checked_at DESC)
)Hypertable with time partitioning. Query last 90 days fast; older data compressed/archived.
Architecture
Scheduler โ Fetch Queue (Kafka) โ Scraper Workers
โ
Price DB (TimescaleDB) โ Price History API
โ
Alert Checker (stream processing)
โ
Notification Service โ Email / Push / SMSSecurity Considerations
| Threat | Mitigation |
|---|---|
| Scraping ToS violation | Prefer official APIs where available; follow robots.txt |
| SSRF via user-submitted URLs | Validate URLs are allowed retailers; block internal IPs |
| Notification spam | Rate limit notifications per user (max 1/product/day); unsubscribe link in emails |
| Price manipulation alerts | Mark alerts triggered by flash sales differently from genuine price drops |
| User email phishing risk | Send from authenticated domain (SPF/DKIM/DMARC); avoid clicking links in notification emails |
Interview Tips
discuss rotating proxies, headless browsers, and API-first preference.
the floor optimization (only run alert query if new price is below lowest target for that product) reduces query volume dramatically.
TimescaleDB or Cassandra is the right choice; don't store in a flat SQL table without partitioning.