System Design Lab›System Design Questions›Design Price Drop Tracker
Design Price Drop Tracker
EasyE-commercecrawlerschedulingnotificationstime-series
Question Overview
Design a service that watches millions of product prices and alerts users when an item hits their target price. The interesting parts are spending a limited crawl budget on the products that matter, storing price history compactly, and matching drops to watchers without duplicate alerts.…
Sign up to see the full question and AI interviewer
Requirements
- Add a product URL with a target price; edit or remove watches at any time
- Fetch prices on a priority-based schedule and record price history for charts
- Notify each watcher once when the price reaches their target, re-arming only after it rises above again
- Refresh popular or heavily watched products at least hourly and the long tail every few days
- Send notifications within 5 minutes of a detected drop, with no duplicates
- Respect retailer rate limits and terms, preferring official product APIs where available
Back-of-the-envelope numbers
- Crawl tiers: 1M hot products hourly (24M/day) + 4M other watched products every 6 h (16M/day) + 45M unwatched every 3 days (15M/day) = 55M fetches/day
- Fetch rate: 55M ÷ 86,400 s ≈ 640 fetches/s, spread across retailers under per-domain rate limits
- Crawl bandwidth: 55M × ~100 KB per page ≈ 5.5 TB/day ≈ 510 Mbps sustained; keep extracted fields, not raw pages
- History: 10% of 55M fetches change price → 5.5M points/day × ~30 bytes ≈ 165 MB/day ≈ 60 GB/year, versus ~600 GB/year storing every observation
- Watches: 20M × ~100 bytes ≈ 2 GB, indexed by (product_id, target_price), about 4 watchers per watched product
- Chart views: 2M visitors × 5 views/day = 10M/day ÷ 86,400 s ≈ 116/s average, ~350/s at 3× peak
Key components
- Product resolver that normalizes pasted URLs into a canonical key (retailer + product ID) so 10K watchers of one item cost one fetch
- Crawl scheduler: priority queue by next_fetch_at, with priority from watcher count, page views, price volatility, and distance to targets
- Fetchers with per-retailer parsers or API clients, per-domain token buckets, and sanity checks that reject implausible prices
- Price history store (e.g. Cassandra or ClickHouse) keyed by product and time, holding change points plus a daily sample
- Drop matcher: on each price change, query watches for that product with target_price >= new price that are still armed
- Notification service keyed by (watch_id, price_change_id) for idempotency, with per-user batching and unsubscribe handling
Common mistakes
- Crawling every product at the same frequency, wasting budget on dead long-tail items while popular ones go stale
- Scanning all 20M watches after every fetch instead of indexing watches by product and target price
- Storing every raw page or every identical observation instead of only price changes
- Trusting parsed prices blindly, so a broken parser reporting $0 triggers a flood of false alerts
- Treating different URL forms of one product as separate products, multiplying fetches and splitting history
- Re-notifying on every fetch while the price stays below target instead of tracking armed and triggered states
- Ignoring retailer rate limits and getting crawler IPs blocked, which stalls every product from that store
Likely follow-ups
- How would you handle Black Friday, when millions of prices change within an hour?
- How would you detect and suppress bogus prices from a broken parser or a retailer glitch?
- How would you support multiple sellers and currencies for the same product?
- How would you support alerts like 20% below the 90-day average efficiently?
- How would you reprioritize crawling if your fetch budget were cut in half?
- How would you adapt if a major retailer started blocking your crawler?
No community solutions yet
Be the first to publish your solution
Practice ‘Design Price Drop Tracker’ with an AI Interviewer
Get scored feedback on your diagram, scalability approach, and trade-offs. Free while we grow — up to 3 full interviews a day.