System Design Lab›System Design Questions›Design Spotify
Design Spotify
MediumReal-time Systemsstreamingcdnencodingstoragesync
Question Overview
Design a music streaming platform serving a 100-million-track catalog to hundreds of millions of listeners. Discuss multi-bitrate encoding, delivering terabits per second of audio through CDNs with fast playback start, syncing playlists across devices, licensed offline downloads, and counting every play accurately for royalties.…
Sign up to see the full question and AI interviewer
Requirements
- Search and browse a 100M-track catalog of artists, albums, and tracks
- Stream at 96, 160, or 320 kbps with fast start and bitrate adaptation to network conditions
- Create, edit, share, and follow playlists that sync across a user's devices within seconds
- Encrypted offline downloads for premium users, with licenses renewed periodically online
- Every play, including offline plays, recorded exactly once for play counts and royalties
- Playback start under 500 ms at p95 and 99.95% playback availability
Back-of-the-envelope numbers
- Plays: 200M DAU × 30 tracks = 6B plays/day ÷ 86,400 s ≈ 69K track starts/s average, ~210K/s at 3× peak
- Concurrent listeners: 6B × 3.5 min = 21B stream-minutes/day ÷ 1,440 min ≈ 14.6M average, ~44M at peak
- Bandwidth: 14.6M × 160 kbps ≈ 2.3 Tbps average, ~7 Tbps at peak, which only a CDN can serve economically
- Catalog: 210 s × (96 + 160 + 320) kbps ≈ 15 MB per track across bitrates × 100M tracks ≈ 1.5 PB, plus ~3.7 PB of lossless masters
- Hot set: the most-played 1M tracks × 15 MB ≈ 15 TB, small enough to keep warm across CDN edge caches
- Play events: 6B/day × ~200 B ≈ 1.2 TB/day ≈ 440 TB/year of raw events for counts and royalties
- Playlists: 4B × 50 entries × ~25 B ≈ 5 TB of playlist entries, sharded by playlist_id
Key components
- Ingestion pipeline that takes label-delivered masters, normalizes loudness, transcodes to each bitrate and codec, encrypts, and stores files in object storage
- CDN in front of origin storage serving audio by byte range via signed URLs; clients start at a lower bitrate and prefetch the next track
- Catalog service with artist, album, and track metadata plus per-country licensing, and a search index with prefix autocomplete
- Playback service that checks entitlement and region, then issues short-lived signed URLs and decryption keys
- Playlist service storing ordered entries by playlist_id with a revision number, so concurrent edits from different devices merge instead of overwriting
- Offline manager that stores encrypted files tied to the device, with licenses that expire unless the app periodically checks in
- Play-event pipeline: clients buffer events with unique IDs and send them to Kafka, where jobs deduplicate and aggregate counts and royalty reports
Common mistakes
- Serving audio from a few origin data centers instead of a CDN, which cannot economically deliver ~7 Tbps at peak
- Storing audio blobs in the metadata database instead of object storage
- Downloading a whole track before playback instead of streaming the first bytes at a lower bitrate and prefetching the next track
- Counting plays on the server when a file is fetched, which counts prefetches and retries and misses offline plays
- Storing offline downloads unencrypted or with permanent licenses, breaking the terms that allow offline listening
- Saving a playlist as one blob rewritten on every edit, so concurrent edits from two devices silently lose changes
- Spending most of the interview on recommendations and never covering streaming, storage, or play-count correctness
Likely follow-ups
- How would you make playback start almost instantly for a track the user has never played?
- How would you prepare for a superstar album released worldwide at midnight?
- How do offline plays get reported for royalties when a device reconnects days later?
- How would you implement collaborative playlists edited by several users at once?
- How would you enforce licensing when a track is available only in some countries?
- How would you add lossless audio without breaking older clients?
No community solutions yet
Be the first to publish your solution
Practice ‘Design Spotify’ with an AI Interviewer
Get scored feedback on your diagram, scalability approach, and trade-offs. Free while we grow — up to 3 full interviews a day.