Designing Data-Intensive Applications
Case 2

Requirements and Constraints

YouTube must ingest massive files, serve globally with low startup latency, and recommend the next clip — all profitably.

Video is the bulk of bytes; metadata is the bulk of queries. Treating them as one system fails — upload and transcode are write pipelines, playback is a CDN problem, home feed is a ranking problem.

In practice

YouTube stores originals in durable object storage (GCS-style), serves via Google's CDN, and keeps metadata in sharded databases. Creators see processing state; viewers see p95 startup time under a few seconds on good networks.

YouTube at scale

A 2 GB upload cannot block a synchronous API thread. Resumable uploads land in object storage; a job queue fans out transcoding workers that write HLS/DASH segments back to the bucket.

Diagram

Why these technologies?

Why Object storage (S3 / GCS)?

Video files are gigabytes — storing them in Postgres or on API servers is impossible at scale. Object storage offers pennies-per-GB durability with multipart/resumable upload APIs. Metadata (title, status) lives elsewhere.

Why Upload API (stateless service)?

Accepts chunked uploads, validates auth, and issues pre-signed URLs so bytes stream directly to object storage without proxying through app servers. Keeps API pods memory-light.

Why Kafka / Pub/Sub?

Transcoding takes minutes and must not block the upload response. An event ('video.uploaded') fans out to worker pools that scale independently. Also feeds moderation, thumbnails, and search indexing.

Why Transcoder workers (FFmpeg fleet)?

CPU-heavy, embarrassingly parallel — one job per resolution. Run on spot instances or batch queues, not on request-serving API pods. Output is immutable HLS/DASH segments back to object storage.

Why CDN?

99% of bytes are video segments served repeatedly — edge caching cuts origin egress cost and startup latency. Adaptive bitrate manifests let players switch quality without new API calls.

Why Metadata database (PostgreSQL / sharded SQL)?

Titles, channels, comments, and processing state need joins, indexes, and transactions. Sharded or regional when single-node Postgres saturates — not replaced by object storage.

Why Stream processor (Flink / Dataflow) for view counts?

Exact per-view counts are unnecessary; approximate aggregation at millions of events/sec is cheaper than row-level UPDATE on a video table.

Key Takeaways
  • Upload: resumable, durable, survive client crashes mid-transfer.
  • Processing: transcode to many resolutions and codecs (ABR ladder).
  • Delivery: edge caching; adaptive bitrate for variable networks.
  • Metadata: title, channel, stats, comments — OLTP with heavy read caching.
  • Recommendations: batch + stream features; personalization is derived data.
  • Copyright and moderation pipelines are async consumers of upload events.
videouploadCDNtranscodingmetadataS3Kafka