Designing Data-Intensive Applications
Case 3

Back-of-the-Envelope Estimation

Message write QPS, connection fan-in, hot storage, and presence memory at WhatsApp scale.

Order-of-magnitude checks catch designs that cannot work. Adjust assumptions for your interview scope — exact numbers matter less than which component becomes the bottleneck.

Assume 2B MAU, 500M DAU, 100B messages/day, avg message 2 KB ciphertext, 200M concurrent connections peak, 7-day server-side retention for offline delivery.

Write traffic

  • Messages: 100B / 86,400 ≈ 1.16M messages/s average.
  • Peak (2× average) ≈ 2.3M writes/s globally — partitioned by chat_id across message store clusters.
  • Group fan-out: 1 write → N deliveries; a 256-member group multiplies connection writes (hybrid pull model essential for large groups).

Read traffic

  • Delivery: roughly 1:1 with sends for online users; offline users read on reconnect (sync from last acked seq).
  • Connection fan-in: 200M concurrent WebSockets — ~2M connections per chat server at 100 machines (order-of-magnitude).
  • Push notifications: ~30% offline → 350M pushes/day ≈ 4,000/s to APNs/FCM.

Storage

  • Hot queue (7-day retention): 100B/day × 7 × 2 KB ≈ 1.4 PB — in practice most messages delivered and trimmed faster; plan hundreds of TB.
  • Long-term: E2E encrypted blobs on device; server storage is routing + short retention, not infinite archive.

Memory & caching

  • Presence (Redis): 500M DAU × 100 B presence key ≈ 50 GB with TTL — sharded Redis clusters.
  • Connection state: 200M × ~4 KB session ≈ 800 GB RAM fleet-wide for routing tables.
  • Per-chat sequence counters: in-memory on chat server partition — negligible vs connection RAM.
typescript — WhatsApp message write QPS
// WhatsApp-scale message writes
const messagesPerDay = 100_000_000_000;
const avgWriteQps = messagesPerDay / 86_400;
const peakWriteQps = avgWriteQps * 2;
const concurrentConnections = 200_000_000;
const hotStorageTb = (messagesPerDay * 7 * 2 * 1024) / 1e12; // 7d × 2KB
console.log({ peakWriteQps: Math.round(peakWriteQps), concurrentConnections, hotStorageTb });
Order-of-magnitude summary

~2.3M peak message writes/s, 200M concurrent connections, ~1 PB order-of-magnitude hot storage, presence in Redis not Postgres. Optimize connection fan-in, not SQL joins.

Key Takeaways
  • ~2.3M peak message writes/s globally — partition by chat_id.
  • 200M concurrent WebSocket connections dominate RAM, not SQL joins.
  • ~1 PB order-of-magnitude hot queue storage with short retention.
  • Presence and typing belong in Redis with TTL, not Postgres.
QPSWebSocketconnectionsstorageRedisfan-out