Case 3
Back-of-the-Envelope Estimation
Message write QPS, connection fan-in, hot storage, and presence memory at WhatsApp scale.
Order-of-magnitude checks catch designs that cannot work. Adjust assumptions for your interview scope — exact numbers matter less than which component becomes the bottleneck.
Assume 2B MAU, 500M DAU, 100B messages/day, avg message 2 KB ciphertext, 200M concurrent connections peak, 7-day server-side retention for offline delivery.
Write traffic
- Messages: 100B / 86,400 ≈ 1.16M messages/s average.
- Peak (2× average) ≈ 2.3M writes/s globally — partitioned by chat_id across message store clusters.
- Group fan-out: 1 write → N deliveries; a 256-member group multiplies connection writes (hybrid pull model essential for large groups).
Read traffic
- Delivery: roughly 1:1 with sends for online users; offline users read on reconnect (sync from last acked seq).
- Connection fan-in: 200M concurrent WebSockets — ~2M connections per chat server at 100 machines (order-of-magnitude).
- Push notifications: ~30% offline → 350M pushes/day ≈ 4,000/s to APNs/FCM.
Storage
- Hot queue (7-day retention): 100B/day × 7 × 2 KB ≈ 1.4 PB — in practice most messages delivered and trimmed faster; plan hundreds of TB.
- Long-term: E2E encrypted blobs on device; server storage is routing + short retention, not infinite archive.
Memory & caching
- Presence (Redis): 500M DAU × 100 B presence key ≈ 50 GB with TTL — sharded Redis clusters.
- Connection state: 200M × ~4 KB session ≈ 800 GB RAM fleet-wide for routing tables.
- Per-chat sequence counters: in-memory on chat server partition — negligible vs connection RAM.
QPSWebSocketconnectionsstorageRedisfan-out