Case 9
Back-of-the-Envelope Estimation
Peak ops/s, active document RAM, and why op logs need compaction.
Order-of-magnitude checks catch designs that cannot work. Adjust assumptions for your interview scope — exact numbers matter less than which component becomes the bottleneck.
Assume 50M docs, 2M DAU, 50k docs edited concurrently peak, 5 editors/doc avg, 2 ops/s/editor, 500 B/op, 5 MB avg asset per doc, 100 KB metadata.
Real-time ops (WebSocket)
- Ops/s: 50k docs × 5 editors × 2 ops/s ≈ 500k ops/s globally.
- Per collab server (10k docs): 50k ops/s — sticky routing by doc_id.
- Broadcast fan-out: 500k ops × 4 peers avg × 500 B ≈ 1 GB/s WebSocket egress fleet-wide.
Storage
- Metadata (Postgres): 50M × 100 KB ≈ 5 TB (+ indexes → 15 TB).
- Assets (S3): 50M × 5 MB ≈ 250 PB total — CDN for export/download; uploads direct to S3.
- Op log (before compaction): 500k ops/s × 500 B × 3600 ≈ 900 GB/hour raw — snapshot every 5 min, compact to 100 KB snapshots.
Memory & caching
- Active doc in RAM (CRDT/OT): 50k × 2 MB working set ≈ 100 GB across collab fleet.
- Presence (Redis): 2M DAU × 150 B ≈ 300 MB with TTL.
- ACL cache: 50k hot docs × 1 KB permission snapshot ≈ 50 MB.
Read traffic
- Doc open: 2M opens/day ≈ 23/s — load snapshot + replay tail of op log since snapshot.
- Export PDF (async): 100k/day ≈ 1/s — separate worker queue, not collab hot path.
ops/sWebSocketS3PostgreSQLcompactionRedis