Designing Data-Intensive Applications
Study Companion

Designing Data-Intensive Applications

An interactive study companion for Martin Kleppmann's classic on reliable, scalable, and maintainable systems.

Start with Chapter 1
Part 1
Foundations of Data Systems

4 chapters

Explore part
Part 2
Distributed Data

5 chapters

Explore part
Part 3
Derived Data

3 chapters

Explore part
System Design
System Design Case Studies

10 case studies

Explore case studies

System Design Case Studies

Real-world architectures for YouTube, WhatsApp, booking systems, voting, multiplayer games, collaboration tools, multistep workflows, and long-running forms — with diagrams and TypeScript examples.

Case 1: System Design Foundations
How to frame requirements, estimate capacity, and reason about trade-offs before drawing boxes.
Case 2: Designing YouTube
Video upload, transcoding pipelines, CDN delivery, and metadata at billions of views.
Case 3: Designing WhatsApp
Low-latency messaging, delivery guarantees, offline users, and end-to-end encryption.
Case 4: Reservation & Booking Systems
Inventory holds, double-booking prevention, payments, and calendar search at Airbnb scale.
Case 5: Voting Systems
One person one vote, tamper resistance, auditability, and handling viral traffic spikes.
Case 6: Multiplayer Games
Real-time state sync, authoritative servers, matchmaking, and regional game shards.
Case 7: Multiplayer Puzzle Games
Turn-based play, async moves, conflict resolution, and leaderboards for Wordle-style games.
Case 8: Multistep Systems & Workflows
Checkout flows, sagas, compensating transactions, and durable workflow engines.
Case 9: Collaboration Tools
Google Docs, Figma, and Notion — presence, operational transforms, and CRDTs.
Case 10: Multistep Forms & Wizards
Loan applications, onboarding wizards, draft persistence, validation, and resume-later UX.

Curriculum

Ch. 1: Reliable, Scalable, and Maintainable Applications
Available
Part 1: Foundations of Data Systems

What we mean by reliability, scalability, and maintainability — and how they shape every data-intensive application.

Open chapter
Ch. 2: Data Models and Query Languages
Available
Part 1: Foundations of Data Systems

Relational, document, and graph models — and the query languages that bring them to life.

Open chapter
Ch. 3: Storage and Retrieval
Available
Part 1: Foundations of Data Systems

The data structures and storage engines that power databases under the hood.

Open chapter
Ch. 4: Encoding and Evolution
Available
Part 1: Foundations of Data Systems

How data is encoded, how schemas evolve, and how information flows between systems.

Open chapter
Ch. 5: Replication
Available
Part 2: Distributed Data

Leaders, followers, quorums, and the consistency trade-offs of replicated data.

Open chapter
Ch. 6: Partitioning
Available
Part 2: Distributed Data

Splitting data across nodes — key-range vs hash partitioning, rebalancing, and routing.

Open chapter
Ch. 7: Transactions
Available
Part 2: Distributed Data

ACID, isolation levels, and the path to serializability.

Open chapter
Ch. 8: The Trouble with Distributed Systems
Available
Part 2: Distributed Data

Networks, clocks, and partial failures — why distributed systems are fundamentally uncertain.

Open chapter
Ch. 9: Consistency and Consensus
Available
Part 2: Distributed Data

Linearizability, ordering, and reaching agreement in distributed systems.

Open chapter
Ch. 10: Batch Processing
Available
Part 3: Derived Data

MapReduce, dataflow engines, and the Unix philosophy applied at scale.

Open chapter
Ch. 11: Stream Processing
Available
Part 3: Derived Data

Event streams, change data capture, and reasoning about time in motion.

Open chapter
Ch. 12: The Future of Data Systems
Available
Part 3: Derived Data

Unbundling databases, correctness, and building applications around dataflow.

Open chapter