Thinking About Data Systems
Most applications today are data-intensive rather than compute-intensive. Understanding what that means is the foundation for every design decision that follows.
When people talk about "the cloud" or "big data," they often focus on raw compute power. In practice, most modern applications are data-intensive: the hard problems come from handling large volumes of data, complex data transformations, or rapid changes — not from crunching numbers on an idle CPU.
Beyond a Single Database
We use the term data system broadly: not just a database, but any combination of components that store, process, or move data. A single product might use a relational database for transactions, a search index for full-text queries, a cache for hot reads, and a message queue for async work. Each tool has strengths and weaknesses.
- Storage engines optimize for different access patterns (point lookups vs scans vs analytics).
- Query languages express intent differently (SQL joins vs document embedding vs graph traversal).
- Consistency models range from strong ACID guarantees to eventual consistency with conflict resolution.
The Three Pillars
Every design choice in this book can be evaluated against three qualities: reliability (it works correctly, even when things go wrong), scalability (it keeps working as load grows), and maintainability (people can work on it productively over time). The rest of this chapter explores each pillar in depth.
