Designing Data-Intensive Applications
Ch. 1

Thinking About Data Systems

Most applications today are data-intensive rather than compute-intensive. Understanding what that means is the foundation for every design decision that follows.

When people talk about "the cloud" or "big data," they often focus on raw compute power. In practice, most modern applications are data-intensive: the hard problems come from handling large volumes of data, complex data transformations, or rapid changes — not from crunching numbers on an idle CPU.

What makes an app data-intensive?

Typical challenges include the database becoming a bottleneck, caching layers multiplying, search indexes to maintain, file storage at scale, analytics pipelines, and inter-service data exchange. The application exists to serve data to users and capture it back.

Beyond a Single Database

We use the term data system broadly: not just a database, but any combination of components that store, process, or move data. A single product might use a relational database for transactions, a search index for full-text queries, a cache for hot reads, and a message queue for async work. Each tool has strengths and weaknesses.

In practice

A typical SaaS backend on AWS might run stateless services in Docker on Kubernetes, store transactional data in PostgreSQL (RDS), cache sessions in Redis (ElastiCache), index products in Elasticsearch, stream domain events through Kafka (MSK), and serve files from S3 — with an nginx or Envoy reverse proxy at the edge.

Airbnb at scale

Airbnb is a textbook data-intensive app: PostgreSQL for bookings and listings, Elasticsearch for search, Kafka for event-driven sync between services, Redis for caching, and Snowflake for analytics. No single database does everything — the architecture composes specialized tools.

Diagram
typescript — Compose specialized data stores — don't force one database to do everything
// Typical service layer composes multiple backends
type DataSystem = {
  oltp: PostgresClient;      // bookings, users
  cache: RedisClient;        // sessions, hot reads
  search: ElasticsearchClient;
  events: KafkaProducer;     // async integration
  warehouse: SnowflakeClient;  // analytics only
};

// Airbnb routes listing search to Elasticsearch,
// but booking payments always hit PostgreSQL
  • Storage engines optimize for different access patterns (point lookups vs scans vs analytics).
  • Query languages express intent differently (SQL joins vs document embedding vs graph traversal).
  • Consistency models range from strong ACID guarantees to eventual consistency with conflict resolution.
Abstract illustration of interconnected data system components

A modern data-intensive application combines specialized components into a cohesive system.

The Three Pillars

Every design choice in this book can be evaluated against three qualities: reliability (it works correctly, even when things go wrong), scalability (it keeps working as load grows), and maintainability (people can work on it productively over time). The rest of this chapter explores each pillar in depth.

Key Takeaways
  • Data-intensive applications are usually limited by data volume, complexity, or speed — not CPU.
  • A "data system" is any system built from multiple components working together to store, process, or move data.
  • Reliability, scalability, and maintainability are not features you bolt on — they shape architecture from day one.
  • There is no one-size-fits-all tool; good engineers combine specialized components thoughtfully.
  • Production stacks combine PostgreSQL, Redis, Kafka, Elasticsearch, and object storage (S3) — each for a different job.
data-intensive applicationreliabilityscalabilitymaintainabilityPostgreSQLRedisKafkaKubernetes