Designing Data-Intensive Applications
Ch. 12

Data Integration

Modern architectures combine specialized tools by deriving data through unbundled pipelines.

The future is not one database to rule them all. Instead, systems specialize: OLTP for transactions, search index for full-text, warehouse for analytics, cache for speed. An event log connects them, and derived data flows keep views consistent.

In practice

A modern e-commerce stack: PostgreSQL for orders, Redis for session cache and rate limiting, Elasticsearch for product search, Kafka for order events, Snowflake for analytics, and S3 for image storage — all synced via CDC and stream processors. Kubernetes and Docker package each service; Envoy or nginx routes traffic at the edge. Authentication flows through OAuth 2.0 / OpenID Connect (Auth0, Keycloak, AWS Cognito).

Airbnb at scale

Listings live in PostgreSQL, search in Elasticsearch, sessions in Redis, analytics in Snowflake — all derived from an immutable Kafka event log via CDC. No single database does everything.

typescript — CDC outbox to Kafka
// Transactional outbox — Airbnb / Shopify event publishing
await prisma.$transaction(async (tx) => {
  const order = await tx.order.create({ data: orderInput });
  await tx.outboxEvent.create({
    data: {
      aggregateId: order.id,
      type: "OrderPlaced",
      payload: JSON.stringify(order),
    },
  });
});
// Separate relay process reads outbox → publishes to Kafka
Diagram
Key Takeaways
  • No single database does everything well — compose specialized systems.
  • Derive views, indexes, and analytics from an immutable event log.
  • Batch and stream processing both feed derived data systems.
  • Unbundling the database separates storage, indexing, querying, and processing.
  • Dataflow-oriented design makes integration explicit in the architecture.
  • Lambda → Kappa architecture with Kafka; microservices with Postgres + Elasticsearch + Redis + Snowflake.
data integrationPostgreSQLKafkaElasticsearchRedisSnowflakeunbundling