Designing Data-Intensive Applications
Case 1

How to Approach System Design

Start with users and constraints — not databases. Clarify functional needs, non-functional targets, and what you can defer.

System design is the practice of mapping product requirements to components that store, move, and transform data. The same DDIA concepts — replication, partitioning, streams — appear in every case study; the skill is knowing which ones matter for each problem.

A repeatable framework

  • Clarify scope: MVP vs full product, geography, scale today vs in 12 months.
  • Functional requirements: core user journeys in plain language.
  • Non-functional: latency targets, durability, consistency, cost, compliance.
  • High-level diagram: clients → API → services → data stores → async pipelines.
  • Deep dives: bottlenecks, failure modes, scaling levers.
  • Trade-offs: what you optimize for and what you explicitly sacrifice.
In practice

Stripe's design reviews start with the user promise ('charge once, never double-charge') before picking Postgres vs DynamoDB. Netflix capacity plans assume regional failure. The framework is universal; the storage engines change.

Airbnb at scale

Before sharding listings search, the team nailed booking invariants: a night cannot be sold twice, payments must match reservations, and cancellations must release inventory. Those constraints drove PostgreSQL transactions and Kafka outbox patterns.

Diagram

Step-by-step walkthrough

Synchronous request path

  • ① HTTPS — Browser or mobile client hits the CDN / reverse proxy (TLS termination, DDoS shield).
  • ② REST/gRPC — Proxy forwards API calls to stateless application pods behind the load balancer.
  • ③ OLTP read/write — API reads and writes authoritative rows in PostgreSQL (orders, accounts, inventory).
  • ④ Cache hot keys — Sessions, rate counters, and hot reads go through Redis with TTL.
  • ⑤ Search query — Full-text and facet queries hit Elasticsearch, not a table scan on Postgres.

Async derived-data path (dashed arrows)

  • ⑥ Publish events — After a successful commit, the API appends a domain event to Kafka (outbox pattern).
  • ⑦ ETL / CDC — Stream consumers copy changes into Snowflake for analytics and BI dashboards.
  • ⑧ Index sync — Search index updates from the same event log — never dual-write to ES and Postgres in one request.
typescript — Requirements checklist as types
// Frame the problem before picking databases
type SystemDesignBrief = {
  actors: ("mobile" | "web" | "admin" | "partner")[];
  functional: string[]; // "upload video", "send message"
  nonFunctional: {
    p99LatencyMs?: number;
    availability?: string; // "99.9%"
    consistency: "strong" | "eventual" | "per-feature";
  };
  scaleToday: { dau: number; peakQps: number };
  invariants: string[]; // "no double booking"
};
Diagram

Why common infrastructure building blocks?

Why CDN (Cloudflare / CloudFront)?

Static assets (JS, CSS, images) are identical for every user — perfect for edge caching. A CDN also absorbs DDoS and terminates TLS close to users, shrinking latency and origin load. You would skip it only for a purely internal API with no browser clients.

Why Load balancer / reverse proxy (nginx, Envoy, ALB)?

Distributes traffic across many stateless API pods, performs health checks, and can enforce rate limits and auth before requests hit application code. Required once you run more than one API instance — which is effectively always in production.

Why API gateway?

Single front door for clients: routing to microservices, JWT validation, request shaping, and API versioning. Alternatives: embed routing in the load balancer (simpler monolith) or use a service mesh (Istio) for east-west traffic between services.

Why PostgreSQL?

Default choice when you need ACID transactions, joins, foreign keys, and ad-hoc queries — bookings, accounts, permissions. Poor fit for billion-row append-only logs, video blobs, or sub-millisecond leaderboard hot paths.

Why Redis?

In-memory speed for hot keys: sessions, rate counters, presence, leaderboards (sorted sets), and short-lived holds. Data is ephemeral or cache-aside — never the sole source of truth for money or inventory unless you accept loss on restart.

Why Kafka (or SQS / Pub/Sub)?

Buffers spikes, decouples producers from consumers, and lets you add subscribers (search index, email, analytics) without changing the write API. Use when work can be asynchronous; skip when the user waits for the result in the same HTTP request.

Why Elasticsearch (or OpenSearch)?

Full-text search, fuzzy matching, facets, and geo filters at scale. Indexes are derived from an OLTP source of truth via CDC — never the system of record for payments or inventory.

Why Object storage (S3 / GCS)?

Cheap, durable storage for large immutable blobs: video, images, PDFs, backups. Not for low-latency row lookups — pair with a database for metadata pointers.

Why Kubernetes (Docker)?

Packages services consistently and scales stateless pods horizontally. Agones extends K8s for game servers; batch workers and functions (Lambda) replace K8s when you want zero cluster management.

Key Takeaways
  • Separate functional requirements (what) from non-functional ones (how well).
  • Identify actors: mobile clients, admins, third-party integrations, batch jobs.
  • Sketch read vs write ratio, consistency needs, and acceptable downtime.
  • Defer deep dives until the happy path and failure modes are clear.
  • Name concrete products (YouTube, WhatsApp) to anchor estimations and trade-offs.
  • Interview and production design both reward structured thinking over memorized diagrams.
requirementsnon-functionaltrade-offsscopefailure modesCDNPostgreSQLRedisKafka