Designing Data-Intensive Applications
Ch. 8

Faults and Partial Failures

In distributed systems, only part of the system can fail — making failures harder to reason about.

On one machine, a crash stops everything — easy to detect. In a distributed system, one node may crash while others continue. A network link may drop while processes on both ends are healthy. This ambiguity makes distributed failure modes uniquely subtle.

In practice

A Kubernetes cluster might lose one pod while others keep serving traffic — but now load is uneven and downstream caches may be stale. An AWS Availability Zone outage takes out some RDS replicas but not the whole region. Docker containers restart on failure, but that does not replace distributed coordination for data consistency.

Netflix at scale

Chaos Monkey deliberately kills instances during business hours to prove partial failures are survivable. A single API pod crash must not corrupt shared state — only reduce capacity until Kubernetes reschedules.

typescript — gRPC deadline on partial failure
// Uber microservices — gRPC deadlines detect slow/dead peers
const deadline = Date.now() + 3_000;
const trip = await tripClient.getTrip(
  { tripId },
  { deadline }, // client aborts after 3s — cannot distinguish slow vs dead
);
// Envoy retry policy + circuit breaker stops cascading timeouts
Key Takeaways
  • Single-machine systems fail entirely; distributed systems have partial failures.
  • Cloud datacenters have better redundancy than single machines but shared failure domains.
  • Supercomputers assume reliable components; cloud systems assume failure is routine.
  • You cannot assume a remote node is alive without evidence.
  • Design for the case where any node, link, or rack can fail independently.
  • Kubernetes pod crashes, AWS AZ outages, and Redis sentinel failover are everyday partial failures.
partial failureKubernetesAWSDockerfault tolerancefailure domain