Designing Data-Intensive Applications
Ch. 1

Reliability

Reliability means the system continues to work correctly — performing the right function at the right time — even when faults occur. Faults are inevitable; failures are not.

Users expect systems to work. Reliability is about meeting that expectation even when the real world intrudes: a disk dies, a deploy introduces a bug, or an operator misconfigures a firewall rule. The goal is not perfection — it is building systems where individual faults do not cascade into user-visible failures.

Fault vs failure

We can tolerate faults (a machine crashes) while preventing failures (the service goes down). Designing for fault-tolerance means assuming things will break and planning accordingly.

Hardware Faults

Hard disks fail, RAM corrupts, power supplies die, and network links flap. In large fleets, failures are routine — not exceptional. Cloud platforms offer VMs with automatic restart, but that does not help if the data on disk is lost or if all replicas share the same power domain.

  • Single-server redundancy (RAID, dual power supplies) reduces but does not eliminate risk.
  • Distributed systems replicate data across machines so one failure does not mean data loss.
  • Correlated failures (shared rack, region outage) require multi-zone or multi-region strategies.

Software Errors

Software bugs can be subtle and long-dormant. A leap second, an edge case in input validation, or a race condition under load can trigger cascading failures across services. Unlike hardware, software faults often affect many machines simultaneously because they all run the same code.

Human Errors

Studies of large outages consistently find human error near the root: a typo in a config, a migration run against production, a certificate that expired unnoticed. Good systems minimize opportunities for error and provide safety nets — staged rollouts, easy rollbacks, and thorough monitoring.

In practice

Amazon RDS Multi-AZ automatically fails over PostgreSQL to a standby on hardware failure. Kubernetes restarts crashed pods and reschedules them onto healthy nodes. Sentry catches software errors before they cascade; PagerDuty and Grafana on-call rotations ensure someone responds. AWS Certificate Manager auto-renews TLS certs — eliminating a common human-error outage vector.

WhatsApp at scale

WhatsApp serves billions of users with Erlang servers designed for fault tolerance — a crashed process does not take down the whole node. Messages are persisted and replicated before ack. The engineering goal matches DDIA's definition: tolerate faults, prevent user-visible failures.

typescript — Circuit breaker — stop cascading failures (Netflix Hystrix pattern)
class CircuitBreaker {
  private failures = 0;
  private open = false;

  async call<T>(fn: () => Promise<T>): Promise<T> {
    if (this.open) throw new Error("Circuit open — fail fast");
    try {
      const result = await fn();
      this.failures = 0;
      return result;
    } catch (err) {
      this.failures++;
      if (this.failures >= 5) this.open = true;
      throw err;
    }
  }
}
Diagram
How important is reliability?

For many business-critical systems (payments, medical records, infrastructure), downtime has serious consequences. Even for less critical apps, unreliability erodes user trust quickly. The investment in reliability should match the cost of failure.

Key Takeaways
  • A fault is one component deviating from spec; a failure is when the system stops delivering its service.
  • Hardware faults (disks, power, network) are common and often correlated in cloud environments.
  • Software errors (bugs, cascading failures) are often worse because they can affect many users at once.
  • Human errors cause most outages — good design makes mistakes hard and recovery easy.
  • Fault-tolerance accepts that faults happen and builds systems that survive them.
  • AWS Multi-AZ RDS, Kubernetes self-healing, and PagerDuty alerting are everyday reliability tooling.
fault tolerancehardware faultcascading failureredundancyAWS RDSKubernetesPostgreSQL