Designing Data-Intensive Applications
Ch. 1

Maintainability

Most of the lifetime cost of software is in maintenance — fixing bugs, adapting to new requirements, and keeping it running. Maintainability is about making that work productive.

Building a system is only the beginning. Over years, teams fix bugs, adapt to regulation, add features, and replace components. If the codebase is a tangled mess, every change is risky and slow. Maintainability is the quality that keeps the total cost of ownership reasonable.

Operability

Operations teams need good visibility into system health: metrics, logs, traces, and clear runbooks for common failures. Good operability means routine tasks (deploys, scaling, config changes) are safe and boring — not heroic.

In practice

Docker containers package services consistently from laptop to production. Kubernetes orchestrates deploys with rolling updates and instant rollbacks. Terraform defines AWS infrastructure as code — reproducible and reviewable in pull requests. OpenTelemetry exports traces to Jaeger or Datadog so on-call engineers debug cross-service failures in minutes, not hours.

Canva at scale

Canva ships features weekly to millions of users by keeping services independently deployable on Kubernetes, using feature flags for safe rollouts, and OpenTelemetry tracing across the design editor, asset pipeline, and collaboration servers. Evolvable architecture lets them swap storage engines without rewriting the editor.

Diagram
typescript — Structured logging + trace IDs for operability
import { trace } from "@opentelemetry/api";

function log(level: string, msg: string, meta: Record<string, unknown> = {}) {
  const span = trace.getActiveSpan();
  console.log(JSON.stringify({
    level,
    msg,
    traceId: span?.spanContext().traceId,
    service: "canva-design-api",
    ...meta,
  }));
}

// Every log line links to a distributed trace in Datadog/Jaeger
  • Expose key metrics: latency, error rate, saturation, and business KPIs.
  • Provide clear documentation for deployment, rollback, and disaster recovery.
  • Automate repetitive tasks to reduce human error.

Simplicity

Complexity slows everyone down. It comes from accidental complexity (poor abstractions, tangled dependencies) and essential complexity (the problem is genuinely hard). Good design minimizes accidental complexity by choosing the right abstractions and interfaces.

Visual metaphor for managing system complexity

Simplicity is not about fewer features — it is about manageable complexity.

Evolvability

Requirements change. A system that cannot adapt becomes a legacy burden. Evolvability means the architecture supports incremental change: you can swap a storage engine, add a new API version, or migrate data formats without a rewrite.

Agility vs maintainability

Agile processes help teams deliver quickly, but without evolvable architecture, speed today becomes paralysis tomorrow. Design for change from the start.

Key Takeaways
  • Operability: make it easy for operations teams to keep the system healthy.
  • Simplicity: manage complexity so new engineers can be productive quickly.
  • Evolvability: make change easy — requirements always change over a system's lifetime.
  • Good abstractions and clear interfaces reduce the cost of change.
  • Technical debt is inevitable; the goal is to keep it manageable.
  • Docker, Kubernetes, Terraform, OpenTelemetry, and CI/CD pipelines are the modern maintainability toolkit.
operabilitytechnical debtabstractionevolvabilityDockerKubernetesTerraformOpenTelemetry