Designing Data-Intensive Applications
Ch. 10

Batch Processing with Unix Tools

The Unix philosophy — composable tools over monolithic programs — scales to big data.

Long before Hadoop, Unix pipelines processed data by chaining simple programs: extract fields with awk, filter with grep, aggregate with sort. Each tool is stateless, reads stdin, writes stdout — a pattern that scales to terabytes when parallelized.

In practice

SREs still debug production with grep and jq on JSON logs. DuckDB queries Parquet files on S3 like a local Unix tool — no cluster required. Fluent Bit and Vector collect container logs on Kubernetes nodes and pipe them to Elasticsearch or S3. The composable-pipeline idea survived; only the scale and storage changed.

Google at scale

Google's early MapReduce paper explicitly credits Unix pipelines as inspiration — small composable tools over monolithic programs. SRE log triage still chains grep, awk, and sort before reaching for a cluster.

typescript — Unix-style log pipeline
// Google SRE log triage — composable Unix-style pipeline
// cat access.log | grep '" 5' | awk '{print $7}' | sort | uniq -c | sort -rn
const lines = await readFile("access.log", "utf8");
const errors = lines.split("\n")
  .filter((l) => l.includes('" 5'))
  .map((l) => l.split(" ")[6]);
const counts = Object.fromEntries(
  [...new Set(errors)].map((p) => [p, errors.filter((e) => e === p).length])
);
Key Takeaways
  • Small tools do one thing well and compose via pipes.
  • stdin/stdout as a uniform interface decouples producers and consumers.
  • Log analysis with awk, grep, and sort handles surprisingly large datasets.
  • Immutability of input files enables retry and recomputation.
  • The same patterns underpin Hadoop, Spark, and modern data pipelines.
  • grep/awk on logs, DuckDB on Parquet, and Fluent Bit pipelines extend the Unix pattern.
Unix philosophyDuckDBFluent Bitbatch processingpipelineimmutability