Designing Data-Intensive Applications
Ch. 10

Beyond MapReduce

Modern batch engines add iterative processing, higher-level APIs, and faster in-memory execution.

MapReduce's rigid map-shuffle-reduce cycle forces disk writes between every stage. Modern engines materialize less, push computation to data, and expose declarative APIs — while retaining fault tolerance through lineage graphs.

In practice

Databricks runs Apache Spark for ETL and ML training over terabytes in memory. Apache Flink unifies batch and stream processing in one runtime. Snowflake separates storage (S3/Azure) from elastic compute warehouses — you write SQL, not MapReduce jobs. dbt layers transformations on top of Snowflake or BigQuery with version-controlled SQL models.

Netflix at scale

Recommendation model training runs on Spark over viewing-history Parquet in S3 — iterative in-memory stages replace MapReduce's disk-bound shuffle between every step.

typescript — Declarative warehouse SQL
// Google BigQuery / Snowflake — declarative OLAP over columnar storage
const revenueByRegion = await snowflake.execute(`
  SELECT d.region, SUM(f.revenue_cents) AS total
  FROM fact_orders f
  JOIN dim_date d ON f.order_date = d.date_key
  WHERE d.year = 2025
  GROUP BY d.region
  ORDER BY total DESC
`);
// Optimizer reads only region + revenue columns — not full rows
Key Takeaways
  • Materializing intermediate state to disk (MapReduce) is slow for iterative algorithms.
  • Spark keeps data in memory across stages for faster iteration.
  • High-level APIs (Spark SQL, DataFrames) decouple logic from execution.
  • Graph processing (Pregel) uses message-passing between vertices.
  • Batch and stream processing are converging into unified engines.
  • Apache Spark, Apache Flink, and Snowflake's elastic compute supersede classic MapReduce.
SparkFlinkSnowflakedbtDataFramedataflow engine