Ch. 10
MapReduce and Distributed Filesystems
MapReduce brought Unix-style batch processing to clusters with automatic parallelization and fault tolerance.
MapReduce wraps the Unix pipeline in a fault-tolerant distributed runtime. Map tasks extract data in parallel; shuffle groups by key; reduce tasks aggregate. The runtime handles scheduling, retries, and data locality.
MapReduceSparkHDFSS3shufflebatch job