Designing Data-Intensive Applications
Ch. 4

Formats for Encoding Data

How you encode data determines compatibility, schema evolution, and cross-language interoperability.

Programs in memory represent data as objects and structs. To send data over a network or store it on disk, you must encode it as bytes. The encoding format determines size, speed, and — critically — whether schemas can evolve without breaking consumers.

Compatibility modes

Forward compatibility: old code reads new data. Backward compatibility: new code reads old data. Field tags (Protobuf) and union schemas (Avro) enable both.

In practice

Public REST APIs return JSON with OpenAPI specs. Internal microservices increasingly use gRPC with Protocol Buffers for compact, typed RPC. Kafka pipelines pair Avro schemas with Confluent Schema Registry so producers and consumers evolve independently during rolling deploys on Kubernetes.

Canva at scale

Design state syncs over WebSockets with compact binary payloads. Internal services use Protocol Buffers so new canvas fields ship without breaking older editor builds — forward compatibility via tagged field numbers.

typescript — Protocol Buffers encoding
// Canva / gRPC — Protocol Buffers with forward-compatible field tags
// message Design { string id = 1; repeated Layer layers = 2; string title = 3; }
const payload = Design.encode({
  id: designId,
  layers: canvasLayers,
  title: "Untitled",
}).finish();
// New field tag 4 added in v2 — old clients ignore unknown tags
Diagram
Key Takeaways
  • JSON and XML are human-readable but verbose and weakly typed.
  • Thrift and Protocol Buffers use tagged field numbers for compact binary encoding.
  • Avro requires a schema for reading — enabling compact encoding without field tags.
  • Forward and backward compatibility let old and new code coexist during deploys.
  • Schema registries manage evolution in production systems.
  • REST APIs use JSON; gRPC uses Protobuf; Kafka ecosystems standardize on Avro with Schema Registry.
JSONProtocol BuffersAvrogRPCschema evolutionserialization