Scalability
Scalability is not a single number — it is how a system copes with increased load. The right strategy depends entirely on what "load" means for your application.
Scalability describes how performance changes when load increases. A system that handles 10 users beautifully but collapses at 10,000 is not scalable. But "scalable" does not mean "infinitely fast" — it means you have a reasonable plan for growth.
Describing Load
Load is not just "number of users." For a social network, it might be posts per second and fan-out reads. For a search engine, it is queries per second and index size. Define the parameters that matter for your system and measure them.
Describing Performance
Throughput tells you how much work a system handles per unit time. Response time (latency) tells you how long an individual request waits. Both matter, but for interactive applications, tail latency — the slowest 1% or 5% of requests — often determines whether users perceive the system as fast.
Approaches for Coping with Load
- Vertical scaling: upgrade CPU, RAM, or disk on one machine. Simple but has hard limits.
- Horizontal scaling: distribute load across many machines. Requires partitioning and coordination.
- Elastic scaling: automatically add/remove capacity. Great for variable load, harder to operate.
- Tuning and optimization: sometimes a 10x improvement comes from a better algorithm, not more hardware.
