Scalability
Scalability is about handling more load without breaking, not raw speed. Covers vertical vs horizontal scaling and what actually bottlenecks systems.
Availability
Availability measures uptime, not correctness. Covers the nines table, SLI/SLO/SLA vocabulary, and the real cost of chasing another nine.
Reliability
Reliability means working correctly over time, not just staying up. Covers failure causes, MTBF vs MTTR, and how it differs from availability.
Single Point of Failure (SPOF)
A single point of failure caps how available a system can ever be. Learn where SPOFs hide in redundant-looking designs and how to remove them.
Latency vs Throughput vs Bandwidth
Latency, throughput, and bandwidth measure three different things — conflating them is a fast way to sound imprecise in a system design interview.
Consistent Hashing
Consistent hashing lets you add or remove servers without reshuffling almost all your data. Includes an interactive hash-ring demo you can run.
CAP Theorem
The CAP theorem forces a choice between consistency and availability during a network partition — here's what it actually means and what it misses.
Failover
Failover automatically switches traffic to a standby when the active component fails. Covers detection, split-brain, and active-active vs passive.
Fault Tolerance
Fault tolerance means a failure stays invisible to users, not just a fast recovery. Covers bulkheading, graceful degradation, and redundancy's limits.