Failure Taxonomy - How Distributed Systems Actually Break

1. Partial vs Total Failure
Most real outages are partial - that's what makes them tricky
  • Partial failure: some nodes down, rest healthy
  • The system must keep working around the gap
  • Total failure: everything down - rare, and oddly simpler
  • Distributed systems fail partially, by default
2. The Core Problem - Slow vs Dead
A silent node and a slow node look identical from outside
  • No response could mean crashed, slow, or lost in transit
  • The network gives no delivery guarantee or deadline
  • A timeout is a guess, not proof of failure
  • You can only ever suspect - never confirm
3. Four Classic Failure Types
Crash
Node stops and never responds again. Easiest to detect - it just goes silent for good.
Omission
A request or response is sent but never arrives. The sender can't tell it was lost.
Timing
Response arrives - just too late. Violates the deadline the caller relied on.
Byzantine
Node responds - with wrong or corrupted data. It can even actively lie.
4. Failure Severity Spectrum
Same four types, ranked by how hard they are to catch
A crash announces itself eventually. Byzantine behavior can look perfectly correct - that's what makes it the hardest class to defend against.
5. How We Detect Failures
  • Heartbeats - periodic "I'm alive" pings
  • Timeouts - no reply within N ms ⇒ suspect it
  • Health checks - actively probe, don't just wait
⏱ Shorter timeout = faster detection, more false alarms. Longer timeout = fewer false alarms, slower reaction. Detection is always a guess with a deadline.
6. Fault-Tolerance Toolkit
Retries
Redundancy
Circuit Breakers
Timeouts
Graceful Degradation
Design FOR failure not against it - assume every call can fail
Seen in the wild: exponential backoff in every major cloud SDK, circuit breakers popularized by Netflix's Hystrix, and Kubernetes liveness/readiness probes doing continuous health checks so a stuck pod gets replaced automatically.
7. Cascading Failures - Avoid vs Do
Avoid:
✘ Unbounded retries
✘ No backoff / no jitter
✘ One shared connection pool
✘ Chains with no timeout
Failures Compound - Contain the Blast Radius
Do:
✔ Exponential backoff + jitter
✔ Circuit breakers
✔ Bulkheads (isolate blast radius)
✔ Timeouts tuned per hop
💡 Retry storms are the #1 self-inflicted outage: everyone retries at once, load spikes, more timeouts fire, more retries follow.