Observability

SYSTEM DESIGN SERIES  Β·  TOPIC 21 OF 23+  Β·  part of the 20 β†’ 21 β†’ 22 chain
1. The Same Slow Request, Seen Two Ways
800ms, four hops - but only one version tells you where it actually went (looping)
No tracing: the request disappears into the system and comes back 800ms later
No tracing: every service is a black box from the outside - no idea which one was slow
No tracing: "it's slow" is all you can say, to whoever's asking
With tracing: the same request carries one ID through every hop
With tracing: each service reports its own span - how long it personally took
With tracing: the 600ms was hiding in one specific service the whole time
2. Monitoring vs Observability
Knowing something's wrong vs knowing why
  • Monitoring - watches predefined signals: "is the dashboard red?"
  • Observability - has enough raw data to ask a question nobody planned for
  • Monitoring tells you that something broke. Observability helps you find out why
  • You can't add observability after the incident starts - the data has to already exist
🩺 Monitoring is a fever thermometer. Observability is the full chart, the labs, and the history - for when "it's hot" isn't enough to go on.
3. The Three Pillars
Different questions, different shapes of data
Logs
Discrete, timestamped events. Rich detail on one moment - expensive to search at scale.
Metrics
Numbers over time. Cheap, great for dashboards and alerts - no per-request story.
Traces
One request's full journey across services. Exactly what zone 1 just showed.
4. The Four Golden Signals
What to actually measure, if you measure nothing else
Latency
How long requests take - the whole subject of topic 01.
Traffic
How much demand is hitting the system, right now.
Errors
The rate of requests that are failing outright.
Saturation
How full the system is - topic 03's resources, watched.
5. Structured Logs & Correlation IDs
The glue that lets you jump between the three pillars
  • Plain text logs are for reading. Structured (JSON) logs are for querying
  • Same fields every time - timestamp, service, level, trace ID
  • That trace ID is the thread: a metric spike, a trace, and the exact log line - all one click apart
  • Without a shared ID, the three pillars are just three separate haystacks
6. Real Observability Tools
A tool per pillar, mostly - plus one trying to unify them
Prometheus / Grafana
(metrics)
ELK / Loki
(logs)
Jaeger / Zipkin
(traces)
OpenTelemetry
(the emerging standard)
Different pillars one shared trace ID stitching them back together
OpenTelemetry's whole pitch is instrument once, export logs/metrics/traces to whichever backend you actually run.
7. Alerting - Symptoms, Not Causes
Page a human for:
βœ” User-facing symptoms - error rate, latency, availability
βœ” Things that are actually urgent, right now
βœ” Signals tied to real impact, not just a number moving
You Can't Fix What You Can't See - And You Shouldn't Watch Everything at Once
Just log and graph:
βœ” Internal causes - CPU, memory, queue depth
βœ” Useful for diagnosis, not for waking anyone up
βœ” Too many low-value alerts train people to ignore all of them
πŸ’‘ Once you can see what's happening, you can safely change what's running. That's Deploys & Rollback - topic 22, next.