Latency, Throughput & Percentiles

1. Latency - Time for ONE Request
How long a single round trip takes, start to finish
  • Measured in milliseconds (ms)
  • Includes network + queueing + processing time
  • One request, one user, one trip
  • Lower number = faster experience
2. Throughput - Requests per Second
How many requests the system handles per unit of time
  • Measured in RPS / QPS
  • Rises with parallelism & more servers
  • Many users, same window of time
  • Higher number = more capacity
3. Latency vs Throughput
1 request
= 120 ms
VS
1000 requests
= 1 sec
🛣 Highway analogy: Latency = time for ONE car to cross the bridge. Throughput = how many cars cross per minute. A highway can be slow (high latency) yet still move thousands of cars/hour if it has enough lanes (high throughput).
★ They're independent axes - fixing one doesn't automatically fix the other.
4. Average vs Percentiles
🚩 Average
Hides outliers. Nine 100ms requests + one 5s request still averages ~590ms - looks fine, isn't.
✅ Percentiles
p50 = median. p95 = 95% were this fast or faster. p99 = the slow tail.
5. Reading Percentile Values
p50
Typical request. Half your users see this or better.
p90
Most users' real experience.
p95
Where slow requests start becoming visible.
p99
The tail. Worst 1-in-100 - often the support ticket.
6. Why Percentiles Matter
Reveals Outliers
Real User Pain
Load Spikes
SLA Risk
Capacity Planning
Tail Latency (p99) What users actually feel - not the average
7. When to Optimize What
Optimize LATENCY when:
✔ User-facing APIs
✔ Checkout / payment flows
✔ Search-as-you-type
✔ Real-time chat / gaming
Use the Right Metric for the Right Job
Optimize THROUGHPUT when:
✔ Batch / ETL jobs
✔ Log & event processing
✔ Analytics pipelines
✔ Background / queued jobs
💡 Optimizing the average can quietly make p99 worse. Always check percentiles before declaring victory.