Average response time: 180 ms. Ticket from support: “the API is slow”. Both true. The average is a diplomat. It offends nobody and tells you nothing.
Latency is a distribution. Our 180 ms hides a p50 of 90 ms and a p99 above four seconds. One request in a hundred is terrible, and with thirty requests per page load, most users hit that unlucky one regularly. The people complaining are not imagining things. They live in the tail, and the average never visits there.
So we started tracing. Nothing fancy: a correlation id generated at the edge, passed into every log line, every SQL comment, every outgoing HTTP header. Plus timing spans, controller, each query, each Redis call, each external API. When a specific slow request comes in, we pull its trace and see the shape of the time.
The shapes were educational. One slow case was a warm path that sometimes missed the cache and rebuilt an expensive aggregate inline, four seconds of SQL nobody remembered writing. Another was an external API that answers in 50 ms except when it answers in three seconds, and our timeout was five. No amount of staring at averages would name these. The trace names them in one screenshot.
It also changed the meetings. Before, they sounded like “Symfony is slow” versus “the database is slow”, two teams pointing at each other’s boxes. Now someone opens a trace, points at span four, and the meeting is over in five minutes.
Sampling keeps it cheap. Trace one percent of everything and one hundred percent of requests slower than a threshold. The slow ones are the whole point.
If you monitor one number, make it p99. Ours was the average, on a dashboard nobody had a reason to look at. It was always green.