Traces, and the span you did not instrument
A trace shows where a request spent its time across every service it touched, and the gap in the middle is usually the answer.
A trace is one request as a tree of timed spans across processes. Where a waterfall shows the browser's view, a trace shows the server side of the same story: this handler took 400ms, of which 380 was one query, of which 300 was waiting for a connection from an exhausted pool. That last sentence is a conclusion you cannot reach from logs and cannot reach from a metric.
The gaps between spans are as informative as the spans. Time inside a request that belongs to no span is time you have not instrumented: waiting in a queue, blocked on the event loop, held up by a pool, or paused for garbage collection. A trace with one enormous span called "handler" is a trace that tells you nothing, which is why instrumenting boundaries, outbound calls, database queries, cache lookups, serialisation of anything large, is worth the effort and instrumenting every function is not.
Sample intelligently. Trace a small percentage of normal traffic and everything that is slow or errored, so the traces you keep are the ones you would want. And propagate the trace context across queue messages and background jobs, otherwise every asynchronous piece of your system is invisible in exactly the place where things go wrong.
You should now be able to
- Read a trace and locate the dominant span
- Interpret a gap between spans
- Instrument the boundaries that matter without tracing everything
Loading…