Learning on Web Dev Open is free for all.

AI-Native Products > The API is the constraintLatency has two numbers
Phase 08The API is the constraint384 of 434

Latency has two numbers

Time to first token and tokens per second are different problems with different fixes, and only one of them is what users feel.

Concept13 minAI pair

Time to first token is how long the user stares at nothing, and it is dominated by how much input has to be processed before generation starts. Tokens per second is how fast the answer then arrives. A response that starts in 400ms and takes six seconds to complete feels dramatically better than one that starts in three seconds and finishes in four, even though the second one is faster overall.

Averages lie about this more than about almost any other metric, because the distribution has a long tail: queue time under load, retries after a rate limit, an unusually long generation. Track p50 and p95, set the target on p95, and treat a p95 regression as a failing test in the same way you would treat a broken build.

The fixes are different for each number. First token improves by shrinking input, caching a stable prompt prefix, and choosing a smaller model for the first hop. Throughput improves by generating less: a shorter format, a structured field instead of prose, streaming so the user can start reading. Only one of these is a prompt engineering problem.

You should now be able to

  • Separate time to first token from total generation time
  • Set a p95 target rather than an average
  • Say which product decisions change which number
Ask the community

Loading…