A model call is a job, not a function call
Slow, occasionally very slow, sometimes wrong, priced per token and rate limited. Everything in this lesson applies to it twice over.
A call to a language model behaves like a slow, flaky third-party API with an unusually wide latency distribution. The median might be two seconds and the ninety-ninth percentile thirty, and there is no timeout value that is both responsive and safe. That is the profile the whole of this lesson was describing: queue it, or stream it, but do not sit inside a request handler waiting for it with a default timeout.
Streaming changes the perception rather than the total, and it is often the right choice for anything a user is watching. Behind it you still need the ordinary discipline: a hard token ceiling so a runaway generation cannot cost you a fortune, retries that respect the provider's rate limit headers rather than fighting them, and a cache keyed on the prompt for the requests that repeat, which is more of them than you would expect.
Treat the output as untrusted input, because it is. Validate it against a schema before it goes near your database, and never interpolate it into a query, a shell command or HTML. A model that has read a user's document and been asked to summarise it is a channel from that user to your parser, and that is a data flow worth drawing on the diagram.
You should now be able to
- Design a backend endpoint that calls a model without blocking
- Handle rate limits and partial failures from a provider
- Decide between streaming to the client and queueing the work
Loading…