Choosing a model is an engineering decision
Routing between a small model and a large one is usually worth more than any prompt you could write.
Tasks are not equally hard. Classification, extraction, routing, short rewrites and structured tagging are handled well by small, fast, cheap models; multi-step reasoning, long synthesis and open-ended code work are not. Sending everything to the largest model is the cloud equivalent of running one instance size for every workload.
The pattern that pays is escalation: run the cheap model first, detect low confidence or a failed validation, and retry with the stronger one. You get the cheap latency and cost most of the time and the strong answer when it matters. It only works if you have a way to detect failure, which is another argument for structured output and evals.
Keep the provider boundary thin, a function that takes messages and returns text or a parsed object, without building an abstraction layer over four providers you do not use. Models change every few months, and the cost of switching should be an afternoon, not a quarter.
You should now be able to
- Match a task to the smallest model that passes your evals
- Design a routing or escalation path between models
- Avoid coupling your code to one provider unnecessarily
Loading…