Classification beats generation
Picking from a list is cheaper, faster, more testable and more reliable than writing a sentence. Use it far more than feels natural.
A model that must pick one of six labels produces a handful of output tokens, returns in a fraction of the time, can be run by a much smaller model, and is graded by comparing to a known answer. A model that must write a paragraph does none of those things. Wherever the product can be satisfied by a decision plus a template, that is almost always the better engineering.
It also gives you the one thing generation cannot: a confusion matrix. You can see which categories are being mixed up, fix the definition or the examples for those specific pairs, and prove the improvement. That is an ordinary supervised evaluation loop, and it turns an argument about prompt wording into a measurement.
You should now be able to
- Reframe a generation task as a classification task
- Measure a classifier against a labelled set
Loading…