Grading prose is a trap
Substring matching cannot tell an assertion from its negation, and that is the whole reason structured output exists.
If your check is that the reply mentions a refund, then "I cannot offer a refund" passes. Substring matching has no notion of negation, scope, condition or subject, so it scores the exact opposite of the intended behaviour identically to the intended behaviour. Every naive eval suite ever written has this bug, usually in four of its five cases.
The fix is to move the decision out of the prose. Have the model return a field, remedy is refund, reship or none, alongside whatever text the user sees, and assert on the field. You are converting a language understanding problem, which you cannot test, into a data validation problem, which you can.
This changes your product design as well as your tests. Once the decision is a field, you can route on it, log it, count it, alert on a distribution change, and render the prose separately or not at all. Prose is the interface; the field is the contract.
You should now be able to
- Explain why string matching fails as a correctness check
- Convert a prose requirement into a checkable field
Loading…