Evals before features
A streaming assistant and a five-case suite where every case passes and four of them should not.
Evals before features
Your AI feature works. You know because you tried it twice and both times it was fine, which is exactly how every broken AI feature in production got shipped.
The brief: a support-reply assistant that streams an answer token by token, plus a small eval suite that decides whether the answer was any good. Ask your model for both. It will give you a beautiful streaming UI and an eval suite that is barely a smoke test.
That asymmetry is the lesson. Streaming is a solved interface problem with a known shape, so models produce it well. Deciding whether a non-deterministic system got the right answer is a judgement problem, so models produce something that looks like a test and asserts almost nothing.
The workbench has a working version of both: a fake stream, and five test cases graded by naive string matching. Every one of them passes. Four of them should not.
Loading…
You should now be able to
- Recognise an eval suite that cannot fail
- Build the seam that lets a suite test a real implementation
- Grade structurally and record cost and latency per run
Loading…