Learning on Web Dev Open is free for all.

AI-Native Products > Evals are the deliverableEvals before features
Phase 08Evals are the deliverable417 of 434

Evals before features

A streaming assistant and a five-case suite where every case passes and four of them should not.

Build55 minAI inside
Phase 08 · Lesson 02

Evals before features

AI: in the product~55 min

Your AI feature works. You know because you tried it twice and both times it was fine, which is exactly how every broken AI feature in production got shipped.

Beat 1: Build it

The brief: a support-reply assistant that streams an answer token by token, plus a small eval suite that decides whether the answer was any good. Ask your model for both. It will give you a beautiful streaming UI and an eval suite that is barely a smoke test.

That asymmetry is the lesson. Streaming is a solved interface problem with a known shape, so models produce it well. Deciding whether a non-deterministic system got the right answer is a judgement problem, so models produce something that looks like a test and asserts almost nothing.

The workbench has a working version of both: a fake stream, and five test cases graded by naive string matching. Every one of them passes. Four of them should not.

Ask the community

Loading…

Workbench

You should now be able to

  • Recognise an eval suite that cannot fail
  • Build the seam that lets a suite test a real implementation
  • Grade structurally and record cost and latency per run
Ask the community

Loading…