Learning on Web Dev Open is free for all.

AI-Native Products > Evals are the deliverableThree kinds of grader and when each is honest
Phase 08Evals are the deliverable419 of 434

Three kinds of grader and when each is honest

Deterministic checks, a model as judge, and a human. They cost different amounts and they lie in different directions.

Concept14 minAI adversary

Deterministic graders: schema valid, field equals expected, contains this identifier, completed under this latency, cost below this ceiling, are free, fast and completely trustworthy for the properties they can express. Push as much of your suite into this category as possible, which is another reason to return structure rather than prose.

A model as judge covers the properties you cannot express: was the tone appropriate, did it use the retrieved sources, is the summary faithful. It is useful and it is biased, towards longer answers, towards its own style, towards whichever option appears first in a comparison. Give it a rubric with concrete criteria rather than asking whether the answer is good, randomise ordering, and calibrate it against human labels before you rely on it.

Humans remain the ground truth and are the reason the other two mean anything. You do not need many: thirty carefully labelled examples used to validate a judge will do more for your suite than three hundred generated cases nobody read.

You should now be able to

  • Choose a grader appropriate to the property being tested
  • Validate a model-as-judge against human labels
  • Explain the biases a judge model brings
Ask the community

Loading…