Three kinds of grader and when each is honest
Deterministic checks, a model as judge, and a human. They cost different amounts and they lie in different directions.
Deterministic graders: schema valid, field equals expected, contains this identifier, completed under this latency, cost below this ceiling, are free, fast and completely trustworthy for the properties they can express. Push as much of your suite into this category as possible, which is another reason to return structure rather than prose.
A model as judge covers the properties you cannot express: was the tone appropriate, did it use the retrieved sources, is the summary faithful. It is useful and it is biased, towards longer answers, towards its own style, towards whichever option appears first in a comparison. Give it a rubric with concrete criteria rather than asking whether the answer is good, randomise ordering, and calibrate it against human labels before you rely on it.
Humans remain the ground truth and are the reason the other two mean anything. You do not need many: thirty carefully labelled examples used to validate a judge will do more for your suite than three hundred generated cases nobody read.
You should now be able to
- Choose a grader appropriate to the property being tested
- Validate a model-as-judge against human labels
- Explain the biases a judge model brings
Loading…