Offline and online, and why you need both
Your suite says it improved. Your users say nothing, which is not the same as agreeing.
Offline evals run a fixed dataset against a fixed pipeline and give you a number you can compare across changes. They are fast, cheap, repeatable and blind to everything your dataset does not contain, which, on the day you ship, is most of what real users will do.
Online measurement watches actual behaviour: acceptance rate on suggestions, edit distance between what you produced and what the user kept, regeneration rate, abandonment, escalation to a human. These are the outcomes; they are noisy and slow, and they are the only evidence your offline suite is measuring anything real. Run both, and when they disagree, believe the online signal and go and fix your dataset.
You should now be able to
- Distinguish offline evaluation from online measurement
- Choose an online signal that is not a vanity metric
Loading…