Shipping non-determinism to real people
Someone who did not choose to be part of an experiment is going to rely on your output. That is the whole ethical content of this phase.
Every feature in this phase puts a system that is confidently wrong some percentage of the time in front of someone who cannot easily tell which case they got. That is acceptable in plenty of contexts and unacceptable in others, and the difference is not the technology: it is who bears the cost of the error and whether they can see it coming. Draft an email: fine. Summarise a medical letter: not without care you probably cannot afford yet.
Two obligations follow, and they are cheap. Be honest in the interface about what the thing is, so nobody mistakes a generated draft for a verified fact. And keep the human decision where the consequences are real: approve, do not auto-send; suggest a category, do not close the ticket. Most of the harm from AI features comes from removing the confirmation step to save a click.
The last one is about you. "The model did it" is not a defence anyone accepts, and it should not be. You chose the model, the prompt, the boundaries and the point at which output becomes action. That is the whole job, and it is why the eval suite, the cost model and the kill switch are the deliverables rather than the demo.
You should now be able to
- Decide what your feature must never do automatically
- Set an honest expectation in the interface
- Take responsibility for an output rather than deferring to the model
Loading…