What a token actually costs
Tokens are not words, input and output are priced differently, and the arithmetic decides your product long before your taste does.
A token is a fragment of text, commonly around four characters of English, but far denser in code, punctuation-heavy JSON and any non-Latin script, where the same meaning can cost several times more. This matters concretely: a 12,000-token JSON blob you paste into every request is a fixed tax on every user action, and it is invisible until the bill arrives.
Input and output are priced separately and usually differ by a factor of three to five, with output the expensive side. That single asymmetry drives more architecture than anything else in this phase: it is why you ask for a classification rather than an essay, why you cap max tokens deliberately rather than defensively, and why summarising a conversation is cheaper than carrying it.
Do the multiplication early and out loud. Eight cents a call sounds like nothing; at a thousand calls a day it is a $2,400 monthly line nobody approved, and at ten thousand it is a decision that should have involved someone senior. A cost model on one page, written before the feature, is the difference between an AI feature and an AI incident.
You should now be able to
- Estimate token counts for a real request
- Explain why output tokens dominate a cost model
- Convert a per-request price into a monthly line item
Loading…