Cost is a first-class constraint
Every architecture has a monthly bill, and the shape of that bill is a design output you can predict before you build.
Cost belongs next to latency and correctness in the constraints section, because it eliminates designs just as firmly. A per-request architecture is cheap at low volume and can be alarming at high volume; a provisioned one is the reverse. Knowing the crossover point for your own numbers is a ten-minute calculation that occasionally changes the whole design, and almost nobody does it before building.
Separate what scales with usage from what does not. Compute, egress, storage growth and per-token model calls scale; a managed database floor, a monitoring seat and a fixed cluster do not. Egress is the one that ambushes people: moving data out of a cloud, or between regions, is priced in a way that makes a chatty cross-region design expensive in a manner no one predicted from the architecture diagram.
Observability and inference are the two modern surprises. Logging every request at full fidelity can cost more than the service producing the logs, which is why sampling is an architectural decision. And a feature calling a large model on every keystroke has a variable cost per user that can exceed what that user pays you, which is a product problem discovered in the billing console rather than in review.
You should now be able to
- Estimate the cost of an architecture before building it
- Identify which cost scales with users and which is fixed
- Recognise the line items that surprise people
Loading…