Constraints are the design
Rough numbers, done on the back of an envelope, that tell you which architectures are already ruled out before you draw anything.
A million daily users doing ten actions each is ten million requests a day, which is around a hundred and twenty a second averaged, and perhaps five hundred at peak. That is not a large number, and knowing it is not large is the point: most systems people over-engineer are running at a scale a single well-indexed database would serve comfortably. Estimation is mostly a tool for talking yourself out of complexity.
Keep a few numbers to hand because they set the shape of everything: memory access in nanoseconds, an SSD read in tens of microseconds, a same-region network round trip around half a millisecond, cross-continent around a hundred and fifty milliseconds, and a cold serverless start in the hundreds. The gap between a memory hit and a transatlantic round trip is roughly a factor of a million, which is why caching and region placement dominate performance conversations.
Storage is the estimate people skip and then discover. Work out bytes per record, multiply by records per day, multiply by retention, and see whether the answer is gigabytes or terabytes. A logging decision that seems harmless is a bill and a query plan a year later, and the arithmetic takes ninety seconds.
You should now be able to
- Estimate throughput, storage and bandwidth to one order of magnitude
- Use an estimate to eliminate an option
- Know the handful of latency numbers worth memorising
Loading…