Somebody always asks what the feature will cost before anyone has written it, and the honest answer is available the same afternoon. The bill is arithmetic: tokens in times the input rate, plus tokens out times the output rate, times the number of calls. Estimates rarely come apart at the multiplication. They come apart in three places where people count the wrong thing.
Take a support-triage feature and count it properly. The system prompt carries the instructions, six category definitions and four worked examples, and runs to around 1,200 tokens. Two similar past tickets are retrieved and pasted in for context, another 600. The ticket itself, a longish email, is 300. So the model reads about 2,100 tokens to produce a category and a one-line reason, which comes back at roughly 40. At 2,000 tickets a day that is 4.2 million tokens in and 80,000 out, and across a month, 126 million in against 2.4 million out.
Sit with that ratio for a second, because it is the whole first lesson. Output is under two per cent of the traffic, and output is the part priced several times higher. Anyone who estimated this feature from the ticket alone, the only text a human actually wrote, undercounted the input sevenfold. Measure a realistic call end to end.
The second place estimates break is the call count. A triage step that looks something up, reads what came back and tries again is one feature making five calls, and each of those calls re-reads everything before it, so the five-call version costs more than five times the one-call version. Count the calls along your worst realistic path.
The third is pricing a model you will not use. The spread between models of similar ability is large and it moves weekly, which is what the cheapest price anyone charges for each level of ability, tracked daily exists to follow, and the same open model is served at quite different rates depending on the host. Price your actual candidate on its actual host, and while you are there note the levers that host offers: caching for the repeated opening of your prompt, a batch endpoint for work nobody is waiting on, and a spend cap per key as the backstop.
The last multiplication is the one this page deliberately leaves to you, because the two numbers it needs move week by week and a figure typed here would be wrong before you read it. They sit on the cards below, live: a rate for every million tokens read, a rate for every million written. Take your 126 million to the first and your 2.4 million to the second, add them together, then double the total for everything you have not thought of yet.
Then there is the structural choice, which changes the slope of the bill. For work that is narrow and repetitive, sorting and extracting and tagging, a small model usually clears the bar, and the gap against a frontier model is a different order of bill. Price both. The comparison takes twenty minutes and it is the only one on this page that can change the answer by a factor rather than a percentage.
Where next: What is a token · Compare model prices · Using a small model as a classifier