The checklist
-
Cost per task, from the bill
The number that matters is what one unit of work costs: one call scored, one question answered. Take it from the bill, because the bill is the only figure that includes everything: retries, failed requests, and work nobody remembers paying for. For example, about $25 a day for about 1,000 calls is about 2.5¢ a call.
How to check it: Open the provider's billing page. Take one month's total and divide it by the tasks completed in that month, counted from your own database. Write the number down. The rest of the audit explains it.
-
Price tables in your own code go stale
Many systems estimate their cost by multiplying token counts by prices typed into the code. Prices change, and nobody updates the numbers. A system we built priced every call at $0.20 and $0.60 per million tokens (input and output) when the real list prices were $0.32 and $3.20, so output was priced 5.3× too low.
How to check it: Find every price in the code and compare it with the provider's current price page. Then compare the code's monthly estimate with the invoice.
-
Your own usage logs can miss calls
Logged token totals don't always cover every request. Requests that error, retries, and side paths added later may never reach the log. In the same system we built, the logging gaps plus the stale prices meant the logs showed well under half the real cost.
How to check it: Count requests in your logs for one day and compare with the provider's usage page for the same day. Reconcile token totals against the invoice. Where they disagree, trust the invoice.
-
Work you pay for but don't use
You pay for every unit of work, including the ones that produce nothing. In call scoring, about half the audio transcribed belonged to calls that were never scored: voicemails, very short calls, and extra recording legs (the same call recorded more than once).
How to check it: Count what goes in against what comes out: recordings transcribed against calls scored, documents indexed against documents ever used. Filter out the junk before the paid step, not after.
-
Input vs output
Output tokens often cost much more than input tokens. In the system we audited, 10× more: $3.20 against $0.32 per million. In call scoring, the coaching the model writes is most of the scoring cost, about 1.5¢ of about 1.7¢ a call. How much you ask the model to write matters more than how long the call was.
How to check it: Pull average input and output tokens per task from the provider's usage page. Multiply each by its price. Then ask whether you need all the output you're paying for.
-
Fixed instructions
Every request carries its instructions (the prompt). If the instructions are long and the same every time, most of your input tokens are repeats. In the call-scoring system, about 90% of each call's input was the same instructions sent with every call. Prompt caching (cheaper rates some providers charge for repeated instructions they have already seen) could cut that. We haven't measured it, so we won't put a number on it.
How to check it: Compare the length of the instructions with the length of a typical task. If instructions dominate, check whether your provider offers prompt caching and what it charges.
-
Retries and silently dropped responses
A request that fails or comes back malformed may still be billed. If the code drops it quietly, you either pay twice (once for the failure, once for the retry) or pay once and get nothing. Nobody notices unless someone counts.
How to check it: Count requests sent against results stored, per day. The gap is retries, errors and dropped responses. Log every failure with its token count.
-
Model choice at list price
Same tokens, different model, very different bill. For one scored call, the same per-call tokens cost about 1.7¢ on Qwen 3.6 27B and about 6¢ on GPT-4o at list prices. That is a price comparison, not a quality test. We go through the numbers in our guide on open-source vs OpenAI call analysis cost.
How to check it: Take your per-task tokens and multiply by the list prices of two or three candidate models. Then run item 10 before you switch.
-
Self-host vs pay per call
A GPU server costs the same whether it is busy or idle. One system's own AWS GPU server cost about $800 a month and was too slow; it moved to pay-per-call at about $750 a month, got faster, and ran a bigger model. Self-hosting pays off in two cases: strict data rules (data can't leave your own cloud account), or enough volume to keep a strong GPU busy.
How to check it: Compare your monthly server cost, idle hours included, with what the same month's tokens would cost at pay-per-call prices. Note how busy the GPU actually is.
-
Quality checks so a cheaper setup isn't worse
Every item above can cut cost. None of them tells you whether the output got worse. A cheaper model, shorter output or a changed prompt can all change results.
How to check it: Keep a fixed set of real examples. Run the cheaper setup on them and have a person check the results before you switch. Repeat whenever you change the model or the prompt.
A worked cost-per-task example
Here is one scored call from a call-scoring system we built, priced at list.
| Part | Tokens | Price per million | Cost |
|---|---|---|---|
| Input | about 5,400 | $0.32 | about 0.17¢ |
| Output | about 4,700 | $3.20 | about 1.5¢ |
| Scoring total | about 1.7¢ | ||
| Transcription and retries | the rest | ||
| All-in from the bill | about 2.5¢ |
Three things to notice. Output is almost all of the scoring cost (item 5). Most of the input is fixed instructions (item 6). And the difference between about 1.7¢ at list and about 2.5¢ from the bill is transcription, retries, and audio that was transcribed but never scored (items 4 and 7). A cost estimate built from scoring tokens alone would have missed that gap.
How to run the audit in an afternoon
- Open three things: the provider's billing page, its usage page, and your own count of tasks completed for the same month. Where your own logs disagree with the provider's usage page, trust the provider (item 3).
- Work out cost per task from the bill (item 1).
- Price one task at list. Average input and output tokens from the usage page, multiplied by current prices (items 2, 5 and 6).
- Compare the two numbers. The gap is other paid steps such as transcription, unused work and retries (items 4 and 7).
- Price two alternative models at list (item 8), and compare self-hosting if you run a server (item 9).
- Write down what you would change, and the fixed set of examples you will test it on first (item 10).
Next step
Want us to check one AI workflow's cost per task?