Guide

Is it cheaper to run AI call analysis on open-source models than on OpenAI?

Yes, for us. Transcribing and scoring a sales call costs us about 2.5¢ all-in on open-source models run through OpenRouter, and at list prices the scoring step alone would cost about 6¢ a call on GPT-4o against about 1.7¢ on the open-source model we use (about 72% less). One caveat before the numbers: this compares price, not quality. We have never run GPT-4o on these calls, so we can't tell you whether it would score them better or worse.

One note on names. Whisper, our transcription model, is an open-source speech-to-text model released by OpenAI, and we run it through OpenRouter. So "open-source vs OpenAI" here means open models against OpenAI's paid API models such as GPT-4o.

The real numbers

Our call-scoring system scores about 1,000 or more sales calls a day for a US financial services company. Here is what we pay, from the real OpenRouter bill. (OpenRouter is a service that gives you one account and one bill for many AI models.)

What we pay (real bill)
Calls scoredabout 1,000+ a day
OpenRouter billabout $25 a day, about $750 a month
All-in cost per call (transcription plus scoring, retries included)about 2.5¢

Here is how one call breaks down at list prices. (Tokens are the units AI providers bill by. Input is what you send the model; output is what it writes back.)

One scored callTokensList priceCost
Scoring: inputabout 5,400 (about 90% fixed instructions)$0.32 per millionabout 0.17¢
Scoring: output (the coaching)about 4,700$3.20 per millionabout 1.5¢
Scoring totalabout 1.7¢
Transcription and retriesthe rest
All-in, from the billabout 2.5¢

Transcription costs more than you might guess. About half the audio we transcribe belongs to calls that never get scored (voicemails, very short calls, extra recording legs), and we still pay for it.

The GPT-4o comparison

GPT-4o's list prices are $2.50 per million input tokens and $10 per million output tokens. Put the same per-call tokens through those prices and scoring one call comes to about 6¢, against about 1.7¢ on Qwen.

Scoring one call, same tokensCost at list price
Qwen 3.6 27B via OpenRouter (what we run)about 1.7¢
GPT-4o (computed, not run)about 6¢
Differenceabout 72% less on open-source

Three limits: it covers scoring only (we didn't price transcription on OpenAI's API), it uses list prices rather than a negotiated rate or an invoice, and it says nothing about quality, because we never ran GPT-4o in production.

What we tried first: our own GPU server

From January to May 2026 the system ran on its own GPU server on AWS, running Whisper for transcription and a smaller open-source model (Qwen 2.5 14B) for scoring, both self-hosted. It cost about $800 a month.

It was too slow: the GPU that made sense to pay for wasn't strong enough. The same kind of models run fast on strong hardware (even a well-specced Mac), but strong cloud GPUs cost much more.

In May 2026 we moved to OpenRouter. The bill is about the same: about $750 a month against about $800. Analysis got faster, and we could use a bigger model (27B instead of 14B; the B is billions of parameters, a rough measure of model size). The win was speed and a bigger model, not a big saving.

Own GPU server (Jan–May 2026)OpenRouter (since May 2026)
ModelsWhisper and Qwen 2.5 14B, self-hostedQwen 3.6 27B for scoring, Whisper large-v3-turbo for transcription
SpeedToo slowFaster
Monthly AI costabout $800about $750

Where the cost actually comes from

Two facts decide the scoring bill.

Output costs 10× input. $3.20 against $0.32 per million tokens. Per call, that is about 1.5¢ for output and about 0.17¢ for input. Almost all the scoring cost is the model writing, not reading.

Most input is fixed instructions. Each call sends about 5,400 input tokens, and about 90% of that is the same scoring instructions sent with every call. The transcript is the small part. Each call gets back about 4,700 output tokens of detailed coaching.

So once you have picked a model, the biggest lever is how much coaching you ask for per call. The length of the call matters less than you would think. A score and one line of reasoning is cheap. Strengths, fixes with the exact words to use, and missed opportunities, which is what we ask for, cost more.

Prompt caching (some providers charge less for instructions they've already seen) could cut the input cost. We haven't measured it, and input is the small part of the bill anyway.

Check your own numbers: our logs showed well under half the real cost

Our system keeps its own cost log. It priced every call from numbers typed into the code: $0.20 per million input tokens and $0.60 per million output tokens.

The real list prices are $0.32 and $3.20, so output was priced 5.3× too low. On top of that, the logged token totals didn't cover every call. Put together, the logs showed well under half of what we were really paying.

Nobody was hiding anything. The prices were out of date, the logging had gaps, and nothing checked either against the bill. The lesson: work out cost per call from your provider's bill, not from your own logs.

When running your own server does make sense

It is the right choice in two cases:

  1. Strict data rules. If audio can't leave your own cloud account, a hosted API is off the table and a self-hosted model is the only option.
  2. Enough volume to keep a strong GPU busy all day. A strong GPU costs the same working or idle, so keeping it busy brings the cost per call down.

Outside those two cases, pay per call.

How to work out your own cost per call

  1. Count scored calls, not recordings. A big share of recordings never gets scored.
  2. Start from the bill. Take the month's total from your provider's billing page, not your own logs, and divide by scored calls. That is your all-in cost per call.
  3. Then price one call at list. Take input and output tokens per call from your provider's usage page and multiply by current list prices. This shows what drives the cost.
  4. Compare the two. If the bill is well above tokens × price, the gap is transcription, retries and audio you transcribe but never score. And re-check when prices change.

How we measured, and the caveats

  • The all-in figure (about 2.5¢) is our real OpenRouter bill divided by calls: about $25 a day for about 1,000 calls a day.
  • The scoring figure (about 1.7¢) is list prices multiplied by typical per-call tokens. There is no invoice-level split by model, so this is the closest we can get.
  • One system, one company, as of the time of writing (September 2026). Your calls, coaching length and prices will differ.

Next step

Want to know what AI call scoring would cost for your team, and whether it grades calls the way your QA team does? That's what our Blind Match does: we score up to 100 calls your QA team already graded, in 2 weeks, for $750.