Case study · AI Sales Coach

Every sales call scored and coached, for about 2.5¢ each

Result: Our AI Sales Coach scores about 1,000 or more recorded sales calls a day for a US financial services company and sends a daily coaching email to about 60 reps. The AI running cost, from the real bill, is about $25 a day, or about 2.5¢ a call all-in.

1,000+calls scored a day
About 60reps coached every weekday
About 2.5¢AI cost per call, from the real bill

The problem

The company's phone sales team makes and takes more than a thousand recorded calls a day. Recording them is easy. Listening to them is not. This is the situation most phone sales teams are in, and it is the one we set out to fix here.

  • Managers can't listen to a thousand calls a day. They hear a handful, usually the ones a rep asks about or a customer complains about.
  • QA can only sample. A small sample is normal, but it leaves most calls unheard.
  • Coaching comes late. By the time a reviewed call reaches the rep, the habit behind it has had weeks to set.
  • Compliance risk goes unseen. A do-not-call request that isn't honored, or a promise a rep shouldn't have made, sits in a recording nobody plays.

What we built

A system that scores every call and turns the scores into coaching, without anyone pressing a button.

  • It pulls and transcribes every call. It collects every recorded call from the phone system automatically and transcribes it.
  • It sets aside non-conversations. Voicemails, wrong numbers, hang-ups under a minute and phone menus are filtered out, so reps aren't marked down for calls that weren't really calls.
  • It scores each real conversation. Every conversation gets a 1–10 score on eight measures: greeting, discovery, product knowledge, rapport, objection handling, closing, persistence (outbound calls only) and script adherence. There is an overall score too. Every score comes with a one-sentence reason, so a rep can see why it was a 6 and not a 9.
  • It writes coaching for each call. Strengths, what to fix with the exact words to use, and missed opportunities.
  • It flags compliance issues separately. Do-not-call requests, pressure tactics, missing disclosures and promises a rep can't make are flagged with a severity: low, medium, high or critical. Flags appear on a dashboard. About two-thirds of calls are analyzed the same day (the median is about 1 hour 45 minutes after the call); the rest are done by the next morning. There is no instant alert to managers. It is a dashboard, not a pager.
  • It emails each rep every weekday morning. Every weekday at 9:00 am ET, each rep with three or more calls on the previous business day gets an email: a win, their scores against the team, and one thing to work on.
  • It reports to managers weekly. Managers get a follow-up report on Monday mornings.
  • It feeds the company's own systems. Call history flows through an API into the company's existing tools.

The numbers

Volume and reach
Calls scoredabout 1,000+ a day
Reps receiving the daily coaching emailabout 60 on a typical weekday
Calls analyzed the same dayabout two-thirds (median about 1 h 45 min after the call)
Remaining calls analyzed bythe next morning
AI running cost
OpenRouter billabout $25 a day, about $750 a month
Per call, all-in (transcription plus scoring)about 2.5¢
Of which scoring, at list pricesabout 1.7¢ (about 5,400 input and 4,700 output tokens per call)

These are the AI model costs (transcription and scoring) only. They do not include our fees.

How we measured

  • The all-in figure is the real OpenRouter bill divided by calls: about $25 a day for about 1,000 calls a day. Both models run through OpenRouter, a service that gives one account and one bill for many AI models.
  • The scoring share is list prices multiplied by typical tokens per call (tokens are the units AI providers bill by). It is a list-price calculation, not an invoice line, because there is no invoice-level split by model.
  • Volumes and timings come from the system's own database.
  • Figures are as of the time of writing, September 2026.

What we did not measure

We would rather say this plainly than let you assume it.

  • Whether our scores match the company's own QA team. No side-by-side test has been run on the same calls. That is exactly what our Blind Match does for a new client.
  • Any effect on sales, close rates or complaints. Not measured, so we don't claim it.
  • QA hours saved. Not measured.
  • How many flags turned out to be real problems. Not measured.

Find out what this would cost for your team

Want to know what AI call scoring would cost for your team, and whether it grades calls the way your QA team does? That's what our Blind Match does: we score up to 100 calls your QA team already graded, in 2 weeks, for $750.