Case study · AI assistant

An AI tutor that stopped guessing

An education company in a regulated industry had an AI tutor that made things up. We rebuilt it. On a fixed set of 150 test messages, wrong "I don't know" replies went from 6 to 0, and a person reading every reply found no made-up answers or numbers. The whole AI bill is under $5 a month.

6 → 0wrong "I don't know" replies
0made-up answers or numbers in 150 test replies
Under $5total AI bill per month

The problem

The company had an AI tutor for its learners. It had two faults. It sometimes made up answers and numbers. And it sometimes said "I don't know" about things that were in the course material. Adding more rules to its instructions did not fix either one. In a regulated industry, a confident wrong answer is worse than no answer.

What we built

We rebuilt it around one rule: the AI only explains, and only from the company's own course content.

  • Simple software, not the AI, asks the practice questions and checks the answers against the real answer key. The AI never decides right or wrong.
  • For each question, the software finds the right passages in the course material. The AI writes its explanation from those passages only.
  • When something is not covered, the software says so straight away, without calling the AI. That costs nothing.
  • Before every release we run the same 150 test messages, and a person reads every reply.

The numbers

What we measuredResult
Wrong "I don't know" replies, on the 35 test messages about terms in the course material6 before, 0 after
Made-up answers or numbers, across all 150 test messages0
Replies to learners, 29 Jul to 24 Sep 2026716 (684 by the AI, 32 by simple software at no cost)
AI fees for those 684 AI replies$2.52, about $0.004 each
Total AI billUnder $5 a month

How we measured

  • We used the same fixed set of 150 test messages, the same AI model and the same setup, before and after the rebuild. After later changes we ran the full set again, and it stayed at 0.
  • Our team wrote the 150 messages to read the way learners really type. They are not real learner messages.
  • The 6 → 0 counts the 35 messages that ask about a term covered in the course material. The other 115 test other things.
  • The $2.52 is the per-reply cost our AI provider reported for successful replies. It leaves out failed attempts, the search step, one-time setup and testing. The model was upgraded once during this period.

What we did not measure

  • Whether learners learn faster or do better. We measured the tutor's replies, not learner results.
  • Real learner messages. Our test set imitates them.
  • One smaller issue remains in 2 of the 150 replies: the tutor says where a topic is covered, then says it can't confirm it. We are fixing it.

Want the same for your documents?

Every AI Assistant starts with Hundred Questions: a working version on a sample of your documents, 100 real questions from your team, each answer checked by a person. $750, one week.