The problem
The company had an AI tutor for its learners. It had two faults. It sometimes made up answers and numbers. And it sometimes said "I don't know" about things that were in the course material. Adding more rules to its instructions did not fix either one. In a regulated industry, a confident wrong answer is worse than no answer.
What we built
We rebuilt it around one rule: the AI only explains, and only from the company's own course content.
- Simple software, not the AI, asks the practice questions and checks the answers against the real answer key. The AI never decides right or wrong.
- For each question, the software finds the right passages in the course material. The AI writes its explanation from those passages only.
- When something is not covered, the software says so straight away, without calling the AI. That costs nothing.
- Before every release we run the same 150 test messages, and a person reads every reply.
The numbers
| What we measured | Result |
|---|---|
| Wrong "I don't know" replies, on the 35 test messages about terms in the course material | 6 before, 0 after |
| Made-up answers or numbers, across all 150 test messages | 0 |
| Replies to learners, 29 Jul to 24 Sep 2026 | 716 (684 by the AI, 32 by simple software at no cost) |
| AI fees for those 684 AI replies | $2.52, about $0.004 each |
| Total AI bill | Under $5 a month |
How we measured
- We used the same fixed set of 150 test messages, the same AI model and the same setup, before and after the rebuild. After later changes we ran the full set again, and it stayed at 0.
- Our team wrote the 150 messages to read the way learners really type. They are not real learner messages.
- The 6 → 0 counts the 35 messages that ask about a term covered in the course material. The other 115 test other things.
- The $2.52 is the per-reply cost our AI provider reported for successful replies. It leaves out failed attempts, the search step, one-time setup and testing. The model was upgraded once during this period.
What we did not measure
- Whether learners learn faster or do better. We measured the tutor's replies, not learner results.
- Real learner messages. Our test set imitates them.
- One smaller issue remains in 2 of the 150 replies: the tutor says where a topic is covered, then says it can't confirm it. We are fixing it.
Want the same for your documents?
Every AI Assistant starts with Hundred Questions: a working version on a sample of your documents, 100 real questions from your team, each answer checked by a person. $750, one week.