Most AI projects are sold on a demo. Someone from the vendor types three questions, the system answers them well, and the room nods. Six weeks after go-live, your staff are quietly checking every answer by hand, which means the system saves nobody any time. The fix is simple, and it belongs in the contract rather than the pitch deck: a measured accuracy threshold, tested against questions your own people wrote, passed before go-live and tied to a payment.
Why the demo tells you almost nothing
A demo is a performance. The questions were chosen by the vendor, tried beforehand and probably tuned until the answers looked good. That is not dishonest. It is just not evidence.
Three answers show that the system can be right. They say nothing about how often it is right across the hundreds of different questions your team will actually ask. They also skip the most dangerous case: the question whose answer is not in your documents at all. A system that confidently invents an answer to that question will do more damage than one that is occasionally wrong about something it can look up.
Demo documents are usually clean. Yours are not. You have scanned PDFs, three versions of the same policy, rate cards in spreadsheets and a procedure manual last updated by someone who left. The only way to know how a system handles your material is to test it on your material.
Build the test set before you sign anything
The test set is a list of questions, each with a known-correct answer and the document it comes from. It is the single most useful thing you can bring to an AI project, and it should come from your side, not the vendor’s.
- Use real questions. Ask the people who answer questions today to write down what they get asked over a week or two. Keep their phrasing, shorthand and spelling. Nobody types a neat, complete sentence into a search box at 4pm on a Friday.
- Write the correct answer next to each one. The answer, and where it lives: document, page or section. Someone who knows the subject should write it, not someone guessing.
- Include questions the system should refuse. Ask about a product you do not sell, a policy you do not have, a client who is not in the files. The right answer is “I can’t find that in the documents.” Anything else is invention.
- Include the hard ones. Questions where the answer changed between versions of a document. Questions that need two documents combined. Questions whose answer is a number, a date or an amount.
- Hold some back. Give the vendor part of the set to develop against. Keep the rest sealed until acceptance testing. If the vendor has seen every question, you are testing their memory of your test, not the system.
As a rough guide, the set should be big enough to cover every type of question you expect, with several examples of each. For a single-department assistant that is often a few hundred questions, not ten.
What to measure
Measure three things, and report them separately.
- Answer correctness. Is the answer right and complete enough to act on? Have your staff grade each one on a simple scale: correct, partly correct, wrong. Agree in advance whether “partly correct” counts as a pass.
- Citation correctness. Does the cited source actually say what the answer says? An answer can be right with the wrong citation, which is luck. Or it can cite the right document and misstate it. Both matter, because your staff will trust the citation and stop checking.
- Refusal. On the questions with no answer in the source, how often does the system say so instead of inventing something? Check the reverse too: a system that refuses questions it could have answered is not much use either.
Do not let these be blended into a single score. A blended number hides exactly the failure you care most about. A system can score well overall while inventing answers to half the questions it should have refused.
What number is reasonable, and how to write it down
There is no universal right threshold. It depends on what the system is used for and what a wrong answer costs.
An internal assistant that helps staff find the right section of the leave policy, where a person reads the source before acting, can live with a lower bar. A wrong answer costs a few minutes. An assistant that quotes contract terms to customers, or tells a driver which rules apply to a load, needs a much higher bar and probably a person reviewing answers before they go out.
The refusal threshold is often the one to set highest. An honest “I don’t know” is cheap. A confident, wrong answer with a real-looking citation is expensive.
Look at the failures, not just the percentage. Say the threshold is 90% correct on a 200-question held-out set. That allows up to 20 wrong answers. If three of those 20 are about safety procedures, the headline number is not the whole story. It is reasonable to name categories where no wrong answers are acceptable.
Then put it in the contract:
- the thresholds for each of the three measures;
- the test set, or how it will be built and who holds the sealed portion;
- who grades the answers and how disagreements are settled;
- a payment milestone released only when the thresholds are met;
- what happens on a fail: the vendor fixes and re-tests at its own cost, with a limit on attempts after which you can walk away.
A vendor who is confident in their work will not object to this. A vendor who resists measuring against your questions is telling you something useful.
Run in parallel, then keep measuring
Passing the test set is the entry ticket, not the finish line. Before you switch over, run the system alongside your current process for a few weeks. Staff answer questions the way they do now, the system answers the same questions, and someone compares. Parallel running finds the questions nobody thought to put in the test set. Add them to it.
After launch, accuracy decays unless someone watches it. Three things cause most of the drift:
- Your content changes. A new price list, an updated policy, a retired procedure. Re-run the test set whenever source documents change materially, and on a fixed schedule even if they don’t.
- The underlying model changes. Vendors update the AI models their systems are built on. The contract should require a re-test against your set before a model change goes live.
- Usage changes. People start asking things nobody planned for. Log questions, answers and citations. Every month, have someone grade a sample. Watch the refusal rate: a sudden rise or fall usually means something has changed underneath.
Someone on your side should own the test set. It is a business asset, and it outlives any one vendor.
A checklist for the vendor meeting
- Will you test the system against questions our staff write, including ones it should refuse?
- Will you report answer correctness, citation correctness and refusal separately?
- Can we hold back part of the test set until acceptance?
- What thresholds do you propose for our use, and why?
- Will you tie a payment milestone to passing those thresholds?
- What happens, and who pays, if the system fails acceptance?
- Can we run it in parallel with our current process before switching over?
- How will we know when accuracy drops after launch?
- Will you re-test against our set before changing the underlying model?
- Who owns the test set and the logs if we part ways?
We build knowledge assistants with citation verification and an evaluation harness from the start; if you are weighing up an AI project and want an independent view first, begin with an AI Operations Audit.