
How to Test AI Search on Your Technical Manuals
Choose real manual questions, then judge the answer and its evidence before you expand an AI search test.
A polished demo can hide a weak lookup. Test permitted documents, approved revisions, known answers, and questions the permitted set cannot answer. Keep the test narrow enough to judge.
Start with one lookup task
Pick one product family and one repeatable question type, such as a compatible part number or a table specification. Keep the test small enough for one product specialist to review.
Write one sentence that defines the task:
Given an approved set of manuals for one product family, find the answer to a support question and show the exact page and region that supports it.
The task does not ask AI to approve a repair, diagnose a hazard, or replace the accountable person.
Freeze the documents and expected answers
Create the test set before anyone runs either method. NIST's AI Risk Management Framework Playbook says to document test sets, metrics, and tool details. It says that documentation enables repeatability and consistency.
For each file, record:
- file name and document identifier
- product model or family
- revision, issue date, and approval status
- whether the file is allowed in the test
- which file wins if two revisions conflict
Have a product specialist write the expected answer and mark the page, cell, label, or note that proves it. Do this before seeing either method's answer.
If source authority is unclear, stop and resolve it. Search cannot fix an approval problem.
Record the current lookup method
Use the same questions, documents, and qualified reviewer for the current method and the AI test.
For each method, record three times:
- Time to find a candidate answer.
- Time for the reviewer to verify the answer against the source.
- Total lookup and verification time.
The reviewer-time score below uses verification time. Read total time alongside it. A fast lookup that takes longer to check may only move the work.
Do not run the baseline first for every question. Before testing, match questions within each type by expected difficulty. Run the baseline first for one question in each pair and the AI method first for the other. Assign the order before anyone sees results. This reduces order bias, but it does not remove learning effects. If the set is too small, record the run order and label it as a limitation.
Run both methods under similar conditions. Give neither method hints that normal users will not have.
Test five kinds of question
Split the set so you can see where each method works and where it breaks.

| Question type | What it tests | Example |
|---|---|---|
| Prose | Direct retrieval from paragraphs | What condition does the manual state for this warranty? |
| Diagram | Labels and visual relationships | Which callout identifies the named component? |
| Table | Row, column, unit, and footnote relationships | Which listed item applies to this model and condition? |
| Revision | Source authority across similar files | Which value appears in the approved revision? |
| Unanswerable | Restraint when evidence is absent | Does the permitted set state this tolerance? |
Use questions from real support work. Keep your team's wording and model numbers, but remove unneeded customer details. Unanswerable questions check whether the tool identifies missing evidence instead of guessing.
Use this scorecard
Copy this table into a spreadsheet. Preserve every raw answer, source reference, and timing record.
| Field | What to record |
|---|---|
| Test ID | Stable ID for the question |
| Question type | Prose, diagram, table, revision, or unanswerable |
| Question | Exact wording used in both methods |
| Expected result | Approved answer, or "not answerable from permitted files" |
| Allowed source | File ID, model, approved revision, and exact evidence page |
| Run order | Which method ran first for this question |
| Baseline result and evidence | Full answer plus file, page, and source region used |
| Baseline times | Lookup time, verification time, and total time |
| Baseline scores | Correctness, source support, revision, and unanswerable behavior when applicable |
| AI result and evidence | Full answer plus file, page, and cited region or excerpt |
| AI times | Lookup time, verification time, and total time |
| AI scores | Correctness, source support, revision, and unanswerable behavior when applicable |
| Reviewer-time score | 0, 1, or 2 using the predeclared verification-time threshold |
| Notes | Failure type, ambiguity, escalation, correction, or order caveat |
Score each method separately for correctness, source support, revision, and unanswerable behavior. Use the reviewer-time row to compare AI verification time with the baseline:
| Measure | 2 | 1 | 0 |
|---|---|---|---|
| Correctness | Matches the approved answer with every required model, part, value, unit, and condition | Useful but incomplete, with no material wrong detail | Material error, unsupported addition, or wrong conclusion |
| Source support | Evidence directly supports the full answer and the cited location is usable | Right document, but the evidence is incomplete or the location is too broad | Missing evidence, wrong source, or evidence contradicts the answer |
| Revision | Uses the approved revision and identifies it clearly | Revision appears right but is not identified clearly | Uses or may use a stale, wrong, or unapproved revision |
| Reviewer time | Verification time is at least the predeclared threshold faster than the baseline | Verification time is faster than the baseline, but by less than the threshold | Verification time is equal to or slower than the baseline, or cannot be verified |
| Unanswerable behavior | Says the permitted set does not support an answer and identifies the gap | Declines to answer but gives a vague or partly wrong reason | Invents an answer or misses evidence that is present |
Do not hide a critical miss inside an average. Fail an answerable case if correctness, source support, or revision scores 0. Fail an unanswerable case if its unanswerable score is 0. Then compare pass rates by question type, reviewer verification time, and total time with the current method.
Run the test without creating an order advantage
- Lock the files, revisions, questions, answer key, rubric, matched pairs, and run order.
- Run each question with its assigned method first, then run the other method.
- Start a clean AI session for each question unless normal use depends on prior context.
- Save the exact answer, source reference, visible evidence, and all three times.
- Have the same reviewer score both methods without changing the rubric.
- Label every miss by cause, such as retrieval, interpretation, source location, revision, or refusal.
- Record order and learning effects as a limitation. Before repeating, record what changed.
If you rewrite a question after seeing results, save it as a new test version.
Worked example: check a parts-list relationship
Graco publishes manual 310662D for UltraMix and HydraMix displacement pumps. This parts-list example uses the manual as its reference answer. It does not report a test of Keyline or another AI tool.
Question: For UltraMix pump part number 248540, which numbered parts in Repair Kit 248438 are also included in Packing Kit 248437?
Expected answer: In Graco manual 310662D, Revision D, June 2019, reference numbers 3 and 15.
Exact evidence: Printed page 12 identifies pump 248540. An asterisk marks Repair Kit 248438 parts, and a dagger marks Packing Kit 248437 parts. Rows 3 and 15 carry both. Row 18 carries only the repair-kit mark. Printed page 16 identifies MM 310662 and states "Revision D, June 2019."
Printed page 5 says that an asterisk marks repair-kit parts and points readers to pages 12 and 13. It does not identify Packing Kit 248437 or show which rows carry both markers.
Under this proposed rubric, a hypothetical answer that gives references 3 and 15, identifies manual 310662D and Revision D, and cites printed pages 12 and 16 earns 2 for correctness, source support, and revision. This is a scoring illustration. It is not an observed AI result.
For an unanswerable variant, permit only printed pages 5 and 16 of manual 310662D. Exclude every other page. Those pages identify the repair-kit marker and revision, but not the packing-kit marker or rows with both markers. The expected result is "not answerable from the permitted pages." This narrow case does not claim that the full manual becomes unanswerable when page 12 is removed.
Set the buying decision before you see the result
Before testing, name the minimum pass rate by question type, maximum critical failures, reviewer-time improvement, decision date, and owner. Set a separate rule for unanswerable cases.
The useful outcome is a decision grounded in the same questions your team handles every week.
If your team has one recurring manual lookup to test, read about Keyline for technical teams or contact us to discuss a paid Keyline beta. Tell us the workflow, the team that owns it, and the document type. Do not send files.