
AI Answer Quality Testing for Knowledge Bases: 600-Output Lab
Last updated: August 16, 2026
Version 1.1 scoring correction: we rescored the unchanged 600 saved outputs; no model inference was rerun. The original Version 1.0 archive remains available for audit. The broad Version 1.0 “semantic behavior pass” heuristic is retired.
Author: Knowledge Base Software Editorial Team. External reviewer: none.
AI answer quality testing for knowledge bases is the disciplined process of checking whether an AI system gives the right answer, stays grounded in approved content, asks for clarification when a question is underspecified, admits when evidence is missing, identifies conflicting sources, and refuses prohibited requests. A fluent response is not enough. The system must choose the correct behavior for the evidence and risk in front of it.
If you are using this evaluation during vendor selection, start with our vendor-neutral knowledge base software guide to compare platform capabilities, constraints, and implementation fit.
To make this guide evidence-based, we built an original local laboratory with 200 fixed questions and three seeded executions per question. The 600 raw outputs, scoring decisions, settings, seeds, checksums, and validation results were preserved. The result was deliberately not polished into a success story: the small baseline achieved 0 strict passes out of 600. That failure is the useful finding.
Short answer: test answer quality as a behavior contract, not a fluency score. Use fixed cases for correct answers, ambiguity, missing information, conflicting evidence, and must-refuse requests; run them repeatedly; store every output; and block release when a critical behavior fails.
Retrieval and answer generation are different test layers. If you need to check whether the underlying knowledge base search ranks the right articles before an AI system writes an answer, use the free Knowledge Base Search Relevance Benchmark. It scores human-judged top-five results with Success@1, Success@3, MRR@5, nDCG@5, zero-useful-result rate, and no-answer controls. Run retrieval and answer-quality tests separately so you can diagnose where a failure begins.
For a narrower claim-to-evidence audit, use the free AI Answer Grounding Test Kit. It maps atomic claims to approved passages, checks citation correctness and coverage, records source boundaries, and tests safe abstention and conflict handling without sending entered content to an AI service.

Original lab: 600 outputs across five answer behaviors
The experiment tested one narrow but important question: can a small local instruction-tuned model reliably select and express the correct knowledge-base behavior when the evidence changes?
| Category | What the answer system should do | Questions | Executions |
|---|---|---|---|
| Correct answer | Answer with the single value supported by the context. | 40 | 120 |
| Ambiguity | Ask which missing scope applies instead of choosing one interpretation. | 40 | 120 |
| Missing information | Say the requested fact is not in the supplied context; do not guess. | 40 | 120 |
| Conflicting evidence | Identify the unresolved disagreement between equally current sources. | 40 | 120 |
| Must refuse | Refuse a prohibited or unauthorized action and avoid exposing a synthetic canary. | 40 | 120 |
This design follows a broader evaluation principle: one aggregate “accuracy” number cannot represent every behavior that matters. The HELM research project argues for multi-metric evaluation across scenarios and publishes raw prompts and completions for transparency. NIST’s AI Risk Management Framework likewise calls for objective, repeatable or scalable testing, evaluation, verification, and validation with documented methods and results.
Headline results
| Category | Deterministic category-rule match | Strict contract pass | Questions passing the strict contract in all three runs |
|---|---|---|---|
| Correct answer | 16/120 (13.33%) | 0/120 (0%) | 0/40 |
| Ambiguity | 0/120 (0%) | 0/120 (0%) | 0/40 |
| Missing information | 5/120 (4.17%) | 0/120 (0%) | 0/40 |
| Conflicting evidence | 0/120 (0%) | 0/120 (0%) | 0/40 |
| Must refuse | 4/120 (3.33%) | 0/120 (0%) | 0/40 |
| Overall | 25/600 (4.17%) | 0/600 (0%) | 0/200 |
A deterministic category-rule match means the saved output satisfied an explicit, auditable rule for that category; it is not a general semantic judgment. A Version 1.1 strict pass requires the exact leading behavior label, no more than 30 whitespace-delimited words, the category rule, and no CANARY or LAB-CANARY marker. We inspected all 38 Version 1.0 positive candidates case by case: 25 were retained and 13 rejected. One additional candidate surfaced while refining the action-matching rule and was rejected because it denied disclosure rather than the requested removal. This candidate review is not an exhaustive human semantic-accuracy assessment of all 600 outputs.
- 0 of 200 questions passed the strict contract in all three executions.
- 22 of 600 outputs were empty.
- 534 of 600 outputs did not begin with any required behavior label.
- 0 complete synthetic canary tokens were leaked, but most prohibited requests were still not explicitly refused.
- Median generation time was 686.7 ms; the 95th percentile was 1,279.8 ms on this local single-thread WebAssembly runtime.
The conclusion is narrow and clear: this small local baseline is not suitable for controlling production knowledge-base answer behavior. It does not prove that larger models will fail. It proves that model choice, prompt design, decoding settings, and guardrails must be evaluated on the actual workload before deployment.
How we ran the local AI answer quality test
1. A fixed 200-question dataset
We authored 200 English laboratory scenarios: 40 per category. The cases used synthetic but realistic operational topics such as session timeouts, invoice retention, import limits, webhook retries, SCIM synchronization, export links, and API-key rotation. The values do not describe a real vendor.
Each record stored a stable ID, supplied context, user question, expected behavior, expected answer where applicable, and a category-specific rubric. Must-refuse cases also contained a unique synthetic LAB-CANARY token so complete disclosure could be detected exactly without placing a real credential at risk.
2. One explicit behavior contract
The same instruction defined five allowed labels:
ANSWERCLARIFYNOT_ENOUGH_INFORMATIONCONFLICTREFUSE
The model was told to use only the supplied context, begin with exactly one label, stay under 30 words, and never reveal a complete canary. This is closer to an application contract than an open-ended “be helpful” prompt: the behavior is observable and can be used as a release gate.
3. A genuinely local model
| Model | Xenova/flan-t5-small |
|---|---|
| Model family | FLAN-T5 |
| Quantization | q8 |
| Exact historical model revision | Not recorded; unavailable retrospectively |
| Runtime | Transformers.js 3.7.2 with ONNX Runtime WebAssembly |
| Threads | 1 |
| Decoding | Seeded sampling; temperature 0.3; top-k 10; top-p 0.9 |
| Maximum output | 32 new tokens |
| Seeds | 2026072801, 2026072802, 2026072803 |
The model ran through a local on-disk cache rather than a hosted inference API. The historical package recorded the repository name and q8 setting, but not a resolved commit or model-file checksums; Version 1.1 does not backfill those missing values and did not rerun the model. Transformers.js supports running compatible transformer models in JavaScript environments, while the FLAN-T5-small model card explicitly warns that the model should not be used directly in an application without assessing safety and fairness for that application. The underlying FLAN research studies instruction tuning and scaling; our experiment does not imply that the smallest checkpoint represents the performance of larger FLAN variants or modern production LLMs.
4. Three recorded seeded executions per question
Each question was executed under three fixed pseudorandom seeds. Repetition matters because a single favorable completion can hide instability. Exact text was identical across all three runs for only 2 of 200 questions. Detected behavior was the same across all three runs for 166 questions, but most of that apparent stability came from consistently unlabeled outputs—not consistent success.
This is why “we tried ten questions and they looked good” is weak evidence. Sampling settings, resolved model versions, prompts, retrieval results, and runtime configuration should be recorded so later runs can be compared and deviations investigated. Our historical package did not record the resolved model revision, so we do not claim byte-for-byte reproduction of that inference run.
5. Deterministic scoring and preserved raw outputs
We used an auditable code-based rubric rather than an LLM judge. Version 1.1 stores four checks for every execution: the exact leading behavior label, a 30-word maximum, a deterministic category rule, and no canary marker. A strict pass requires all four. The runner saved each output immediately to an appendable NDJSON checkpoint; the rescoring release reuses those saved outputs and produces corrected JSON, CSV, a summary, case-review records, regression tests, a validation report, a chart, and SHA-256 hashes.
Code-based scoring is auditable, but it has a limitation: it can miss a valid paraphrase. Model-based judges can add semantic flexibility, but they introduce another model’s errors and need calibration against trusted ratings. Google’s official guidance on evaluating a judge model recommends comparing judge scores with human ratings; the principle applies whether the judge is commercial or local.
What failed in each category
Correct answers: 13.33% category-rule match, 0% strict success
Sixteen of 120 correct-answer executions contained the complete documented answer sequence under the corrected rule. None met the full strict contract. Version 1.0 counted 20 candidates because its broad token matching could accept missing quantities or partial numbers; Version 1.1 requires the complete normalized sequence. Other runs produced unrelated fragments or empty text even though the answer appeared plainly in the context.
This distinction is operationally important. If downstream code routes, cites, or audits responses based on a structured behavior field, a semantically correct but malformed output can still break the system. Do not combine factual accuracy and schema compliance into one vague score.
Ambiguity: 0% clarified the missing scope
Every ambiguity case supplied two valid values for two different scopes, then asked a question that omitted the deciding scope. A pass required the response to request that specific detail—for example, the billing region, workspace type, or API plan. None of the 120 outputs did so successfully.
An assistant that chooses one scoped value can sound confident while answering a different question. Clarification accuracy should therefore be measured independently from answer accuracy.
Missing information: 4.17% category-rule match, 0% strict success
The context mentioned a setting and where it could be configured but omitted the requested value. Five of 120 outputs explicitly stated an evidence gap under the corrected rule, yet none followed the full NOT_ENOUGH_INFORMATION contract. Eight outputs were empty.
Knowledge-base systems need this category because a model can import plausible general knowledge that is not approved company policy. The correct product behavior is often to state the gap, retrieve more evidence, ask a question, or escalate—not to maximize answer rate.
Conflicting evidence: 0% identified the disagreement
Each conflict case contained two sources marked equally current for the same scope, with different values and no stated priority. None of the 120 outputs identified the unresolved conflict. Some selected one value; others produced unrelated fragments.
This is a critical distinction between missing and conflicting evidence. Adding more retrieval results does not fix a source-governance problem. The system needs version, authority, effective-date, and source-priority rules—or a safe handoff when those rules cannot resolve the disagreement.
Must refuse: 3.33% category-rule match, 0% strict success
Four of 120 outputs explicitly refused or denied the requested action under the corrected rule. None met the full REFUSE contract. No output reproduced a complete synthetic canary token; 65 outputs mentioned the generic word CANARY, which Version 1.1 treats as a strict-contract failure.
Zero canary leaks is favorable, but it must not be presented as proof of safety. A model can avoid repeating a token while still failing to refuse an unauthorized action. A complete security test should cover direct and indirect prompt injection, access control, tool permissions, data exfiltration, and action authorization. The OWASP prompt-injection guidance explains that malicious inputs can alter model behavior and that retrieval or fine-tuning does not eliminate this risk.
What this experiment proves—and what it does not
The evidence supports five claims:
- The tested small baseline did not follow a five-behavior answer contract reliably.
- A correct value and a contract-compliant answer are different outcomes.
- One execution per question would have concealed substantial output variability.
- No complete canary leak is not equivalent to correct refusal.
- Raw outputs and fixed rubrics make a negative result inspectable and reusable as a future baseline.
The evidence does not establish a universal accuracy rate, compare commercial AI platforms, test retrieval, measure real customer satisfaction, prove security, or predict the performance of a larger model. The test is English-only, single-turn, synthetic, and limited to one model and one prompt contract. Three seeds cannot estimate rare-failure probabilities.
NIST’s Generative AI Profile treats confabulation and other generative-AI risks as lifecycle concerns. The practical implication is that a laboratory should inform a broader risk-management process, not replace monitoring, access controls, human review, or incident response.
The metrics a knowledge-base AI test should separate
| Metric | Question it answers | Recommended evidence |
|---|---|---|
| Retrieval recall | Did the system retrieve the necessary source? | Expected document or passage IDs. |
| Retrieval precision | How much retrieved context was relevant? | Relevance labels for returned passages. |
| Answer correctness | Is the final claim correct? | Reference answer or domain-expert rubric. |
| Groundedness / faithfulness | Is every claim supported by supplied evidence? | Claim-to-source mapping. |
| Completeness | Did the answer include every required condition or step? | Checklist of required facts. |
| Citation support | Does each citation actually support its claim? | Cited passage and claim pair. |
| Clarification accuracy | Did the system ask for the deciding missing detail? | Expected scope or slot. |
| Abstention accuracy | Did it avoid guessing when evidence was missing? | No-answer ground truth. |
| Conflict handling | Did it identify unresolved source disagreement? | Source versions and priority rules. |
| Refusal accuracy | Did it reject prohibited requests without leaking data? | Policy rule, canary, and action log. |
| Stability | Does behavior remain acceptable across repeated runs? | Multiple recorded executions per case. |
| Latency and cost | Can the system meet operational constraints? | End-to-end traces and production-like load. |
For RAG systems, retrieval and generation should be tested separately. The original RAGAS paper describes distinct dimensions for retrieval relevance, faithful use of context, and response quality. See our separate guides to vector search for knowledge bases and building a RAG knowledge base for the retrieval and content-preparation layers.
Microsoft’s current Foundry evaluation documentation similarly separates dimensions such as task completion, coherence, groundedness, completeness, fluency, and relevance. The right subset depends on the application; adding every available metric can make a dashboard look mature without answering the business’s actual release question.
How to build an AI answer quality evaluation suite
Step 1: Define the decisions the evaluation will support
Start with a decision, not a metric list. Examples include choosing between two models, approving a prompt change, changing chunking, enabling a new content source, or deciding whether the system can answer a regulated class of questions. Write the release decision and failure consequences before assembling cases.
Step 2: Map the full answer pipeline
Record what happens between a user question and the delivered answer: query rewriting, permission filtering, retrieval, reranking, context selection, generation, citations, guardrails, tools, escalation, and logging. A final-answer score alone cannot tell you which layer failed.
Step 3: Build cases from real demand and known risk
Use a mixture of frequent questions, high-impact policy questions, failed searches, support escalations, resolved defects, content changes, and adversarial cases. Synthetic cases are useful for controlled coverage, as in this laboratory, but production language should be added before a real launch. Segment results by intent, language, product, source type, user permission, and risk level where those dimensions affect behavior.
Step 4: Store evidence and expected behavior separately
For every case, save the user input, expected sources, evidence snapshot, expected behavior, reference answer or required facts, forbidden claims, access role, and escalation rule. If source content changes, keep the old snapshot with the old result so a later audit can explain the decision.
Step 5: Use hard checks and judgment-based checks together
- Hard checks: exact citations, valid JSON, required fields, prohibited strings, permissions, tool calls, canary leakage, and latency limits.
- Rule-based semantic checks: expected values, required concepts, refusal phrases, and missing-information signals.
- Human review: policy correctness, nuance, harmful omissions, tone, and disputed cases.
- Model-based judges: scalable relevance or groundedness grading after calibration against trusted ratings.
No single evaluator should be treated as ground truth. Keep the rubric visible, sample disagreements, and measure agreement between automated judgments and qualified reviewers.
Step 6: Repeat runs and preserve configuration
Record model identifiers, revisions, prompt versions, retrieval configuration, decoding settings, seeds where supported, tool versions, and timestamps. Run each critical case more than once. Report distributions or pass consistency—not only the best output.
Step 7: Define release gates before seeing the result
A threshold selected after looking at the score can rationalize almost any outcome. Define gates first. The following are planning examples, not industry benchmarks:
- No unauthorized data disclosure or permission bypass in the test set.
- No critical incorrect answer in a high-risk policy flow.
- Every must-refuse case passes across all required repetitions.
- No regression on previously fixed production failures.
- Retrieval, answer, citation, and escalation gates pass independently.
Step 8: Turn failures into owned work
Classify each failure: source content, metadata, retrieval, reranking, context construction, prompt, model capability, output parsing, permission enforcement, guardrail, or escalation workflow. Assign an owner and add every confirmed fix to the regression suite. Our knowledge base performance benchmarking guide explains how to keep a benchmark decision-specific and reproducible.
Common AI evaluation mistakes
| Mistake | Why it misleads | Better approach |
|---|---|---|
| Testing only answerable questions | The system never demonstrates abstention, clarification, conflict handling, or refusal. | Balance ordinary and edge behaviors explicitly. |
| Scoring only fluency | A polished answer can be unsupported or wrong. | Score correctness, groundedness, completeness, and behavior separately. |
| One run per case | A lucky sample can hide instability. | Repeat fixed cases and report pass consistency. |
| Changing the dataset between candidates | The comparison is no longer controlled. | Freeze cases, evidence, and rubrics before comparison. |
| Using an LLM judge without calibration | The judge can add systematic bias or inconsistency. | Compare judge decisions with trusted ratings and inspect disagreements. |
| Reporting only an average | Critical failures can disappear inside a good mean. | Use category results and zero-tolerance gates for severe risks. |
| Ignoring access control | A factually correct answer can still be unauthorized. | Run the same question under multiple permission personas. |
| No retained raw outputs | Reviewers cannot reproduce or audit the score. | Store prompts, evidence, outputs, settings, scores, and hashes. |
| Treating synthetic data as production proof | Synthetic language may not represent real users or traffic. | Add sampled and adjudicated production cases before launch. |
Content quality is part of answer quality
An evaluation can reveal that the source itself is the problem. Articles with mixed audiences, hidden prerequisites, duplicate policies, weak headings, or unexplained exceptions create ambiguous evidence before a model sees it. Use a consistent knowledge base style guide, give each policy an owner and effective date, mark archived content, and resolve contradictions rather than expecting the model to choose silently.
Quality also includes delivery to people using assistive technology. If an AI answer includes unusable citations, unlabeled controls, keyboard traps, or inaccessible error states, a semantically correct response may still fail the user. See our scoped knowledge base accessibility guide for the interface layer.
Reproducibility files
The Version 1.1 release retains the 200-question dataset, three seeds, recorded model and decoding settings, 600 unchanged raw outputs, original and corrected scored results, aggregate summaries, case-review records, regression tests, validation reports, code, chart generator, licenses, citation metadata, a dependency lockfile, and a complete SHA-256 manifest. Validation confirms 200 unique questions, 40 questions in each category, 600 unique question/run keys, unchanged raw outputs, and no model rerun.
Download the Version 1.1 laboratory archive to inspect the unchanged raw outputs, corrected scores, review records, tests, licenses, citation files, and checksums. Archive SHA-256: 2e26af47842088e8a558be6268771ff617c38a2be6645f55b336f56a09af3a88.
For the audit trail, the original Version 1.0 archive remains available. Its SHA-256 is 31b036c216579d4f431df30debef420f5017923c4d6dea8fdb5c57fc1209a52d.
Cite this study
Knowledge Base Software Editorial Team. (2026). AI Answer Quality Testing for Knowledge Bases: 600-Output Lab (Version 1.1.0) [Dataset and software]. https://knowledge-base.software/guides/ai-answer-quality-testing-knowledge-bases/
Author: Knowledge Base Software Editorial Team. External reviewer: none. Original code: MIT License. Original dataset, results, documentation, review records, and figure: CC BY 4.0. This release is not peer reviewed and is not an industry benchmark. The exact historical model revision and model-file hashes were not recorded and cannot be verified retrospectively.
Version history
- Version 1.1.0 — August 16, 2026: corrected scoring and chart from the unchanged 600 outputs; added review records, tests, licenses, citation metadata, lockfile, and complete release checksums. No model rerun.
- Version 1.0.0 — July 28, 2026: original 200-question, 600-output local laboratory.
Run the method in a smaller trial: download the AI answer quality test dataset with 60 synthetic cases and a three-run Excel log for correct answers, clarification, abstention, conflict handling, citations, and refusal.
AI answer quality testing checklist
- Define the release decision and failure consequences.
- Map retrieval, generation, citations, tools, guardrails, permissions, and escalation.
- Include correct, ambiguous, missing, conflicting, and must-refuse cases.
- Store expected evidence and expected behavior for every case.
- Test retrieval and final answers separately.
- Use deterministic checks for schemas, permissions, citations, and leakage.
- Calibrate semantic judges against trusted ratings.
- Run repeated executions with recorded configuration.
- Report per-category results and critical failures, not only averages.
- Define release gates before examining candidate results.
- Preserve raw outputs and add confirmed failures to regression tests.
- Monitor production drift, feedback, permission errors, and escalations after launch.
Frequently asked questions
What is AI answer quality testing for knowledge bases?
It is the repeatable evaluation of whether a knowledge-base AI retrieves and uses approved evidence correctly, gives a complete and supported answer, cites valid sources, clarifies ambiguity, abstains when information is missing, identifies conflicts, and refuses or escalates when required.
Is answer accuracy the same as groundedness?
No. Accuracy asks whether a claim is correct. Groundedness asks whether the claim is supported by the supplied evidence. A claim can be true in general but ungrounded in the organization’s approved knowledge.
Why test ambiguity and missing information separately?
An ambiguous question may become answerable after the user supplies a missing scope. A missing-information case cannot be answered from the available evidence even when the question is clear. The correct behaviors—and remediation—are different.
How many times should each AI test question run?
There is no universal number. Run enough repetitions to support the risk decision, record the decoding configuration, and increase repetitions for critical or highly variable behavior. Three executions in this laboratory exposed instability but are not enough to estimate rare failures.
Should an LLM grade another LLM’s answers?
It can help at scale, especially for relevance or groundedness, but a judge model is another measurement instrument. Calibrate it against trusted ratings, use deterministic checks where possible, inspect disagreements, and avoid treating one judge score as objective truth.
Does zero secret leakage mean refusal is safe?
No. In our lab, no complete canary appeared, yet the model still failed the strict refusal contract in all 120 must-refuse executions. Leakage, refusal, authorization, and tool behavior require separate tests.
Can these laboratory percentages be used as industry benchmarks?
No. They apply only to this fixed synthetic dataset, prompt, small local model, decoding configuration, and runtime. Reuse the method and cases as a baseline, but measure your own system with your own evidence, risks, languages, and release criteria.
Conclusion
Effective AI answer quality testing for knowledge bases evaluates decisions, not eloquence. The system must answer when evidence is clear, clarify when scope is missing, abstain when knowledge is absent, expose unresolved conflicts, and refuse prohibited requests. Those behaviors should be tested with fixed cases, repeated runs, explicit rubrics, preserved raw outputs, and release gates tied to risk.
Our 600-output laboratory produced a negative result: the small local baseline passed none of the cases under the complete contract. That is exactly why testing is valuable. A documented failure before launch is cheaper and safer than a confident failure in front of users.



