Free human-reviewed evaluation kit – No sign-up
AI Answer Grounding Test Kit
Audit AI knowledge base answers claim by claim. Map each factual statement to approved evidence, check citations and source boundaries, test safe no-answer behavior, and export a review record your team can rerun.
Nothing is sent to an API or saved by this page. Download the report before leaving if you need to keep it.
Grounding audit report
Claim support breakdown
Partial support, unsupported additions, and contradictions remain visible instead of being hidden inside one overall score.
Priority review queue
Resolve the highest-risk behavior, source-boundary, contradiction, unsupported-claim, and citation failures first.
Evidence by test case
| Question | Expected to observed | Claims | Full support | Correct citation coverage | Source boundary | Finding |
|---|
Citation diagnostics
| Metric | Result | Denominator | What it tells you |
|---|
What an AI answer grounding test measures
Grounding asks a narrow, auditable question: does each factual claim in the generated answer follow from evidence the system was allowed to use? The test should not be confused with truth checking, retrieval quality, helpfulness, or business outcomes.
Claim grounding
Break the answer into atomic claims and map each one to a passage that fully supports its meaning and conditions.
Evidence alignment
Check whether a citation exists, points to the right passage, and actually supports the attached claim.
Safe behavior
Confirm that the assistant clarifies, abstains, refuses, surfaces conflict, or escalates when the evidence demands it.
Evidence hygiene
Record approval, currentness, and access separately. A stale source can support a claim while still making the answer unsafe.
Run the kit as a repeatable evidence audit
Use the same frozen test set to compare a baseline with a bounded model, prompt, retrieval, citation, or guardrail change.
Record the model, prompt, corpus or retrieved context, permissions, language, and date.
Store the original question, answer, evidence pack, expected behavior, and risk.
Label full, partial, unsupported, contradicted, or non-factual; then judge citations separately.
Export the evidence, address the priority queue, and rerun the same cases under the same rubric.
A strict claim and citation rubric
Keep the unit of review small. One sentence may contain several claims with different support outcomes.
| Label | Use it when | Strict result |
|---|---|---|
| Fully supported | The complete atomic claim follows from one or more identified passages, including qualifiers, scope, dates, audience, and permissions. | Passes the claim-support gate. |
| Partly supported | Some of the claim follows from the evidence, but a material condition or detail is missing or changed. | Fails strict grounding; split the claim or correct the answer. |
| Unsupported | No allowed passage supports the factual statement. | Fails; may indicate fabrication or evidence leakage. |
| Contradicted | The allowed evidence directly conflicts with the claim. | Release-blocking when material. |
| Non-factual | The text is a greeting, transition, question, or other statement that does not assert a checkable fact. | Excluded from claim denominators. |
Common grounding failures and the first place to investigate
| Observed failure | Investigate first | Possible action |
|---|---|---|
| Unsupported detail appears in an otherwise correct answer | Compound claims, prompt incentives, model prior knowledge, missing evidence | Split claims, constrain the answer, retrieve the missing source, or require abstention. |
| A citation exists but does not support the nearby claim | Citation placement, chunk mapping, source ID handling, post-generation citation attachment | Bind citations at claim generation time and verify the exact passage. |
| The assistant answers when the evidence is insufficient | Confidence or evidence gate, fallback prompt, no-answer training cases | Add explicit abstention tests and a safe escalation route. |
| The answer chooses one of two current conflicting sources | Source governance, recency metadata, conflict detection, authority rules | Resolve the source conflict or require the assistant to surface it. |
| A claim is grounded to restricted content | Identity context, permission filters, cache boundaries, tool authorization | Treat this as an access-control incident, not a relevance problem. |
| Grounding is high but the answer is still wrong for users | Source accuracy, currentness, completeness, answer relevance, task success | Repair the source and run separate quality and user-outcome tests. |
How to build a useful grounding test set
Start with de-identified production questions, then add designed boundary cases. Keep each output and evidence pack frozen so changes remain comparable.
Answerable questions
Include direct facts, multi-step procedures, scoped policy questions, and answers that require two or more sources.
Clarification cases
Use questions missing a product, plan, role, location, version, or object that changes the correct answer.
No-answer controls
Include plausible requests for which the allowed evidence contains no complete answer.
Conflict and refusal
Test contradictory sources, restricted data, destructive actions, and requests outside the assistant’s authority.
Do not tune a candidate on the same cases and then present the final score as independent evidence. Preserve a held-out set for consequential release decisions, and repeat human review when the source snapshot or policy boundary changes.
Methodology and limitations
This kit applies a human-reviewed, claim-level evidence audit. Google Cloud’s grounding documentation describes perfect grounding as every claim being wholly supported by supplied facts, while Microsoft’s RAG evaluation guidance separates groundedness from retrieval, relevance, and response completeness. NIST’s Generative AI Profile recommends documented testing, evaluation, verification, and validation across the AI lifecycle.
- The report applies only to the entered questions, outputs, evidence, permissions, judgments, and date.
- Human reviewers can disagree; document the rubric and adjudicate uncertain high-risk cases.
- A grounded answer can repeat an inaccurate or stale source. Grounding is not a factual-truth guarantee.
- The kit does not inspect model internals, retrieve documents, call an LLM judge, or automatically verify real-world facts.
- The kit does not measure search ranking. Use the Knowledge Base Search Relevance Benchmark for ranked retrieval.
- The kit does not prove task completion, user satisfaction, ticket deflection, or financial impact.
Primary references: Google Cloud Check Grounding, Microsoft RAG evaluators, and the NIST AI RMF Generative AI Profile.
AI answer grounding test FAQ
What is AI answer grounding?
Grounding is the degree to which factual claims in an AI answer are supported by the evidence the system was permitted to use for that response.
Is grounding the same as factual accuracy?
No. A claim can faithfully reflect a stale, incomplete, or incorrect source. Source accuracy and currentness need their own governance checks.
What is an atomic claim?
An atomic claim is one checkable factual assertion. Split sentences when names, numbers, dates, conditions, audiences, or permissions could receive different evidence judgments.
How is the fully supported claim rate calculated?
It is the number of factual claims labeled fully supported divided by all judged factual claims. Partly supported claims do not pass the strict numerator.
How are citation correctness and completeness different?
Correctness uses citations that were supplied as its denominator and asks whether they support the claim. Completeness uses claims that require citations and asks whether a citation was supplied. Correct citation coverage requires both.
How should no-answer cases be tested?
Record an evidence pack that is deliberately insufficient and set the expected behavior to abstain, clarify, refuse, surface conflict, or escalate. Do not invent claim-support scores for an answer that correctly makes no factual claim.
Can the scores be compared across different assistants?
Only when the assistants receive the same questions, evidence boundary, permission context, expected behaviors, and human rubric. Otherwise the numbers describe different tests.
Does the tool send answers or sources to an AI service?
No. The tool is a local browser worksheet and calculator. It does not call an AI model or upload the entered test data.
Does this replace a full AI answer quality evaluation?
No. It covers grounding, citations, source boundaries, and expected evidence behavior. Use the broader AI answer quality testing methodology for relevance, completeness, ambiguity, conflict, refusal, repeatability, and release design.
