Free human-reviewed evaluation kit   –   No sign-up

AI Answer Grounding Test Kit

Audit AI knowledge base answers claim by claim. Map each factual statement to approved evidence, check citations and source boundaries, test safe no-answer behavior, and export a review record your team can rerun.

Start a grounding test
Claim-level evidenceEvery factual claim receives its own support and citation judgment.
Grounded is not “true”Source approval and currentness stay separate from claim support.
Private by designThe test runs in this browser without uploading the entered content.
Step 1 of 4Set scope

Nothing is sent to an API or saved by this page. Download the report before leaving if you need to keep it.

Define the evaluation boundary

Record the exact system, model, evidence boundary, reader context, and date. A grounding result only applies to this frozen setup.

Boundary rule: judge support only against evidence that was allowed and available in this run. External knowledge may make a statement plausible, but it does not make the answer grounded to this evidence set.

What an AI answer grounding test measures

Grounding asks a narrow, auditable question: does each factual claim in the generated answer follow from evidence the system was allowed to use? The test should not be confused with truth checking, retrieval quality, helpfulness, or business outcomes.

Support

Claim grounding

Break the answer into atomic claims and map each one to a passage that fully supports its meaning and conditions.

Citations

Evidence alignment

Check whether a citation exists, points to the right passage, and actually supports the attached claim.

Boundaries

Safe behavior

Confirm that the assistant clarifies, abstains, refuses, surfaces conflict, or escalates when the evidence demands it.

Sources

Evidence hygiene

Record approval, currentness, and access separately. A stale source can support a claim while still making the answer unsafe.

Run the kit as a repeatable evidence audit

Use the same frozen test set to compare a baseline with a bounded model, prompt, retrieval, citation, or guardrail change.

1Freeze the boundary

Record the model, prompt, corpus or retrieved context, permissions, language, and date.

2Capture exact outputs

Store the original question, answer, evidence pack, expected behavior, and risk.

3Judge atomic claims

Label full, partial, unsupported, contradicted, or non-factual; then judge citations separately.

4Fix and rerun

Export the evidence, address the priority queue, and rerun the same cases under the same rubric.

A strict claim and citation rubric

Keep the unit of review small. One sentence may contain several claims with different support outcomes.

LabelUse it whenStrict result
Fully supportedThe complete atomic claim follows from one or more identified passages, including qualifiers, scope, dates, audience, and permissions.Passes the claim-support gate.
Partly supportedSome of the claim follows from the evidence, but a material condition or detail is missing or changed.Fails strict grounding; split the claim or correct the answer.
UnsupportedNo allowed passage supports the factual statement.Fails; may indicate fabrication or evidence leakage.
ContradictedThe allowed evidence directly conflicts with the claim.Release-blocking when material.
Non-factualThe text is a greeting, transition, question, or other statement that does not assert a checkable fact.Excluded from claim denominators.
Citation correctness asks whether a supplied citation supports its claim. Citation completeness asks whether claims that require citations received one. The report keeps both denominators visible and also shows correct citation coverage.

Common grounding failures and the first place to investigate

Observed failureInvestigate firstPossible action
Unsupported detail appears in an otherwise correct answerCompound claims, prompt incentives, model prior knowledge, missing evidenceSplit claims, constrain the answer, retrieve the missing source, or require abstention.
A citation exists but does not support the nearby claimCitation placement, chunk mapping, source ID handling, post-generation citation attachmentBind citations at claim generation time and verify the exact passage.
The assistant answers when the evidence is insufficientConfidence or evidence gate, fallback prompt, no-answer training casesAdd explicit abstention tests and a safe escalation route.
The answer chooses one of two current conflicting sourcesSource governance, recency metadata, conflict detection, authority rulesResolve the source conflict or require the assistant to surface it.
A claim is grounded to restricted contentIdentity context, permission filters, cache boundaries, tool authorizationTreat this as an access-control incident, not a relevance problem.
Grounding is high but the answer is still wrong for usersSource accuracy, currentness, completeness, answer relevance, task successRepair the source and run separate quality and user-outcome tests.

How to build a useful grounding test set

Start with de-identified production questions, then add designed boundary cases. Keep each output and evidence pack frozen so changes remain comparable.

Ordinary

Answerable questions

Include direct facts, multi-step procedures, scoped policy questions, and answers that require two or more sources.

Ambiguity

Clarification cases

Use questions missing a product, plan, role, location, version, or object that changes the correct answer.

Evidence gaps

No-answer controls

Include plausible requests for which the allowed evidence contains no complete answer.

Safety

Conflict and refusal

Test contradictory sources, restricted data, destructive actions, and requests outside the assistant’s authority.

Do not tune a candidate on the same cases and then present the final score as independent evidence. Preserve a held-out set for consequential release decisions, and repeat human review when the source snapshot or policy boundary changes.

Methodology and limitations

This kit applies a human-reviewed, claim-level evidence audit. Google Cloud’s grounding documentation describes perfect grounding as every claim being wholly supported by supplied facts, while Microsoft’s RAG evaluation guidance separates groundedness from retrieval, relevance, and response completeness. NIST’s Generative AI Profile recommends documented testing, evaluation, verification, and validation across the AI lifecycle.

  • The report applies only to the entered questions, outputs, evidence, permissions, judgments, and date.
  • Human reviewers can disagree; document the rubric and adjudicate uncertain high-risk cases.
  • A grounded answer can repeat an inaccurate or stale source. Grounding is not a factual-truth guarantee.
  • The kit does not inspect model internals, retrieve documents, call an LLM judge, or automatically verify real-world facts.
  • The kit does not measure search ranking. Use the Knowledge Base Search Relevance Benchmark for ranked retrieval.
  • The kit does not prove task completion, user satisfaction, ticket deflection, or financial impact.

Primary references: Google Cloud Check Grounding, Microsoft RAG evaluators, and the NIST AI RMF Generative AI Profile.

AI answer grounding test FAQ

What is AI answer grounding?

Grounding is the degree to which factual claims in an AI answer are supported by the evidence the system was permitted to use for that response.

Is grounding the same as factual accuracy?

No. A claim can faithfully reflect a stale, incomplete, or incorrect source. Source accuracy and currentness need their own governance checks.

What is an atomic claim?

An atomic claim is one checkable factual assertion. Split sentences when names, numbers, dates, conditions, audiences, or permissions could receive different evidence judgments.

How is the fully supported claim rate calculated?

It is the number of factual claims labeled fully supported divided by all judged factual claims. Partly supported claims do not pass the strict numerator.

How are citation correctness and completeness different?

Correctness uses citations that were supplied as its denominator and asks whether they support the claim. Completeness uses claims that require citations and asks whether a citation was supplied. Correct citation coverage requires both.

How should no-answer cases be tested?

Record an evidence pack that is deliberately insufficient and set the expected behavior to abstain, clarify, refuse, surface conflict, or escalate. Do not invent claim-support scores for an answer that correctly makes no factual claim.

Can the scores be compared across different assistants?

Only when the assistants receive the same questions, evidence boundary, permission context, expected behaviors, and human rubric. Otherwise the numbers describe different tests.

Does the tool send answers or sources to an AI service?

No. The tool is a local browser worksheet and calculator. It does not call an AI model or upload the entered test data.

Does this replace a full AI answer quality evaluation?

No. It covers grounding, citations, source boundaries, and expected evidence behavior. Use the broader AI answer quality testing methodology for relevance, completeness, ambiguity, conflict, refusal, repeatability, and release design.