Free browser-based evaluation tool – No sign-up
Knowledge Base Search Relevance Benchmark
Build a repeatable query set, judge the top five results, and measure whether your help center or internal knowledge base puts useful articles first. Get Success@1, Success@3, MRR@5, nDCG@5, no-answer handling, and a prioritized improvement queue.
Nothing is uploaded by this tool. Avoid entering confidential queries, customer data, credentials, or restricted document titles.
Search relevance baseline
This report describes the entered query set. It is not a universal vendor grade. Compare it with later runs that use the same scope and judgment method.
Priority improvement queue
Failures are ordered by business priority and how far the first useful result fell. Review the evidence before changing content or ranking.
Performance by query intent
Small groups can move sharply between runs. Treat an intent row with fewer than five queries as directional.
| Intent | Queries | Success@1 | Success@3 | MRR@5 | nDCG@5 |
|---|
Per-query evidence
| Query | Intent | First useful | Success@1 | nDCG@5 | Diagnosis |
|---|
How to use this result
What a knowledge base search benchmark measures
A search relevance benchmark is a fixed set of information needs, queries, ranked results, and human judgments. It measures whether useful knowledge appears early enough to help a reader – not whether a page merely contains the same words.
Immediate success
Success@1 answers the simplest operational question: did the first result provide a useful answer?
Recoverable search
Success@3 shows whether a useful answer is visible before most readers must scan deeply.
Graded ranking quality
MRR rewards an early first useful result; nDCG also rewards putting the best result ahead of partial answers.
Safe no-answer behavior
Control queries expose search experiences that confidently present an unsupported answer when none exists.
A four-step offline evaluation method
The workflow adapts established information-retrieval evaluation ideas to a practical knowledge base test you can run without connecting an analytics account or search API.
Record the search surface, version, date, cutoff, and evaluator rules.
Build a balanced set from de-identified logs, support cases, tasks, and known risks.
Rate each result 0 – 3 for usefulness and identify target matches where applicable.
Keep the test stable, change the system, rerun it, and inspect query-level differences.
Methodology basis: NIST describes a retrieval test collection as documents, topics, and relevance judgments; Elasticsearch exposes ranking evaluation over typical queries and rated documents. This tool uses a smaller, transparent top-five workflow for operational knowledge bases. See NIST’s TREC overview and the Elasticsearch ranking evaluation reference.
How the search relevance metrics work
No single number tells the whole story. Read the metrics together and keep the per-query evidence available for every reported change.
| Metric | Question answered | Calculation in this tool | Main limitation |
|---|---|---|---|
| Success@1 | Was rank 1 useful? | Share of answer-expected queries with grade 2 – 3 at rank 1. | Ignores useful results below rank 1 and differences between grades 2 and 3. |
| Success@3 | Was a useful result visible early? | Share with at least one grade 2 – 3 result in ranks 1 – 3. | Does not distinguish rank 1 from rank 3. |
| MRR@5 | How early did the first useful result appear? | Mean of 1 divided by the first useful rank; zero when none appears. | Ignores all results after the first useful one. |
| nDCG@5 | Did the ranking put the strongest results first? | Uses gains of 0, 1, 3, and 7 for grades 0 – 3, discounts lower ranks, then normalizes by the ideal order of the judged top five. | Cannot measure relevant documents that were never retrieved or judged. |
| Safe no-answer handling | Did search avoid a false answer? | Share of control queries that clearly abstained or routed safely without claiming an answer. | Depends on a deliberate set of valid no-answer controls. |
Build a query set that reflects reader risk
Start with real patterns, then add coverage intentionally. A benchmark dominated by easy branded queries can improve while users still fail on urgent tasks.
Common searches
Sample high-volume query families after removing personal data and confidential context.
Zero-result and reformulated searches
Include wording that led readers to retry, abandon, or contact support.
High-risk tasks
Over-sample security, billing, access, recovery, and policy needs where a wrong result costs more.
No-answer controls
Test unsupported requests, impossible plan assumptions, and questions that should route to a human or authoritative system.
Turn failures into focused improvements
| Observed failure | Investigate first | Possible action |
|---|---|---|
| No useful result in the top five | Content coverage, index inclusion, permissions, language, query interpretation | Create or repair the answer, fix indexing, add synonyms, or correct access rules. |
| Useful result appears at ranks 2 – 5 | Title, headings, metadata, field weighting, freshness, duplicate candidates | Clarify the target article and tune ranking with a controlled rerun. |
| Related pages outrank the exact answer | Keyword overlap, generic titles, navigation pages, overly broad content | Improve specificity, consolidate overlap, or adjust boosts and filters. |
| Unsafe result on a no-answer query | Confidence threshold, fallback behavior, answer generation, source constraints | Add abstention rules and a safe escalation or navigation path. |
| Offline score rises but users still escalate | Snippet clarity, page usability, article accuracy, task completion, latency | Pair the benchmark with click, task-success, deflection, and support evidence. |
This tool evaluates ranked article retrieval. Use AI answer quality testing for generated answers and citations, and knowledge base performance benchmarking for adoption, self-service, and operating outcomes.
Knowledge base search relevance benchmark questions
What is a knowledge base search relevance benchmark?
It is a repeatable offline test that runs a fixed set of queries, records ranked results, and applies consistent human relevance judgments. It helps compare search configurations or releases using the same evidence.
How many queries should I test?
The tool accepts three or more so you can learn the workflow. A decision-quality benchmark normally needs a larger, representative sample across intents, frequency, and business risk. Expand it until important query groups are no longer represented by one or two examples.
Who should rate result relevance?
Use people who understand the reader need and the authoritative content. Give evaluators written rating guidance, calibrate on a shared sample, and review disagreements for high-impact queries. Do not treat automated labels as unquestioned ground truth.
What is the difference between MRR and nDCG?
MRR uses only the rank of the first useful result. nDCG uses every judged result in the cutoff and gives more credit when highly relevant results appear above partial answers. Read both because they expose different ranking behavior.
Why test queries with no correct answer?
A knowledge search can look helpful while presenting an irrelevant result with too much confidence. No-answer controls measure whether the experience abstains or offers a safe route when the knowledge base cannot support the request.
Can I compare two knowledge base software products with this tool?
Yes, if both products search the same content, permissions, language, query set, and result cutoff, and the judgments are applied consistently. If the collections or scopes differ, the result describes the full search experience rather than the ranking algorithm alone.
Does a high offline relevance score prove customer self-service success?
No. It shows that judged results rank well for the sampled queries. Validate the live experience with accessibility, latency, snippets, clicks, task completion, article accuracy, repeat searches, escalation, and support outcomes.
Does this tool send my search data anywhere?
The benchmark calculations and exports run in your browser and require no account. Do not enter personal data, secrets, customer records, or restricted document titles; the website itself may still use ordinary hosting, security, or analytics services under its published policies.
How is relevance testing different from functional search testing?
Functional testing checks whether the search box, filters, links, and permissions operate. Relevance testing checks whether the ranked results satisfy the information need and put the strongest answer early enough.
Does this benchmark evaluate generated AI or RAG answers?
No. It evaluates ranked retrieval of knowledge base articles or documents. Generated-answer accuracy, source support, citation quality, and abstention need a separate answer-quality test.
Related knowledge base planning tools
Published by the Knowledge Base Software Editorial Team. The methodology is designed for transparent comparison, not certification or a guaranteed business outcome.
