Free browser-based evaluation tool – No sign-up

Knowledge Base Search Relevance Benchmark

Build a repeatable query set, judge the top five results, and measure whether your help center or internal knowledge base puts useful articles first. Get Success@1, Success@3, MRR@5, nDCG@5, no-answer handling, and a prioritized improvement queue.

Start a benchmark
Transparent metricsEvery score has a visible definition and per-query evidence.
Human relevance judgmentsYou decide what is useful; the tool does not invent ground truth.
Local calculationYour entries are processed in this browser, with no account required.
Step 1 of 4Set scope

Nothing is uploaded by this tool. Avoid entering confidential queries, customer data, credentials, or restricted document titles.

Define a comparable search test

Record the search surface and version you are testing. Reuse the same query set and judgment rules after a configuration, content, or software change.

Use a name you can recognize when comparing runs.
Comparison rule: keep the query set, top-five cutoff, rating scale, and evaluator guidance stable between runs. Change one search or content variable at a time where practical.

What a knowledge base search benchmark measures

A search relevance benchmark is a fixed set of information needs, queries, ranked results, and human judgments. It measures whether useful knowledge appears early enough to help a reader – not whether a page merely contains the same words.

Rank 1

Immediate success

Success@1 answers the simplest operational question: did the first result provide a useful answer?

Top 3

Recoverable search

Success@3 shows whether a useful answer is visible before most readers must scan deeply.

Order

Graded ranking quality

MRR rewards an early first useful result; nDCG also rewards putting the best result ahead of partial answers.

Restraint

Safe no-answer behavior

Control queries expose search experiences that confidently present an unsupported answer when none exists.

A four-step offline evaluation method

The workflow adapts established information-retrieval evaluation ideas to a practical knowledge base test you can run without connecting an analytics account or search API.

1Freeze the scope

Record the search surface, version, date, cutoff, and evaluator rules.

2Sample real needs

Build a balanced set from de-identified logs, support cases, tasks, and known risks.

3Judge top results

Rate each result 0 – 3 for usefulness and identify target matches where applicable.

4Compare and diagnose

Keep the test stable, change the system, rerun it, and inspect query-level differences.

Methodology basis: NIST describes a retrieval test collection as documents, topics, and relevance judgments; Elasticsearch exposes ranking evaluation over typical queries and rated documents. This tool uses a smaller, transparent top-five workflow for operational knowledge bases. See NIST’s TREC overview and the Elasticsearch ranking evaluation reference.

How the search relevance metrics work

No single number tells the whole story. Read the metrics together and keep the per-query evidence available for every reported change.

MetricQuestion answeredCalculation in this toolMain limitation
Success@1Was rank 1 useful?Share of answer-expected queries with grade 2 – 3 at rank 1.Ignores useful results below rank 1 and differences between grades 2 and 3.
Success@3Was a useful result visible early?Share with at least one grade 2 – 3 result in ranks 1 – 3.Does not distinguish rank 1 from rank 3.
MRR@5How early did the first useful result appear?Mean of 1 divided by the first useful rank; zero when none appears.Ignores all results after the first useful one.
nDCG@5Did the ranking put the strongest results first?Uses gains of 0, 1, 3, and 7 for grades 0 – 3, discounts lower ranks, then normalizes by the ideal order of the judged top five.Cannot measure relevant documents that were never retrieved or judged.
Safe no-answer handlingDid search avoid a false answer?Share of control queries that clearly abstained or routed safely without claiming an answer.Depends on a deliberate set of valid no-answer controls.

Build a query set that reflects reader risk

Start with real patterns, then add coverage intentionally. A benchmark dominated by easy branded queries can improve while users still fail on urgent tasks.

Frequency

Common searches

Sample high-volume query families after removing personal data and confidential context.

Failure

Zero-result and reformulated searches

Include wording that led readers to retry, abandon, or contact support.

Impact

High-risk tasks

Over-sample security, billing, access, recovery, and policy needs where a wrong result costs more.

Restraint

No-answer controls

Test unsupported requests, impossible plan assumptions, and questions that should route to a human or authoritative system.

Turn failures into focused improvements

Observed failureInvestigate firstPossible action
No useful result in the top fiveContent coverage, index inclusion, permissions, language, query interpretationCreate or repair the answer, fix indexing, add synonyms, or correct access rules.
Useful result appears at ranks 2 – 5Title, headings, metadata, field weighting, freshness, duplicate candidatesClarify the target article and tune ranking with a controlled rerun.
Related pages outrank the exact answerKeyword overlap, generic titles, navigation pages, overly broad contentImprove specificity, consolidate overlap, or adjust boosts and filters.
Unsafe result on a no-answer queryConfidence threshold, fallback behavior, answer generation, source constraintsAdd abstention rules and a safe escalation or navigation path.
Offline score rises but users still escalateSnippet clarity, page usability, article accuracy, task completion, latencyPair the benchmark with click, task-success, deflection, and support evidence.

This tool evaluates ranked article retrieval. Use AI answer quality testing for generated answers and citations, and knowledge base performance benchmarking for adoption, self-service, and operating outcomes.

Knowledge base search relevance benchmark questions

What is a knowledge base search relevance benchmark?

It is a repeatable offline test that runs a fixed set of queries, records ranked results, and applies consistent human relevance judgments. It helps compare search configurations or releases using the same evidence.

How many queries should I test?

The tool accepts three or more so you can learn the workflow. A decision-quality benchmark normally needs a larger, representative sample across intents, frequency, and business risk. Expand it until important query groups are no longer represented by one or two examples.

Who should rate result relevance?

Use people who understand the reader need and the authoritative content. Give evaluators written rating guidance, calibrate on a shared sample, and review disagreements for high-impact queries. Do not treat automated labels as unquestioned ground truth.

What is the difference between MRR and nDCG?

MRR uses only the rank of the first useful result. nDCG uses every judged result in the cutoff and gives more credit when highly relevant results appear above partial answers. Read both because they expose different ranking behavior.

Why test queries with no correct answer?

A knowledge search can look helpful while presenting an irrelevant result with too much confidence. No-answer controls measure whether the experience abstains or offers a safe route when the knowledge base cannot support the request.

Can I compare two knowledge base software products with this tool?

Yes, if both products search the same content, permissions, language, query set, and result cutoff, and the judgments are applied consistently. If the collections or scopes differ, the result describes the full search experience rather than the ranking algorithm alone.

Does a high offline relevance score prove customer self-service success?

No. It shows that judged results rank well for the sampled queries. Validate the live experience with accessibility, latency, snippets, clicks, task completion, article accuracy, repeat searches, escalation, and support outcomes.

Does this tool send my search data anywhere?

The benchmark calculations and exports run in your browser and require no account. Do not enter personal data, secrets, customer records, or restricted document titles; the website itself may still use ordinary hosting, security, or analytics services under its published policies.

How is relevance testing different from functional search testing?

Functional testing checks whether the search box, filters, links, and permissions operate. Relevance testing checks whether the ranked results satisfy the information need and put the strongest answer early enough.

Does this benchmark evaluate generated AI or RAG answers?

No. It evaluates ranked retrieval of knowledge base articles or documents. Generated-answer accuracy, source support, citation quality, and abstention need a separate answer-quality test.