Vector Search for Knowledge Bases: Original Lab Results

Vector search for knowledge bases can retrieve the right article when a visitor’s words do not match the documentation. But “semantic” does not automatically mean “better.” In our original local lab, vector-only search ranked the labeled answer first for 90.3% of 72 fixed queries. The pre-set hybrid and reranking methods both performed worse.

We tested BM25, dense vectors, Reciprocal Rank Fusion, and a transparent second-stage reranker against the same 48-document support corpus. This guide publishes the complete method, results, failure cases, and practical decisions you can use when building or buying knowledge base search.

Short answer: use labeled queries from your own knowledge base to choose a retrieval method. In this controlled test, vector search handled paraphrases far better than BM25, but equal-weight hybrid fusion diluted that advantage. BM25 remained dramatically faster and perfect on the exact keyword queries.

Original vector search laboratory comparing Recall at 1, median latency, and keyword, paraphrase, and scenario performance for BM25, vector, hybrid RRF, and hybrid with a fixed reranker.
Knowledge Base Software laboratory, July 28, 2026. The results come from 48 fixed documents and 72 pre-labeled queries; local latency includes query embedding but excludes network and hosted-search latency.

What is vector search for knowledge bases?

Vector search converts a query and each searchable passage into numerical embeddings. A retrieval system then compares the query vector with document vectors and ranks the nearest items. The purpose is to capture meaning rather than require identical words.

For example, an article may say “a failed subscription payment is retried,” while a customer searches for “the dunning schedule after a card is declined.” A lexical system sees limited word overlap. A suitable embedding model may place both texts close together because they describe the same intent.

The underlying idea is not that vectors “understand” an article as a person does. An embedding model maps text into a fixed-dimensional space learned from its training data. Documents that the model considers similar receive nearby vectors. The Sentence-BERT paper describes an architecture for producing sentence embeddings that can be compared with cosine similarity. The all-MiniLM-L6-v2 model card identifies the model used in our lab as a 384-dimensional encoder intended for clustering and semantic search.

Vector search, keyword search, hybrid search, and reranking

  • Keyword search ranks documents from literal term evidence. BM25 rewards useful term matches while accounting for term rarity and document length.
  • Vector search ranks documents by embedding similarity, so it can connect paraphrases that share little vocabulary.
  • Hybrid search runs lexical and vector retrieval, then combines their rankings or scores.
  • Reranking takes a smaller candidate set and applies a second scoring stage intended to improve the final order.

Hybrid search is attractive because keyword and semantic retrieval fail differently. Exact error codes, product names, dates, and policy identifiers often favor lexical search. Natural-language paraphrases can favor vectors. Microsoft’s hybrid search documentation describes the same broad pattern: full-text and vector queries run in parallel and a fusion method produces one result list.

Our original vector-search experiment

This was a laboratory retrieval test, not a vendor benchmark and not a study of customer search logs. We wrote a synthetic product-support corpus so that every document, distractor, query, and relevance label could be published and reproduced without exposing private data.

The fixed corpus and labeled queries

The corpus contained 48 short English documents across identity, billing, developer APIs, data management, publishing, mobile apps, administration, and search. Twenty-four documents contained a canonical answer. Each had a same-topic distractor that was plausible but did not contain the labeled answer.

We wrote 72 answerable queries before the final run:

  • 24 keyword queries used important terms from the answer, such as “HTTP 429 Retry-After exponential backoff full jitter.”
  • 24 paraphrase queries expressed the same need with different language, such as “What is the dunning schedule when a subscription charge is declined?”
  • 24 scenario queries described a user’s task or problem, such as “How do I make downloaded timestamps use Europe/Paris instead of universal time?”

Every query had one pre-labeled grade-3 answer document. The labels were part of the frozen test fixture; they were not changed after seeing the rankings. Corpus and query files have published SHA-256 checksums in the laboratory record.

The four matched retrieval methods

1. BM25. The lexical baseline used k1=1.2 and b=0.75. Titles were concatenated twice before the body to apply the same simple title boost to all documents. The original Okapi at TREC-3 publication is an early primary reference for the Okapi probabilistic retrieval work that led to BM25.

2. Dense vector search. We used Xenova/all-MiniLM-L6-v2 at pinned revision 751bff37182d3f1213fa05d7196b954e230abad9, with q8 model weights, mean pooling, L2 normalization, and exhaustive cosine scoring. The model generated 384-dimensional vectors. Document embeddings were created once; generating the query embedding was included in query latency.

3. Hybrid RRF. We fused the complete BM25 and vector rankings with Reciprocal Rank Fusion using k=60. For every document, the system added 1 / (60 + rank) from each ranking. This avoids combining raw BM25 and cosine scores, which have incompatible scales. The method follows the formulation in the original RRF paper by Cormack, Clarke, and Büttcher; Microsoft also documents the RRF scoring process.

4. Hybrid plus a fixed feature reranker. The top 20 RRF candidates were rescored with a formula fixed before the final run: 30% normalized BM25, 25% vector cosine, 20% IDF-weighted query-term coverage, 10% title coverage, 7.5% adjacent-query-bigram rate, 5% term proximity, and 2.5% identifier coverage.

This fourth method is a transparent deterministic feature reranker. It is not a neural cross-encoder, and the results should not be used to judge cross-encoder products. We chose an inspectable formula so every candidate score could be exported and audited.

Metrics and timing boundary

Recall@1 asks whether the labeled answer appeared first. Recall@3 and Recall@5 ask whether it appeared in the first three or five results. MRR@10 rewards an earlier first relevant result within the top ten; NIST’s TREC description of mean reciprocal rank defines reciprocal rank from the position of the first correct answer. nDCG@10 captures ranked gain with a discount for later positions.

The local timing includes BM25 scoring, query embedding where applicable, exhaustive cosine scoring, fusion, and reranking. It excludes network travel, HTTP handling, JSON serialization, a hosted vector database, and access-control filters. One warm-up embedding ran before measurement, and the final record contains one local timing observation per query. Treat latency as a within-lab comparison, not a production service-level benchmark.

Results: vector-only ranked first in this lab

MethodRecall@1Recall@3MRR@10nDCG@10Median local latencyp95 local latency
BM2572.2%91.7%0.8220.8590.3 ms1.1 ms
Vector90.3%98.6%0.9420.95723.6 ms35.5 ms
Hybrid RRF83.3%94.4%0.8970.92024.2 ms36.8 ms
Hybrid + fixed reranker81.9%93.1%0.8840.90925.4 ms38.9 ms
One final local pass over 72 frozen queries. Higher relevance metrics are better; lower latency is better.

Vector search put the labeled answer first for 65 of 72 queries. Hybrid RRF did so for 60, the fixed reranker for 59, and BM25 for 52. Vector also achieved the best Recall@3, MRR@10, and nDCG@10 in this dataset.

The result does not prove that vector search is universally superior. It says that this model, corpus, and query distribution favored vector retrieval. It also falsifies a tempting assumption: adding lexical results and a reranker did not automatically improve this system.

Query style explains much of the difference

MethodKeyword Recall@1Paraphrase Recall@1Scenario Recall@1
BM25100.0%50.0%66.7%
Vector100.0%87.5%83.3%
Hybrid RRF100.0%79.2%70.8%
Hybrid + fixed reranker100.0%62.5%83.3%

All four methods were perfect on the 24 keyword-form queries. The separation came from the 48 paraphrase and scenario queries. BM25 ranked only half of the paraphrase answers first. Vector search reached 87.5%. That is the strongest evidence in this lab for using embeddings: visitors do not need to know the vocabulary used by the writer.

The fixed reranker shows the opposite lesson. Its literal term-coverage and title features helped scenario questions relative to plain RRF, but paraphrase Recall@1 fell to 62.5%. A reranker that rewards lexical overlap can undo the semantic signal you added vectors to capture.

Why hybrid and reranking lost here

RRF is a fusion rule, not a guarantee. Our equal-weight configuration assumed BM25 and vector rankings deserved the same influence. That assumption was wrong for this query mix because the vector ranker was materially stronger on paraphrases. Equal fusion pulled some good vector results down when BM25 disagreed.

One example was “Can I retrieve paid billing records after closing the workspace?” Vector search chose a general workspace-deletion article first and placed the labeled invoice-retention article second. The question mixes billing, retention, and closure, so its semantic neighborhood contains several plausible documents. Hybrid fusion did not resolve the ambiguity because the lexical ranking also favored nearby retention language.

Another query asked which response header tells a client how long to pause after exceeding an API quota. The labeled answer was the Retry-After article, but vector search ranked a related rate-limit-header article first. That “distractor” genuinely discusses quota headers, making the single-label judgment stricter than some users might be. This exposes an evaluation issue as well as a retrieval issue: production test sets should permit multiple relevance grades when several articles are useful.

The fixed reranker had a different failure mode. For the paraphrase “A teammate typed the wrong password many times; when can they try again?”, it promoted the password-reset distractor over the account-unlock answer. Literal overlap with “password” was strong, but the task intent was lockout recovery. A learned reranker trained on representative judgments might handle this better; our deterministic formula did not.

When should a knowledge base use vector search?

Vector retrieval is worth testing when the words in customer questions regularly differ from the words in your documentation. Typical signals include failed searches that are obvious paraphrases, conversational queries, internal terminology that differs by team, and support questions that describe symptoms rather than feature names.

  • Choose a BM25-first baseline when exact identifiers dominate, the corpus is small and well titled, latency is extremely constrained, or you do not yet have a labeled evaluation set.
  • Test vector search when paraphrases, natural-language questions, or symptom descriptions are common.
  • Test hybrid search when the same query stream contains both semantic questions and exact codes, names, versions, or policy numbers.
  • Add a reranker only when initial retrieval has enough candidate recall and a labeled set can show that the second stage improves the final positions.

Do not select a method from a generic benchmark alone. Your content structure, terminology, languages, query distribution, and definition of relevance determine the winner. A model that performs well on general sentence similarity may miss your product codes or policy language.

How to implement vector search in a knowledge base

1. Define the retrieval unit

Decide whether one searchable item is an article, a section, or a smaller passage. Whole articles preserve context but can blend unrelated topics into one vector. Small chunks improve topical focus but may lose prerequisites, scope, or warnings. A useful support chunk should usually stand on its own: include the article title, section heading, product or plan scope, and enough text to answer one intent.

Keep the displayed article as the destination even if search indexes sections. Store a stable article ID, section anchor, locale, product version, visibility rules, and last-updated timestamp with every vector.

2. Create one embedding policy

Use the same model and compatible preprocessing for document and query embeddings. Record the model name, revision, vector dimension, pooling, normalization, distance metric, maximum input length, and truncation behavior. Re-embedding only half an index with a different model creates incomparable vectors.

Pin a model revision in reproducible tests. Our lab recorded the repository commit, q8 weights, mean pooling, normalization, and cosine scoring. “We use MiniLM” is not enough information to recreate a result.

3. Enforce permissions and scope

Search must never retrieve content a user cannot access. Store tenant, workspace, role, collection, locale, status, and product-version metadata in the index. Apply authorization filtering inside the retrieval path, not as a best-effort cleanup after confidential passages have already been sent to an answer generator.

4. Build lexical and semantic baselines separately

Measure BM25 and vector search independently before fusing them. This tells you whether hybrid search adds complementary evidence or merely mixes a strong ranker with a weak one. Preserve the top candidates and raw component ranks for failure analysis.

5. Tune fusion on held-out labels

Equal-weight RRF is a sensible baseline because it combines ranks instead of incompatible raw score scales. It is not sacred. Test candidate depth, RRF constant, lexical/vector weights where supported, and query-specific routing. Use a development set for tuning and a separate test set for the final report; otherwise you will overfit the same questions used to judge success.

6. Separate retrieval from answer quality

A correct search result can still produce a wrong generated answer, and a weak top result can sometimes be rescued from deeper context. Track retrieval metrics before evaluating a RAG or AI-answer layer. At minimum, log the query, candidate IDs, ranks, model revision, filters, selected context, answer citations, and user outcome.

How to evaluate your own knowledge base search

A useful evaluation set is more important than a sophisticated retrieval stack. Start with real, consented search or support intents when available, remove personal data, and ask knowledgeable reviewers to identify every useful document. If real logs are unavailable, write a clearly labeled synthetic set and include hard same-topic distractors as we did.

  1. Freeze a content snapshot and give every searchable unit a stable ID.
  2. Create keyword, paraphrase, scenario, typo, identifier, multilingual, and no-answer query groups that reflect your users.
  3. Use graded relevance when more than one document is useful.
  4. Keep tuning queries separate from final test queries.
  5. Report Recall@k and nDCG or MRR alongside latency, index cost, and update delay.
  6. Inspect failures by query group instead of relying on one aggregate score.
  7. Repeat measurements on the actual service path, including network, permissions, filters, and approximate indexing.

For a help center, Recall@1 reflects whether the first result answers the question. Recall@3 matters when users can quickly scan a few options or when a RAG system passes several candidates downstream. nDCG is useful when multiple documents have different usefulness grades. No-answer precision is essential if your interface may claim that nothing relevant exists.

Performance and operating costs

BM25 took a median 0.3 milliseconds in this in-process, 48-document lab. Vector search took 23.6 milliseconds because it generated a query embedding locally. Hybrid and reranking added small scoring costs on top. These numbers are not a capacity forecast; a production architecture adds transport, queues, authentication, filters, approximate-nearest-neighbor lookup, observability, and regional latency.

Document embedding is usually an indexing cost, while query embedding is a request-time cost. Cache only when privacy and query repetition justify it. Re-embed changed chunks, not the entire corpus, but schedule a complete rebuild when changing models or incompatible preprocessing. Keep an old index available until the new one passes evaluation and traffic checks.

Index size is driven by document count, dimensions, numeric precision, metadata, and approximate-index overhead. Before choosing a database, test filtered recall and update behavior as well as raw nearest-neighbor speed. A fast vector index that applies permissions incorrectly or leaves stale content searchable is not a suitable knowledge base.

What this experiment does not prove

  • It does not prove vector search will beat BM25 on your content.
  • It does not compare vector databases or approximate indexes.
  • It does not test a neural cross-encoder reranker.
  • It does not test no-answer decisions, typos, multilingual retrieval, long-document chunking, or freshness.
  • It does not use customer queries, click-through behavior, or human satisfaction scores.
  • Its local one-pass timings are not production latency benchmarks.

The complete test fixture, settings, per-query rankings, reranker feature values, CSV results, and limitations are retained with the publication record. That transparency matters more than presenting one method as a universal winner.

Decision checklist

  • Do we have at least a small labeled set of real or explicitly synthetic queries?
  • Are failures mostly vocabulary mismatch, exact identifiers, ambiguous intent, missing content, or access filters?
  • Is the searchable unit self-contained and linked to a stable article destination?
  • Is the embedding model, revision, preprocessing, and distance metric recorded?
  • Can permissions, locale, product version, and publication status be enforced in retrieval?
  • Have BM25 and vector baselines been measured before hybrid tuning?
  • Does reranking improve a held-out set rather than only the tuning queries?
  • Are relevance, latency, index freshness, and no-answer behavior monitored after launch?

Frequently asked questions

Does vector search replace keyword search in a knowledge base?

Not automatically. Vector search is useful for paraphrases and conceptual similarity. Keyword search remains strong for exact names, codes, dates, and specialized terms. Test both on the same labeled queries before deciding whether to use one method, hybrid retrieval, or query routing.

Why did hybrid search perform worse in this experiment?

Equal-weight RRF gave BM25 and vector rankings the same influence even though vector retrieval was stronger on this query distribution. Fusion therefore pulled down some correct vector results. Hybrid search is a configuration to evaluate, not a guaranteed improvement.

What embedding model did the lab use?

We used Xenova’s ONNX conversion of all-MiniLM-L6-v2 at pinned repository revision 751bff37182d3f1213fa05d7196b954e230abad9, q8 weights, mean pooling, L2 normalization, and cosine similarity. This is one small English model, not a claim that it is best for every knowledge base.

How many queries are enough to evaluate search?

There is no universal minimum. A small set can reveal obvious failures, but reliable product decisions need enough examples across your important query types and languages. Report the count for every segment, preserve a held-out test set, and widen the set as new failures appear.

Should I add a reranker?

Add one when the first retrieval stage usually contains the right answer but orders candidates poorly, and when a held-out evaluation shows that the reranker improves the positions that matter. Our fixed feature reranker did not pass that test overall.

Final recommendation

Start with evidence, not architecture fashion. Build a BM25 baseline, test an appropriate embedding model on the same frozen labels, and analyze failures by query style. Add hybrid fusion only when the two rankers contribute complementary wins. Add reranking only when it improves held-out results.

In our original vector search for knowledge bases lab, vector-only retrieval was the best quality configuration and BM25 was the speed baseline. The hybrid and fixed reranking stages made the ranking worse. Publishing that result is the point of the experiment: the correct search stack is the smallest one your own evidence can justify.