RAG Knowledge Base: Original Lab & Implementation Guide

Original laboratory study — July 28, 2026: We tested two versions of the same support knowledge: four long, conventionally written source documents and 16 short documents prepared specifically for retrieval-augmented generation (RAG). Both conditions used the same facts, questions, BM25 retriever, local generator, and abstention rule. The complete method and all 64 results are reported below, including a negative result: better retrieval did not improve answer accuracy with this small baseline model.

A RAG knowledge base is a collection of source material prepared so a retrieval system can find the right evidence and give it to a language model before the model answers. The practical goal is not to make documents “AI-friendly” in the abstract. It is to make the evidence unit findable, understandable outside its original page, attributable, current, and safe to use.

This guide explains the architecture, content design, evaluation process, failure modes, and governance controls behind a production RAG knowledge base. It also publishes our own controlled comparison of raw documentation versus RAG-ready content rather than repeating unverified benchmark numbers.

RAG knowledge base: the short definition

Retrieval-augmented generation combines two forms of memory: knowledge stored in a model’s parameters and external evidence retrieved at query time. The original RAG paper describes this as combining parametric and non-parametric memory for generation. In a support or documentation system, the external memory is commonly an index of approved articles, policies, procedures, and product facts.

  • Ingestion collects and cleans approved sources.
  • Chunking turns sources into retrievable evidence units.
  • Indexing stores terms, vectors, metadata, or a combination.
  • Retrieval selects candidate passages for a user question.
  • Generation produces an answer constrained by the supplied evidence.
  • Guardrails decide when to answer, qualify, cite, or abstain.
  • Evaluation measures retrieval and answer quality separately.

The foundational research is available in Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”. It is useful background, but a production knowledge base also needs permissions, freshness, traceability, and explicit refusal behavior that are specific to the organization using it.

Our original lab: raw documentation versus RAG-ready content

We built a reproducible local experiment to isolate one question: does rewriting the same facts as smaller, self-contained knowledge units change retrieval and answer performance? This is a content-structure test, not a vendor comparison and not a benchmark of modern production language models.

Laboratory results comparing raw and RAG-ready knowledge base content across retrieval, prompt size, answer accuracy, abstention, and generation latency
Raw versus RAG-ready content in the Knowledge Base Software laboratory, July 28, 2026. Higher is better for retrieval and answer measures; lower is better for prompt size and latency.

Experimental design

  • Corpus: 16 synthetic but realistic support topics covering accounts, billing, security, exports, integrations, and administration.
  • Raw condition: four long documents containing all 16 topics.
  • RAG-ready condition: 16 short, self-contained documents containing the same facts, with descriptive titles and no dependency on surrounding sections.
  • Question set: 64 questions—48 answerable questions and 16 questions for which the corpus contained no answer.
  • Retriever: the same deterministic BM25 implementation in both conditions.
  • Generator: the same locally executed FLAN-T5-base quantized model in both conditions, with sampling disabled.
  • Context: only the top-ranked item was passed to the generator.
  • Abstention rule: return no answer when query-term coverage was below 0.5.
  • Environment: one Windows workstation; all retrieval and generation ran locally after model download.

Every answer was graded against predefined acceptable fact strings. Retrieval was scored independently from generation so that a correct passage with an incorrect answer was not counted as an end-to-end success. Timing values are local laboratory timings, not cloud service-level claims.

Measured results

MetricRaw documentsRAG-ready documentsObserved change
Documents indexed416Finer evidence units
Hit@3, answerable questions95.8%100.0%+4.2 percentage points
Correct item ranked first85.4%100.0%+14.6 percentage points
Answer accuracy37.5%35.4%−2.1 percentage points
Correct abstention, no-answer questions100.0%100.0%No change
Mean prompt length1,729.6 characters406.5 characters76.5% shorter
Generation latency, median4,402 ms1,697 ms61.5% lower
Generation latency, p959,997 ms5,089 ms49.1% lower
Retrieval latency, p950.58 ms0.38 ms0.20 ms lower
Results from 64 fixed questions. Percentages are calculated from 48 answerable or 16 no-answer questions as appropriate. Values are rounded for readability.

What the experiment shows—and what it does not

Preparing content for RAG solved the retrieval problem in this corpus: the correct evidence ranked first for every answerable question. It also reduced the average context passed to the model by about three quarters, which was associated with substantially lower local generation latency.

However, answer accuracy did not improve. It decreased slightly from 37.5% to 35.4%. Inspection of the outputs showed that the small baseline generator could still omit, distort, or poorly format a fact even when the correct evidence was supplied. This is the most important result in the study: retrieval quality and answer quality are different layers, and improving one does not guarantee the other.

The 100% abstention result applies only to the fixed no-answer set and the stated lexical coverage rule. It should not be generalized to unseen questions. Likewise, this experiment does not compare embeddings, rerankers, commercial systems, or current frontier models. For retrieval-method comparisons, see our vector search for knowledge bases laboratory guide. For a broader test set that separates correct answers, ambiguity, conflict, missing evidence, and refusal, see AI answer quality testing for knowledge bases.

How to structure content for a RAG knowledge base

1. Make each evidence unit self-contained

A retriever may return a section without the paragraphs before it. Replace context-dependent headings such as “Exceptions” or “What happens next?” with headings that name the subject. Repeat the product, plan, role, region, or policy name when it is required to interpret the rule.

Weak chunk: “It is retained for seven years.”

Self-contained chunk: “Archived billing invoices are retained for seven years after the invoice date.”

2. Split by meaning before splitting by size

Keep a decision, procedure, limitation, or policy together when possible. Microsoft’s official Azure AI Search chunking guidance distinguishes fixed-size, structure-aware, semantic, and combined approaches, and recommends treating starting sizes as parameters to test rather than universal rules. Headings and document structure provide useful boundaries; overlap can preserve context but also duplicates text and increases index size.

3. Put the answer near the descriptive heading

Lead with the direct answer, then add prerequisites, steps, examples, and exceptions. A short answer separated from its subject is fragile; a long introduction before the operative fact wastes retrieved context. For procedures, use one ordered sequence and state the expected result.

4. Store metadata that can change retrieval

  • stable document and chunk identifiers;
  • canonical source URL;
  • product, feature, audience, plan, region, and language;
  • owner and review date;
  • effective date and version;
  • access-control labels;
  • superseded or archived status.

Metadata is not decoration. It lets the system exclude expired policies, respect permissions, filter to the correct product version, and return a source a reviewer can inspect.

5. Write explicit exceptions and conflicts

Do not hide exceptions in notes that can be separated from the rule. State when a procedure does not apply and what source takes precedence. If two policies legitimately differ by plan or region, encode the scope in both the content and metadata. If the sources actually conflict, the answer layer should surface the conflict or abstain—not silently choose one.

6. Preserve citations and provenance

Every answerable unit should map back to an approved source. Store the URL, title, section, version, and effective date needed to reproduce the evidence. Citations help users verify answers and help maintainers diagnose whether a failure came from stale content, retrieval, or generation.

A practical RAG knowledge base architecture

LayerPrimary responsibilityFailure to test
Source governanceApproved ownership, versions, permissions, retentionStale, unauthorized, or contradictory source
Parsing and chunkingPreserve meaning while creating retrievable unitsLost headings, tables, lists, or exception scope
IndexRepresent text, vectors, and metadataMissing fields, duplicate versions, filter leakage
RetrieverReturn relevant candidatesCorrect evidence absent from top-k
RerankerImprove candidate orderRelevant passage demoted
Prompt assemblyProvide sufficient, bounded evidenceTruncation, injection, unrelated context
GeneratorAnswer from evidence in the requested formUnsupported claim, omission, calculation error
Policy layerAnswer, qualify, cite, escalate, or refuseAnswering when evidence is absent or restricted
Evaluation and monitoringDetect regressions by question classAggregate score hides dangerous failures

How to evaluate a RAG knowledge base

Do not reduce evaluation to one “accuracy” number. Use a fixed, versioned question set and report each layer separately.

  1. Retrieval coverage: Is an acceptable evidence unit present in top-k?
  2. Ranking: How often is the best unit ranked first, and what is its reciprocal rank?
  3. Grounded correctness: Does the answer match the approved fact and remain within the evidence?
  4. Citation correctness: Does the cited source actually support the claim?
  5. Completeness: Are required steps, conditions, and exceptions present?
  6. Abstention: Does the system refuse or escalate when evidence is missing, conflicting, or restricted?
  7. Operational cost: Track context size, retrieval latency, generation latency, and model usage.
  8. Slice results: Report by product, language, document type, query class, and risk level.

For risk management, document the intended use, foreseeable misuse, measurement limitations, review process, and incident response. The NIST Generative AI Profile is a primary reference for integrating generative-AI risks into an organization’s broader risk-management program.

Common RAG failure modes and fixes

Observed failureLikely layerFirst diagnostic
Correct source never appearsIngestion, chunking, or retrievalConfirm the fact was indexed and inspect query–chunk terms and metadata filters
Correct source appears below noiseRankingCompare lexical, vector, hybrid, and reranked order on the same candidates
Correct passage is first but answer is wrongGeneration or promptTest an extractive answer, stricter prompt, and stronger model before rewriting the index
Answer mixes two product versionsMetadata and governanceFilter by version and remove superseded chunks
System answers an unanswerable questionPolicy and abstentionAdd no-answer and conflict cases; tune threshold on a held-out set
Citation does not support the sentencePrompt, citation mapping, or generationValidate claims against cited spans, not only document URLs
Restricted content leaksAuthorizationApply permission filtering before retrieval and test with identities, not only documents

Implementation checklist

  • Define the answerable scope and risk levels before indexing.
  • Inventory canonical sources, owners, permissions, and update paths.
  • Remove navigation, repeated boilerplate, and obsolete versions during ingestion.
  • Build self-contained chunks around meaningful sections.
  • Preserve titles, hierarchy, lists, tables, dates, and source URLs.
  • Choose lexical, vector, or hybrid retrieval using a labeled question set.
  • Add reranking only after measuring the baseline.
  • Require evidence-aware citations and test whether they support each claim.
  • Design explicit behavior for missing, conflicting, and restricted information.
  • Version the corpus, prompts, models, thresholds, questions, and scoring rules.
  • Review failures by category; never optimize only the aggregate score.
  • Re-run the suite when content, index settings, retrieval, model, or policy changes.

Frequently asked questions

Does a RAG knowledge base prevent hallucinations?

No. Retrieval can supply evidence, but the model can still misread it, combine it incorrectly, omit a condition, or answer when no adequate evidence exists. Test retrieval, generation, citations, and abstention separately.

What is the best chunk size for RAG?

There is no universal best size. Start with boundaries that preserve a complete meaning unit, then test chunk size and overlap on your own questions. The right choice depends on document structure, query type, embedding limits, retriever, and generator context.

Should a RAG knowledge base use vector search?

Not automatically. Lexical search can be strong for exact product terms, error codes, and policy wording. Vector search helps with paraphrases and semantic similarity. Hybrid retrieval often deserves testing, but the decision should come from a labeled evaluation rather than a default architecture diagram.

What did this laboratory experiment find?

RAG-ready content raised top-1 retrieval from 85.4% to 100%, reduced mean prompt length by 76.5%, and reduced median local generation latency by 61.5%. It did not improve answer accuracy with the small FLAN-T5-base baseline: accuracy changed from 37.5% to 35.4%.

Bottom line

A reliable RAG knowledge base is an evidence system, not a pile of articles attached to a chatbot. Structure each unit so it survives retrieval, preserve scope and provenance, enforce permissions before retrieval, and test refusal alongside correctness. Our lab found that better content structure materially improved retrieval efficiency—but also demonstrated why a retrieval win must never be reported as an answer-quality win without measuring the generator.


Laboratory transparency: Test date: July 28, 2026. Corpus: 16 synthetic support topics. Questions: 64. Retriever: deterministic BM25. Generator: FLAN-T5-base, quantized, local, deterministic. Raw outputs and scoring files were retained by Knowledge Base Software. This study is a scoped comparison of two content structures and should not be interpreted as a vendor benchmark or a claim about modern production LLMs.