
RAG Knowledge Base: Original Lab & Implementation Guide
Original laboratory study — July 28, 2026: We tested two versions of the same support knowledge: four long, conventionally written source documents and 16 short documents prepared specifically for retrieval-augmented generation (RAG). Both conditions used the same facts, questions, BM25 retriever, local generator, and abstention rule. The complete method and all 64 results are reported below, including a negative result: better retrieval did not improve answer accuracy with this small baseline model.
A RAG knowledge base is a collection of source material prepared so a retrieval system can find the right evidence and give it to a language model before the model answers. The practical goal is not to make documents “AI-friendly” in the abstract. It is to make the evidence unit findable, understandable outside its original page, attributable, current, and safe to use.
This guide explains the architecture, content design, evaluation process, failure modes, and governance controls behind a production RAG knowledge base. It also publishes our own controlled comparison of raw documentation versus RAG-ready content rather than repeating unverified benchmark numbers.
RAG knowledge base: the short definition
Retrieval-augmented generation combines two forms of memory: knowledge stored in a model’s parameters and external evidence retrieved at query time. The original RAG paper describes this as combining parametric and non-parametric memory for generation. In a support or documentation system, the external memory is commonly an index of approved articles, policies, procedures, and product facts.
- Ingestion collects and cleans approved sources.
- Chunking turns sources into retrievable evidence units.
- Indexing stores terms, vectors, metadata, or a combination.
- Retrieval selects candidate passages for a user question.
- Generation produces an answer constrained by the supplied evidence.
- Guardrails decide when to answer, qualify, cite, or abstain.
- Evaluation measures retrieval and answer quality separately.
The foundational research is available in Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”. It is useful background, but a production knowledge base also needs permissions, freshness, traceability, and explicit refusal behavior that are specific to the organization using it.
Our original lab: raw documentation versus RAG-ready content
We built a reproducible local experiment to isolate one question: does rewriting the same facts as smaller, self-contained knowledge units change retrieval and answer performance? This is a content-structure test, not a vendor comparison and not a benchmark of modern production language models.

Experimental design
- Corpus: 16 synthetic but realistic support topics covering accounts, billing, security, exports, integrations, and administration.
- Raw condition: four long documents containing all 16 topics.
- RAG-ready condition: 16 short, self-contained documents containing the same facts, with descriptive titles and no dependency on surrounding sections.
- Question set: 64 questions—48 answerable questions and 16 questions for which the corpus contained no answer.
- Retriever: the same deterministic BM25 implementation in both conditions.
- Generator: the same locally executed FLAN-T5-base quantized model in both conditions, with sampling disabled.
- Context: only the top-ranked item was passed to the generator.
- Abstention rule: return no answer when query-term coverage was below 0.5.
- Environment: one Windows workstation; all retrieval and generation ran locally after model download.
Every answer was graded against predefined acceptable fact strings. Retrieval was scored independently from generation so that a correct passage with an incorrect answer was not counted as an end-to-end success. Timing values are local laboratory timings, not cloud service-level claims.
Measured results
| Metric | Raw documents | RAG-ready documents | Observed change |
|---|---|---|---|
| Documents indexed | 4 | 16 | Finer evidence units |
| Hit@3, answerable questions | 95.8% | 100.0% | +4.2 percentage points |
| Correct item ranked first | 85.4% | 100.0% | +14.6 percentage points |
| Answer accuracy | 37.5% | 35.4% | −2.1 percentage points |
| Correct abstention, no-answer questions | 100.0% | 100.0% | No change |
| Mean prompt length | 1,729.6 characters | 406.5 characters | 76.5% shorter |
| Generation latency, median | 4,402 ms | 1,697 ms | 61.5% lower |
| Generation latency, p95 | 9,997 ms | 5,089 ms | 49.1% lower |
| Retrieval latency, p95 | 0.58 ms | 0.38 ms | 0.20 ms lower |
What the experiment shows—and what it does not
Preparing content for RAG solved the retrieval problem in this corpus: the correct evidence ranked first for every answerable question. It also reduced the average context passed to the model by about three quarters, which was associated with substantially lower local generation latency.
However, answer accuracy did not improve. It decreased slightly from 37.5% to 35.4%. Inspection of the outputs showed that the small baseline generator could still omit, distort, or poorly format a fact even when the correct evidence was supplied. This is the most important result in the study: retrieval quality and answer quality are different layers, and improving one does not guarantee the other.
The 100% abstention result applies only to the fixed no-answer set and the stated lexical coverage rule. It should not be generalized to unseen questions. Likewise, this experiment does not compare embeddings, rerankers, commercial systems, or current frontier models. For retrieval-method comparisons, see our vector search for knowledge bases laboratory guide. For a broader test set that separates correct answers, ambiguity, conflict, missing evidence, and refusal, see AI answer quality testing for knowledge bases.
How to structure content for a RAG knowledge base
1. Make each evidence unit self-contained
A retriever may return a section without the paragraphs before it. Replace context-dependent headings such as “Exceptions” or “What happens next?” with headings that name the subject. Repeat the product, plan, role, region, or policy name when it is required to interpret the rule.
Weak chunk: “It is retained for seven years.”
Self-contained chunk: “Archived billing invoices are retained for seven years after the invoice date.”
2. Split by meaning before splitting by size
Keep a decision, procedure, limitation, or policy together when possible. Microsoft’s official Azure AI Search chunking guidance distinguishes fixed-size, structure-aware, semantic, and combined approaches, and recommends treating starting sizes as parameters to test rather than universal rules. Headings and document structure provide useful boundaries; overlap can preserve context but also duplicates text and increases index size.
3. Put the answer near the descriptive heading
Lead with the direct answer, then add prerequisites, steps, examples, and exceptions. A short answer separated from its subject is fragile; a long introduction before the operative fact wastes retrieved context. For procedures, use one ordered sequence and state the expected result.
4. Store metadata that can change retrieval
- stable document and chunk identifiers;
- canonical source URL;
- product, feature, audience, plan, region, and language;
- owner and review date;
- effective date and version;
- access-control labels;
- superseded or archived status.
Metadata is not decoration. It lets the system exclude expired policies, respect permissions, filter to the correct product version, and return a source a reviewer can inspect.
5. Write explicit exceptions and conflicts
Do not hide exceptions in notes that can be separated from the rule. State when a procedure does not apply and what source takes precedence. If two policies legitimately differ by plan or region, encode the scope in both the content and metadata. If the sources actually conflict, the answer layer should surface the conflict or abstain—not silently choose one.
6. Preserve citations and provenance
Every answerable unit should map back to an approved source. Store the URL, title, section, version, and effective date needed to reproduce the evidence. Citations help users verify answers and help maintainers diagnose whether a failure came from stale content, retrieval, or generation.
A practical RAG knowledge base architecture
| Layer | Primary responsibility | Failure to test |
|---|---|---|
| Source governance | Approved ownership, versions, permissions, retention | Stale, unauthorized, or contradictory source |
| Parsing and chunking | Preserve meaning while creating retrievable units | Lost headings, tables, lists, or exception scope |
| Index | Represent text, vectors, and metadata | Missing fields, duplicate versions, filter leakage |
| Retriever | Return relevant candidates | Correct evidence absent from top-k |
| Reranker | Improve candidate order | Relevant passage demoted |
| Prompt assembly | Provide sufficient, bounded evidence | Truncation, injection, unrelated context |
| Generator | Answer from evidence in the requested form | Unsupported claim, omission, calculation error |
| Policy layer | Answer, qualify, cite, escalate, or refuse | Answering when evidence is absent or restricted |
| Evaluation and monitoring | Detect regressions by question class | Aggregate score hides dangerous failures |
How to evaluate a RAG knowledge base
Do not reduce evaluation to one “accuracy” number. Use a fixed, versioned question set and report each layer separately.
- Retrieval coverage: Is an acceptable evidence unit present in top-k?
- Ranking: How often is the best unit ranked first, and what is its reciprocal rank?
- Grounded correctness: Does the answer match the approved fact and remain within the evidence?
- Citation correctness: Does the cited source actually support the claim?
- Completeness: Are required steps, conditions, and exceptions present?
- Abstention: Does the system refuse or escalate when evidence is missing, conflicting, or restricted?
- Operational cost: Track context size, retrieval latency, generation latency, and model usage.
- Slice results: Report by product, language, document type, query class, and risk level.
For risk management, document the intended use, foreseeable misuse, measurement limitations, review process, and incident response. The NIST Generative AI Profile is a primary reference for integrating generative-AI risks into an organization’s broader risk-management program.
Common RAG failure modes and fixes
| Observed failure | Likely layer | First diagnostic |
|---|---|---|
| Correct source never appears | Ingestion, chunking, or retrieval | Confirm the fact was indexed and inspect query–chunk terms and metadata filters |
| Correct source appears below noise | Ranking | Compare lexical, vector, hybrid, and reranked order on the same candidates |
| Correct passage is first but answer is wrong | Generation or prompt | Test an extractive answer, stricter prompt, and stronger model before rewriting the index |
| Answer mixes two product versions | Metadata and governance | Filter by version and remove superseded chunks |
| System answers an unanswerable question | Policy and abstention | Add no-answer and conflict cases; tune threshold on a held-out set |
| Citation does not support the sentence | Prompt, citation mapping, or generation | Validate claims against cited spans, not only document URLs |
| Restricted content leaks | Authorization | Apply permission filtering before retrieval and test with identities, not only documents |
Implementation checklist
- Define the answerable scope and risk levels before indexing.
- Inventory canonical sources, owners, permissions, and update paths.
- Remove navigation, repeated boilerplate, and obsolete versions during ingestion.
- Build self-contained chunks around meaningful sections.
- Preserve titles, hierarchy, lists, tables, dates, and source URLs.
- Choose lexical, vector, or hybrid retrieval using a labeled question set.
- Add reranking only after measuring the baseline.
- Require evidence-aware citations and test whether they support each claim.
- Design explicit behavior for missing, conflicting, and restricted information.
- Version the corpus, prompts, models, thresholds, questions, and scoring rules.
- Review failures by category; never optimize only the aggregate score.
- Re-run the suite when content, index settings, retrieval, model, or policy changes.
Frequently asked questions
Does a RAG knowledge base prevent hallucinations?
No. Retrieval can supply evidence, but the model can still misread it, combine it incorrectly, omit a condition, or answer when no adequate evidence exists. Test retrieval, generation, citations, and abstention separately.
What is the best chunk size for RAG?
There is no universal best size. Start with boundaries that preserve a complete meaning unit, then test chunk size and overlap on your own questions. The right choice depends on document structure, query type, embedding limits, retriever, and generator context.
Should a RAG knowledge base use vector search?
Not automatically. Lexical search can be strong for exact product terms, error codes, and policy wording. Vector search helps with paraphrases and semantic similarity. Hybrid retrieval often deserves testing, but the decision should come from a labeled evaluation rather than a default architecture diagram.
What did this laboratory experiment find?
RAG-ready content raised top-1 retrieval from 85.4% to 100%, reduced mean prompt length by 76.5%, and reduced median local generation latency by 61.5%. It did not improve answer accuracy with the small FLAN-T5-base baseline: accuracy changed from 37.5% to 35.4%.
Bottom line
A reliable RAG knowledge base is an evidence system, not a pile of articles attached to a chatbot. Structure each unit so it survives retrieval, preserve scope and provenance, enforce permissions before retrieval, and test refusal alongside correctness. Our lab found that better content structure materially improved retrieval efficiency—but also demonstrated why a retrieval win must never be reported as an answer-quality win without measuring the generator.
Laboratory transparency: Test date: July 28, 2026. Corpus: 16 synthetic support topics. Questions: 64. Retriever: deterministic BM25. Generator: FLAN-T5-base, quantized, local, deterministic. Raw outputs and scoring files were retained by Knowledge Base Software. This study is a scoped comparison of two content structures and should not be interpreted as a vendor benchmark or a claim about modern production LLMs.



