Independent research · Repeatable methods · Preserved evidence
Knowledge Base Software Research Lab
We run controlled experiments on how knowledge systems retrieve, migrate, present, and generate answers. Every study identifies its scope, test date, fixture, versions, scoring rules, observed results, preserved evidence, and limitations—so you can inspect the result before using the conclusion.
Independent scope: Knowledge Base Software is an editorial resource, not a software vendor. Vendor documentation may be a source, but vendor claims do not become measured findings. A laboratory result applies only to the named setup; it is not a universal product score.
These figures describe separate experiments. They must not be combined into one market benchmark.
Start with the question
Choose the evidence you need
Open the study that matches the decision in front of you. Each result remains attached to its fixture, version, language, and test boundary.
Can an AI answer safely from approved evidence?
Inspect answer, clarification, abstention, conflict, and refusal behavior across fixed cases.
Open the AI lab 02Which retrieval method ranks the answer first?
Compare BM25, vectors, hybrid fusion, and a transparent reranker on one frozen corpus.
Open the search lab 03Can a plugin find and preserve the content?
Review matched search and export-import tasks across five free WordPress plugins.
Open the plugin lab 04Does the interface expose a sound structural baseline?
See repeatable accessibility checks and the manual tests that automation cannot replace.
Open the accessibility auditAI Answer Quality Testing: 600-Output Lab
A fixed 200-question English dataset tested five required behaviors—answer, clarify, abstain when evidence is missing, expose unresolved conflict, and refuse prohibited requests—across three seeded executions per case.
Observed result: the FLAN-T5-small q8 baseline achieved 0 strict passes out of 600. Version 1.1 corrected the deterministic scoring of the unchanged outputs to 25 category-rule matches; no model inference was rerun. The failure is the finding: fluent text did not satisfy the complete behavior contract.
Boundary: this is a failure analysis of one small local model and one prompt contract, not a benchmark of production AI systems.
Original research
Search, portability, and accessibility evidence
These studies use different fixtures and answer different questions. Read each limitation before applying its result to your own stack.
Vector Search for Knowledge Bases
BM25, dense vectors, Reciprocal Rank Fusion, and a transparent fixed reranker were evaluated against the same 48-document support corpus and 72 pre-labeled queries.
Observed: vector-only retrieval achieved 90.3% Recall@1, compared with 72.2% for BM25. Both pre-set hybrid configurations ranked fewer labeled answers first than vector-only retrieval.
- 48 documents
- 72 labeled queries
- 4 retrieval methods
Limit: one synthetic English corpus, one embedding model, and local timing—not a vendor or production-latency benchmark.
Five WordPress Plugins Tested
Five free plugins were installed on isolated WordPress Playground sites. Each received the same 20-article corpus, 14 queries, three runs per query, and export-delete-import checks where the free edition exposed them.
Observed: BasePress ranked the expected article first for 12 of 12 answerable queries. BetterDocs restored all 20 documents and four categories in its verified CSV round trip.
- 5 plugins
- 20 articles
- 14 queries × 3
Limit: tested free editions, versions, synthetic corpus, and local WordPress Playground with SQLite.
Knowledge Base Accessibility Audit
Six live templates were inspected against 14 repeatable structural checks, along with target-size candidates and horizontal-overflow proxies.
Observed: all six templates passed the measured structural checks at the tested viewport, while the FAQ and search templates produced the largest manual-review queues.
- 6 templates
- 14 checks
- July 29 rerun
Limit: this is not a WCAG conformance claim; keyboard traversal, 200% zoom, and real assistive-technology testing still require humans.
More published experiments
Four additional studies, with three public artifact packages overall
These studies extend the lab into RAG structure, performance, search experience, and ROI sensitivity. Download labels appear only where a public package is actually available.
RAG Knowledge Base: Structure and Answer Quality
Compared the same facts in four long documents and 16 RAG-ready units across 64 questions, with retrieval evaluated separately from generation.
Knowledge Base Performance Benchmarking
Used matched content implementations, frozen query labels, fixed BM25 settings, and 10,000 bootstrap resamples to separate useful differences from noise.
Knowledge Base Search Experience Audit
Audited four public search states and recorded response evidence, markup, accessible names, destinations, and checksums instead of reducing UX to one score.
Knowledge Base ROI Sensitivity Lab
Ran 27 disclosed synthetic scenarios while separating cash outcomes from capacity effects and preserving the assumptions, result files, manifest, and checksums.
Public artifact packages are available from the AI answer-quality study, the search-experience audit, and the ROI sensitivity laboratory. The other studies disclose method and retained evidence but do not claim a public download.
Use the method with your content
Turn published evidence into your own acceptance test
A laboratory result is a starting hypothesis. Re-run the decisive checks with your articles, queries, language, permissions, plan, integrations, and failure costs before buying or releasing a system.
Search Relevance Benchmark
Score top-five results with labeled queries, no-answer controls, Success@1, MRR@5, and nDCG@5.
Benchmark your searchAI Answer Grounding Test Kit
Map atomic claims to approved passages and check citations, coverage, conflict handling, and safe abstention.
Audit answer groundingArticle Quality & Accessibility Checker
Review structure, findability, instructions, evidence, accessibility signals, and governance metadata.
Check an articleExport & Recovery Validator
Inspect a real export for content, hierarchy, metadata, attachments, redirects, and recovery evidence.
Validate an export25-Task Trial Test Plan
Run the same authoring, search, AI, permissions, migration, and export tasks against every shortlisted product.
Download the 25-task knowledge base software trial test planResearch method
What makes a result usable
A percentage without a boundary is difficult to use. Our method shows what was tested, how the result was produced, and which decisions the evidence can and cannot support.
- Start with one decisionRetrieval relevance, migration integrity, answer behavior, and accessibility require different tests.
- Freeze representative inputsDocuments, queries, expected behavior, distractors, and scoring rules are fixed before the final run when possible.
- Record the exact setupDate, version, plan or edition, environment, settings, language, corpus, and timing boundary stay with the result.
- Measure observable behaviorRanks, restored records, supported claims, false positives, and raw outputs take priority over impressions.
- Preserve failures and evidenceNegative results remain visible; every study states which files are downloadable and which evidence is retained.
- Publish limits and correctionsA rerun records the changed date, version, fixture, and method instead of silently changing the old boundary.
Next research protocol
Knowledge Base Software Benchmark 2026
The next study will compare a small platform set under the same synthetic organization, content, roles, tasks, queries, and scoring rules. This section describes a protocol in development—not published results.
- Search and AI answer quality
- Permissions and public/private boundaries
- Migration, export, and recovery
- Mobile and accessibility screening
- Exact plan, version, region, date, screenshots, and raw result files
Research FAQ
How to use these findings
What is the Knowledge Base Software Research Lab?
It is a collection of original experiments about retrieval, AI answer behavior, portability, migration, accessibility, and other measurable parts of operating or evaluating a knowledge base. It is separate from the category guide and vendor-comparison pages.
Are these studies rankings of the best knowledge base software?
No. Each study answers a narrow question under a named setup. A product or method can perform well on one test and still fail another requirement. Use the finding to design your own trial rather than turning it into one universal score.
Can these percentages be used as industry benchmarks?
No. A percentage applies only to the stated corpus, queries, versions, settings, language, environment, and scoring rules. It can provide a reproducible baseline, but it does not describe the entire market or predict performance with your content.
Can readers reproduce the experiments?
Reproducibility varies by study. The AI answer-quality study provides a downloadable archive containing its dataset, outputs, scripts, settings, and checksums. Other studies disclose their setup, measurements, retained evidence, and limits; we do not describe evidence as downloadable when no public file exists.
What does “tested” mean on this site?
It means a real experiment or task was performed and its setup, date, inputs, measurements, result, and limitations are disclosed. Reading documentation or a pricing page is labeled vendor-documented research, not hands-on testing.
How are commercial relationships handled?
Advertising, affiliate relationships, and vendor contact do not determine which experiment is run, how evidence is scored, or whether a negative result is published. Relevant relationships are disclosed, and vendors or readers may report factual errors through the public corrections process.
Use evidence at the right stage
Move from a market claim to a testable buying decision
Use the Research Lab when you need methods, measurements, failure cases, or reusable test fixtures. If you are still defining the category or building a shortlist, start with the main knowledge base software guide, then convert the relevant findings into requirements and trial gates.
