Independent research · Repeatable methods · Preserved evidence

Knowledge Base Software Research Lab

We run controlled experiments on how knowledge systems retrieve, migrate, present, and generate answers. Every study identifies its scope, test date, fixture, versions, scoring rules, observed results, preserved evidence, and limitations—so you can inspect the result before using the conclusion.

Independent scope: Knowledge Base Software is an editorial resource, not a software vendor. Vendor documentation may be a source, but vendor claims do not become measured findings. A laboratory result applies only to the named setup; it is not a universal product score.

Narrow questionsEach test has a defined decision, setup, and boundary.
Frozen inputsFixtures, labels, and scoring rules are fixed before the final run when possible.
Failures stay visibleNegative outputs and false positives are not edited into a success story.
Limits are part of the resultEvery study states what it did not establish.

These figures describe separate experiments. They must not be combined into one market benchmark.

Measured in the labRaw ZIP availableNegative result preserved

AI Answer Quality Testing: 600-Output Lab

A fixed 200-question English dataset tested five required behaviors—answer, clarify, abstain when evidence is missing, expose unresolved conflict, and refuse prohibited requests—across three seeded executions per case.

Observed result: the FLAN-T5-small q8 baseline achieved 0 strict passes out of 600. Version 1.1 corrected the deterministic scoring of the unchanged outputs to 25 category-rule matches; no model inference was rerun. The failure is the finding: fluent text did not satisfy the complete behavior contract.

Boundary: this is a failure analysis of one small local model and one prompt contract, not a benchmark of production AI systems.

200fixed questions written before the final run
3×seeded executions for every question
5answer-behavior categories
25/600deterministic category-rule matches in Version 1.1
0/600 strict passesRaw outputs, scores, settings, scripts, validation results, and checksums are preserved in the archive.

Original research

Search, portability, and accessibility evidence

These studies use different fixtures and answer different questions. Read each limitation before applying its result to your own stack.

Measured in the lab

Vector Search for Knowledge Bases

BM25, dense vectors, Reciprocal Rank Fusion, and a transparent fixed reranker were evaluated against the same 48-document support corpus and 72 pre-labeled queries.

Observed: vector-only retrieval achieved 90.3% Recall@1, compared with 72.2% for BM25. Both pre-set hybrid configurations ranked fewer labeled answers first than vector-only retrieval.

  • 48 documents
  • 72 labeled queries
  • 4 retrieval methods

Limit: one synthetic English corpus, one embedding model, and local timing—not a vendor or production-latency benchmark.

Observed hands-on

Five WordPress Plugins Tested

Five free plugins were installed on isolated WordPress Playground sites. Each received the same 20-article corpus, 14 queries, three runs per query, and export-delete-import checks where the free edition exposed them.

Observed: BasePress ranked the expected article first for 12 of 12 answerable queries. BetterDocs restored all 20 documents and four categories in its verified CSV round trip.

  • 5 plugins
  • 20 articles
  • 14 queries × 3

Limit: tested free editions, versions, synthetic corpus, and local WordPress Playground with SQLite.

Bounded structural audit

Knowledge Base Accessibility Audit

Six live templates were inspected against 14 repeatable structural checks, along with target-size candidates and horizontal-overflow proxies.

Observed: all six templates passed the measured structural checks at the tested viewport, while the FAQ and search templates produced the largest manual-review queues.

  • 6 templates
  • 14 checks
  • July 29 rerun

Limit: this is not a WCAG conformance claim; keyboard traversal, 200% zoom, and real assistive-technology testing still require humans.

More published experiments

Four additional studies, with three public artifact packages overall

These studies extend the lab into RAG structure, performance, search experience, and ROI sensitivity. Download labels appear only where a public package is actually available.

Controlled lab study

RAG Knowledge Base: Structure and Answer Quality

Compared the same facts in four long documents and 16 RAG-ready units across 64 questions, with retrieval evaluated separately from generation.

64 questions4 vs 16 content unitsEvidence retained
Read the RAG lab
Controlled lab study

Knowledge Base Performance Benchmarking

Used matched content implementations, frozen query labels, fixed BM25 settings, and 10,000 bootstrap resamples to separate useful differences from noise.

Matched implementations10,000 resamplesHashes published
Inspect the performance lab
Observed public auditRaw ZIP available

Knowledge Base Search Experience Audit

Audited four public search states and recorded response evidence, markup, accessible names, destinations, and checksums instead of reducing UX to one score.

100 result cards396 links108 destinations
Open the search UX audit
Deterministic modelRaw ZIP available

Knowledge Base ROI Sensitivity Lab

Ran 27 disclosed synthetic scenarios while separating cash outcomes from capacity effects and preserving the assumptions, result files, manifest, and checksums.

27 scenariosAssumptions CSVManifest included
Review the ROI laboratory

Public artifact packages are available from the AI answer-quality study, the search-experience audit, and the ROI sensitivity laboratory. The other studies disclose method and retained evidence but do not claim a public download.

Use the method with your content

Turn published evidence into your own acceptance test

A laboratory result is a starting hypothesis. Re-run the decisive checks with your articles, queries, language, permissions, plan, integrations, and failure costs before buying or releasing a system.

Browse all free tools
Retrieval

Search Relevance Benchmark

Score top-five results with labeled queries, no-answer controls, Success@1, MRR@5, and nDCG@5.

Benchmark your search
AI answers

AI Answer Grounding Test Kit

Map atomic claims to approved passages and check citations, coverage, conflict handling, and safe abstention.

Audit answer grounding
Content quality

Article Quality & Accessibility Checker

Review structure, findability, instructions, evidence, accessibility signals, and governance metadata.

Check an article
Portability

Export & Recovery Validator

Inspect a real export for content, hierarchy, metadata, attachments, redirects, and recovery evidence.

Validate an export

Research method

What makes a result usable

A percentage without a boundary is difficult to use. Our method shows what was tested, how the result was produced, and which decisions the evidence can and cannot support.

  1. Start with one decisionRetrieval relevance, migration integrity, answer behavior, and accessibility require different tests.
  2. Freeze representative inputsDocuments, queries, expected behavior, distractors, and scoring rules are fixed before the final run when possible.
  3. Record the exact setupDate, version, plan or edition, environment, settings, language, corpus, and timing boundary stay with the result.
  4. Measure observable behaviorRanks, restored records, supported claims, false positives, and raw outputs take priority over impressions.
  5. Preserve failures and evidenceNegative results remain visible; every study states which files are downloadable and which evidence is retained.
  6. Publish limits and correctionsA rerun records the changed date, version, fixture, and method instead of silently changing the old boundary.

Next research protocol

Knowledge Base Software Benchmark 2026

The next study will compare a small platform set under the same synthetic organization, content, roles, tasks, queries, and scoring rules. This section describes a protocol in development—not published results.

  • Search and AI answer quality
  • Permissions and public/private boundaries
  • Migration, export, and recovery
  • Mobile and accessibility screening
  • Exact plan, version, region, date, screenshots, and raw result files

Research FAQ

How to use these findings

What is the Knowledge Base Software Research Lab?

It is a collection of original experiments about retrieval, AI answer behavior, portability, migration, accessibility, and other measurable parts of operating or evaluating a knowledge base. It is separate from the category guide and vendor-comparison pages.

Are these studies rankings of the best knowledge base software?

No. Each study answers a narrow question under a named setup. A product or method can perform well on one test and still fail another requirement. Use the finding to design your own trial rather than turning it into one universal score.

Can these percentages be used as industry benchmarks?

No. A percentage applies only to the stated corpus, queries, versions, settings, language, environment, and scoring rules. It can provide a reproducible baseline, but it does not describe the entire market or predict performance with your content.

Can readers reproduce the experiments?

Reproducibility varies by study. The AI answer-quality study provides a downloadable archive containing its dataset, outputs, scripts, settings, and checksums. Other studies disclose their setup, measurements, retained evidence, and limits; we do not describe evidence as downloadable when no public file exists.

What does “tested” mean on this site?

It means a real experiment or task was performed and its setup, date, inputs, measurements, result, and limitations are disclosed. Reading documentation or a pricing page is labeled vendor-documented research, not hands-on testing.

How are commercial relationships handled?

Advertising, affiliate relationships, and vendor contact do not determine which experiment is run, how evidence is scored, or whether a negative result is published. Relevant relationships are disclosed, and vendors or readers may report factual errors through the public corrections process.

Use evidence at the right stage

Move from a market claim to a testable buying decision

Use the Research Lab when you need methods, measurements, failure cases, or reusable test fixtures. If you are still defining the category or building a shortlist, start with the main knowledge base software guide, then convert the relevant findings into requirements and trial gates.

Research Hub v1.0 · Updated August 10, 2026 · Results apply only to their stated scope, environment, versions, and scoring rules.