Buying guide

Read chatbot performance claims without mistaking a demo for proof

Last materially reviewed 2026-09-25

Quick answerAsk what was measured, against which sources and under which conditions before using a vendor’s result in your buying decision.
Likely to work well when

✓ Public documentation maintainers

✓ Small software teams evaluating a chatbot

✓ Editors building a repeatable answer review

Important limitations

— Private account actions

— Guaranteed factual correctness

— Replacing source owners with a widget

What to know

Identify the object behind the number

A claim about accuracy, resolution or time saved needs a defined task, sample and measurement method. Ask whether the result came from a controlled test, selected customer examples or an estimate. A polished demo may demonstrate that one configured workflow can produce a useful answer. It does not establish how the same product will perform on your documentation, your version conflicts or questions that lack a supported answer.

What to know

Look for the missing denominator

A percentage without the number and type of cases leaves important uncertainty. Were difficult questions excluded? Did a human correct the result before it was counted? Was an unanswered question treated as a success because no support ticket followed? These questions do not imply misconduct. They identify what you need to know before comparing the claim with another vendor or with an ordinary documentation-search baseline.

What to know

Separate product facts from comparative judgments

An official source is appropriate for a documented feature or current plan limit, but a vendor-authored score for its own product is not an independent benchmark. Our SiteGPT comparison sources are labelled accordingly. We do not convert their ratings into our own quality scores. Verify the relevant capability directly in an authorized evaluation and keep the actual question, response and expected source passage available for review.

What to know

Use the claim to design a test

If a vendor emphasizes grounded answers, include a citation-relevance case. If it emphasizes freshness, include a changed-source case. Translate the promise into an observable task instead of repeating the headline. Keep your conclusion conditional on the scope and results you actually observed. Record untested promises as untested, and do not infer lower support costs or better sales from an answer that merely sounds convincing.

Source boundary

The evidence behind this buying guidance

This guide draws on SiteGPT vendor-authored Chatbase comparison; not an independent benchmark, LangChain: evaluation datasets and reference answers. Merchant-controlled records describe the provider’s own capabilities, terms or standards; they do not independently validate those claims. These records do not establish independent confirmation of the product claims.

Verify any current price, plan limit, label direction, compatibility rule, or commercial term that would materially change the decision. The dated source ledger shows the underlying records so this conclusion can be checked and updated.

Sources used for this page

These records support the facts and comparisons above. Merchant-controlled records are labelled so you can separate product claims from independent evidence.

  1. SiteGPT vendor-authored Chatbase comparison; not an independent benchmark — Merchant documentation · sitegpt.ai · Merchant-controlled · checked 2026-09-25
  2. LangChain: evaluation datasets and reference answers — Platform documentation · docs.langchain.com · Merchant-controlled · checked 2026-09-25