✓ Public documentation maintainers
✓ Small software teams evaluating a chatbot
✓ Editors building a repeatable answer review
— Private account actions
— Guaranteed factual correctness
— Replacing source owners with a widget
Identify the object behind the number
A claim about accuracy, resolution or time saved needs a defined task, sample and measurement method. Ask whether the result came from a controlled test, selected customer examples or an estimate. A polished demo may demonstrate that one configured workflow can produce a useful answer. It does not establish how the same product will perform on your documentation, your version conflicts or questions that lack a supported answer.
Look for the missing denominator
A percentage without the number and type of cases leaves important uncertainty. Were difficult questions excluded? Did a human correct the result before it was counted? Was an unanswered question treated as a success because no support ticket followed? These questions do not imply misconduct. They identify what you need to know before comparing the claim with another vendor or with an ordinary documentation-search baseline.
Separate product facts from comparative judgments
An official source is appropriate for a documented feature or current plan limit, but a vendor-authored score for its own product is not an independent benchmark. Our SiteGPT comparison sources are labelled accordingly. We do not convert their ratings into our own quality scores. Verify the relevant capability directly in an authorized evaluation and keep the actual question, response and expected source passage available for review.
Use the claim to design a test
If a vendor emphasizes grounded answers, include a citation-relevance case. If it emphasizes freshness, include a changed-source case. Translate the promise into an observable task instead of repeating the headline. Keep your conclusion conditional on the scope and results you actually observed. Record untested promises as untested, and do not infer lower support costs or better sales from an answer that merely sounds convincing.
The evidence behind this buying guidance
This guide draws on SiteGPT vendor-authored Chatbase comparison; not an independent benchmark, LangChain: evaluation datasets and reference answers. Merchant-controlled records describe the provider’s own capabilities, terms or standards; they do not independently validate those claims. These records do not establish independent confirmation of the product claims.
Verify any current price, plan limit, label direction, compatibility rule, or commercial term that would materially change the decision. The dated source ledger shows the underlying records so this conclusion can be checked and updated.
Sources used for this page
These records support the facts and comparisons above. Merchant-controlled records are labelled so you can separate product claims from independent evidence.
- SiteGPT vendor-authored Chatbase comparison; not an independent benchmark — Merchant documentation · sitegpt.ai · Merchant-controlled · checked 2026-09-25
- LangChain: evaluation datasets and reference answers — Platform documentation · docs.langchain.com · Merchant-controlled · checked 2026-09-25