Practical guide

Separate critical chatbot failures from wording defects

Last materially reviewed 2026-09-25

Quick answerClassify failures by their consequences for the reader, not by how awkward the sentence sounds.
What to know

Define severity before seeing results

List failures that would prevent public use of the pilot, such as invented destructive steps, unauthorized account claims or materially wrong permissions. Separate those from missing context and minor clarity issues. The classification should reflect the actual reader task. Avoid assigning severity only after a preferred product performs badly; that makes the acceptance rule a defense of the product rather than a useful quality boundary.

What to know

Keep one serious error visible

Do not let many easy successful answers cancel a critical failure in an average score. Record the failed case, expected evidence and required repair as a separate release condition. A percentage can summarize a test set, but it cannot explain whether the remaining errors are trivial or consequential. Include the count and nature of unresolved cases whenever you discuss overall performance.

What to know

Use a small fictional triage

An answer with a slightly long introduction may need editing. An answer missing a required administrator role needs factual repair. An answer falsely claiming to have deleted a workspace violates the read-only boundary. These are different problems even if all receive a reader downvote. The example illustrates classification, not a measured failure distribution for any vendor or a universal risk scale.

What to know

Give every held case an owner

Record who will correct the source, configuration or interface and what evidence will close the issue. If the cause is unknown, preserve that status rather than changing it to low priority. Retest the affected case and related boundary cases after a repair. If the team cannot maintain the review burden, reduce the assistant’s remit or retain ordinary documentation access instead of launching an unreviewable promise.

  • Assign each held case a consequence, an owner and a reopening condition rather than merely marking it with a colored severity label.
Continue when useful

Next: A practical rubric for checking documentation answers

Check correctness, source support, version and useful limits separately; fluent prose cannot compensate for a material error.

Open A practical rubric for checking documentation answers →

Sources used for this page

These records support the facts and comparisons above. Merchant-controlled records are labelled so you can separate product claims from independent evidence.

  1. LangChain: evaluation datasets and reference answers — Platform documentation · docs.langchain.com · Merchant-controlled · checked 2026-09-25