Define severity before seeing results
List failures that would prevent public use of the pilot, such as invented destructive steps, unauthorized account claims or materially wrong permissions. Separate those from missing context and minor clarity issues. The classification should reflect the actual reader task. Avoid assigning severity only after a preferred product performs badly; that makes the acceptance rule a defense of the product rather than a useful quality boundary.
Keep one serious error visible
Do not let many easy successful answers cancel a critical failure in an average score. Record the failed case, expected evidence and required repair as a separate release condition. A percentage can summarize a test set, but it cannot explain whether the remaining errors are trivial or consequential. Include the count and nature of unresolved cases whenever you discuss overall performance.
Use a small fictional triage
An answer with a slightly long introduction may need editing. An answer missing a required administrator role needs factual repair. An answer falsely claiming to have deleted a workspace violates the read-only boundary. These are different problems even if all receive a reader downvote. The example illustrates classification, not a measured failure distribution for any vendor or a universal risk scale.
Give every held case an owner
Record who will correct the source, configuration or interface and what evidence will close the issue. If the cause is unknown, preserve that status rather than changing it to low priority. Retest the affected case and related boundary cases after a repair. If the team cannot maintain the review burden, reduce the assistant’s remit or retain ordinary documentation access instead of launching an unreviewable promise.
- Assign each held case a consequence, an owner and a reopening condition rather than merely marking it with a colored severity label.
Sources used for this page
These records support the facts and comparisons above. Merchant-controlled records are labelled so you can separate product claims from independent evidence.
- LangChain: evaluation datasets and reference answers — Platform documentation · docs.langchain.com · Merchant-controlled · checked 2026-09-25