Practical guide

Retest chatbot answers after a source or configuration change

Last materially reviewed 2026-09-25

Quick answerKeep a frozen baseline and rerun the cases affected by a change, plus important boundary checks.
What to know

Record the changed object

Was the source edited, a model changed, an instruction revised or a connector added? Write the specific change and the reason. Avoid changing several unrelated controls at once if you want to understand the result. Preserve the earlier source list, question set and outputs. A before-and-after comparison is only meaningful when you know which configuration produced each answer and which evidence was available at the time.

What to know

Choose relevant cases without cherry-picking

Rerun the task directly affected by the change and adjacent cases that could regress. Include a version question and an out-of-scope request when those boundaries matter. Do not test only the previously failed prompt and declare the whole system repaired. Also do not rerun a huge collection by default if a small scoped check can answer the decision within the team’s review budget.

What to know

Compare reasons, not just scores

An answer can keep the same overall rating while changing from one unsupported claim to another. Compare the actual text, citation and missing requirements. If the new wording is shorter but omits a necessary qualification, that is not an unambiguous improvement. Record both benefits and new defects so a later owner can understand why the change was retained, revised or reversed.

What to know

Keep failed runs in the record

A repaired result should not overwrite the original failure. Preserve the sequence and clearly label the final accepted configuration. Budget every actual question attempt, including investigation and repeat checks, according to the product’s current metering. If an important regression remains unresolved, hold the affected public behavior or return to the last acceptable setup rather than accumulating exceptions that no one can explain.

  • Label the source snapshot, question-set version and configuration for every run; a timestamp alone does not identify what was actually tested.
Continue when useful

Next: Keep evaluation questions separate from tuning questions

Reserve a distinct set of questions so repeated tuning does not become the only evidence of improvement.

Open Keep evaluation questions separate from tuning questions →

Sources used for this page

These records support the facts and comparisons above. Merchant-controlled records are labelled so you can separate product claims from independent evidence.

  1. LangChain: evaluation datasets and reference answers — Platform documentation · docs.langchain.com · Merchant-controlled · checked 2026-09-25