Practical guide

Keep evaluation questions separate from tuning questions

Last materially reviewed 2026-09-25

Quick answerReserve a distinct set of questions so repeated tuning does not become the only evidence of improvement.
What to know

Decide the split before tuning

Create development questions for configuration work and separate evaluation questions for the later decision. Keep the expected source references for both. The split should cover the same kinds of reader tasks without copying every phrase. There is no universal split ratio in this guide; choose enough variety to expose your important failure modes within the review effort you can sustain.

What to know

Keep reserved cases out of the instructions

If a held-out question becomes a worked prompt example or a special response rule, it is no longer independent of that tuning. Record the change and move it into the development set. Do not quietly keep calling it held-out because the result now looks better. This is especially important when a small team remembers every difficult question and unintentionally optimizes the setup around it.

What to know

Run the decision check once per frozen configuration

Record the source collection and relevant settings before the evaluation. If the result reveals a serious issue, repair it and describe the next test as a new evaluation, not a corrected version of the original pass. Preserve failed outputs. A clean record helps distinguish genuine improvement from selective reporting and prevents an attractive dashboard from hiding the cost of repeated attempts.

What to know

Add new cases from new problems

After launch, real questions may reveal a missing task or a novel ambiguity. Add an appropriately sanitized case and its expected evidence to the regression set. Keep its origin and date without retaining unnecessary personal details. The goal is broader coverage of meaningful failures, not a larger score denominator. A small honest set is more useful than a large set of near-identical questions.

  • Label each case development, held-out or regression, and record when its status changes so evaluation independence is not silently lost.
Continue when useful

Next: Build a small, decision-focused chatbot question set

Write questions with expected evidence and failure conditions, not just a list of prompts that sound realistic.

Open Build a small, decision-focused chatbot question set →

Sources used for this page

These records support the facts and comparisons above. Merchant-controlled records are labelled so you can separate product claims from independent evidence.

  1. LangChain: evaluation datasets and reference answers — Platform documentation · docs.langchain.com · Merchant-controlled · checked 2026-09-25