Decide the split before tuning
Create development questions for configuration work and separate evaluation questions for the later decision. Keep the expected source references for both. The split should cover the same kinds of reader tasks without copying every phrase. There is no universal split ratio in this guide; choose enough variety to expose your important failure modes within the review effort you can sustain.
Keep reserved cases out of the instructions
If a held-out question becomes a worked prompt example or a special response rule, it is no longer independent of that tuning. Record the change and move it into the development set. Do not quietly keep calling it held-out because the result now looks better. This is especially important when a small team remembers every difficult question and unintentionally optimizes the setup around it.
Run the decision check once per frozen configuration
Record the source collection and relevant settings before the evaluation. If the result reveals a serious issue, repair it and describe the next test as a new evaluation, not a corrected version of the original pass. Preserve failed outputs. A clean record helps distinguish genuine improvement from selective reporting and prevents an attractive dashboard from hiding the cost of repeated attempts.
Add new cases from new problems
After launch, real questions may reveal a missing task or a novel ambiguity. Add an appropriately sanitized case and its expected evidence to the regression set. Keep its origin and date without retaining unnecessary personal details. The goal is broader coverage of meaningful failures, not a larger score denominator. A small honest set is more useful than a large set of near-identical questions.
- Label each case development, held-out or regression, and record when its status changes so evaluation independence is not silently lost.
Sources used for this page
These records support the facts and comparisons above. Merchant-controlled records are labelled so you can separate product claims from independent evidence.
- LangChain: evaluation datasets and reference answers — Platform documentation · docs.langchain.com · Merchant-controlled · checked 2026-09-25