When a model answers “How long do I have to return an item?” by retrieving chunk #14 first, it scores well. If #14 only appears in position 4, or not at all, its score drops.
The quality of your golden dataset sets the ceiling on how much you can trust the results. A careful set of 30 questions beats a careless set of 200.
Generated vs hand-written questions
- Generated with AI
- Written by hand
Truvec can write questions for you. It gives an AI model batches of your chunks and asks for realistic questions answered by those chunks, with one to three expected chunks each.Pros: fast, covers the whole corpus, and a great way to start.Watch out for:
- Questions that copy the wording of the chunk. They are too easy, because matching words is trivial.
- Questions that are vague or could be answered by many chunks.
- Missing expected chunks: another chunk also answers the question but wasn’t listed. A model that retrieves it looks wrong when it isn’t.
What makes a good test question
It sounds like a real user
It sounds like a real user
Real users write “can I get my money back after 2 weeks”, not “What is the duration of the return eligibility period?”. Use their words, including informal phrasing.
It doesn't copy the document
It doesn't copy the document
If the question repeats the exact phrasing of the chunk, any model finds it by matching words, and the test tells you little. Paraphrase.
Its expected chunks are complete
Its expected chunks are complete
List every chunk that genuinely answers the question. If the answer is split across two chunks, list both.
It has a clear answer in the documents
It has a clear answer in the documents
Avoid questions that the documents only partly answer, or that need outside knowledge.
The set covers the important topics
The set covers the important topics
Include easy and hard questions, frequent and rare topics, and every document type. A set that only covers one document measures only that document.
How many questions?
With few questions, one question can move a score by several points. If two models are within a few points of each other on a small set, treat them as tied.
Committing a dataset
While a dataset is a draft, you can edit, add and delete its questions. When you commit it, it becomes read-only and can be used for evaluations. Locking the questions guarantees that every model you evaluate is graded against exactly the same answer key, so results stay comparable over time.When your documents change
A golden dataset refers to specific chunks. If you add or remove documents afterwards:- New chunks may also answer existing questions. They aren’t listed as expected, so a model that retrieves them is counted as wrong, and scores drop for no real reason.
- Removed chunks disappear from the answer key, and the affected questions become impossible to get right.
Build your golden dataset
Step-by-step guide to generating, reviewing and committing a dataset.