A golden dataset is the answer key for your evaluations. This guide walks through creating one. For why it matters and what makes a good question, read Golden datasets first.
You need documents in the project before creating a dataset. See Upload documents.

1. Generate the first questions

1

Open Golden Set and click Create Dataset

2

Describe the dataset

Enter a name and an optional description, for example “Customer questions, v1”.
3

Choose the number of questions

Up to 20 on Free, 100 on Starter and 200 on Team Pro. Start with 20 to 50. You’ll review every question, so generate a number you can realistically check.
4

Click Generate Dataset

Truvec sends your chunks to an AI model in batches, spread across the whole corpus, and asks for realistic questions answered by those chunks, with one to three expected chunks each. Malformed answers from the AI are retried automatically. Generation takes from a few seconds to a minute or two.

2. Review every question

Open the dataset. Each row shows a question, its expected chunks and a status.
  • Click a chunk chip (for example #14 · returns.pdf) to read the full chunk. Use the arrows to step through all the expected chunks of the question. Hovering a chip shows a preview.
  • Check the status column. Valid means all expected chunks exist. Missing chunk means one of them was deleted with its document.
For each question, ask yourself:
If not, edit the chunk list or delete the question.
Add it to the expected chunks. Otherwise, a model that retrieves it is counted as wrong. Use the Playground to search for other chunks that answer the question.
Rewrite questions that copy the document’s wording or sound unnatural.
“What is the policy?” could match dozens of chunks. Make it specific or delete it.
To edit a question, click the pencil icon. To delete it, click the trash icon.

3. Add your own questions

Click Add Entry, type the question and enter the expected chunk IDs separated by commas, for example 14, 15. To find chunk IDs, search the Ingested Chunks table on the Data Upload page (the ID column), or run the question in the Playground, which shows the ID of every chunk it retrieves.
The best questions come from real users: support tickets, search logs, chat transcripts, sales calls. Even ten real questions make your dataset much more representative.

4. Commit the dataset

When the questions are right, click Commit Dataset. The dataset becomes read-only and appears in the Evaluation and Comparison pages. Only committed datasets can be evaluated. This ensures that results from different runs are graded on exactly the same questions.
A committed dataset can’t be edited. To change questions later, create a new dataset. Keeping the old one lets you still compare with earlier results.

Keeping datasets up to date

If you add or remove documents after creating a dataset, it shows a Corpus changed warning. New chunks may answer existing questions without being listed as expected, which lowers scores unfairly. After significant document changes, review the dataset or generate a new one.