Do

Use real questions

Mix AI-generated questions with questions from support tickets, search logs or chat history. Real phrasing is what your system will face.

Review the golden dataset

Read every generated question and its expected chunks before committing. A wrong answer key produces wrong conclusions.

Match your production setup

Use the same chunking and the same top k as your RAG system. If you already have chunks, upload them as pre-chunked data.

Read the misses

Scores tell you how much, misses tell you why. Ten minutes in the Retrieval Inspector is worth more than another decimal of MRR.

Keep a baseline

Include your current model in every comparison, so you always know whether a change is an improvement.

Change one thing at a time

To compare models, keep chunking fixed. To compare chunking, keep the model fixed and use separate projects.

Avoid

With 20 questions, one question moves hit rate by 5 points. Treat close scores as ties, or add questions.
Questions that reuse the document’s exact words are easy for every model and hide real differences. Paraphrase, and include hard cases.
New chunks can answer existing questions without being listed as expected, which lowers scores unfairly. Upload all documents first, then build the dataset. Watch for the Corpus changed badge.
It biases retrieval toward one group of documents. Keep one set of settings per file type. Truvec warns you when you’re about to mix them.
Each model has its own scale. Compare rankings and metrics, not raw scores.
A small quality gain rarely justifies a large increase in cost or latency. See Choosing a model.

A checklist before you decide

  • Documents are a representative sample, and all of them are uploaded.
  • Chunking matches production, with one set of settings per file type.
  • The golden dataset was reviewed, includes real questions and has 50+ questions for close calls.
  • Top k matches what your system sends to the LLM.
  • Your current model is included as a baseline.
  • You read the misses of the top models.
  • You weighed cost, latency and storage, not just quality.