Do
Use real questions
Mix AI-generated questions with questions from support tickets, search logs or chat history. Real phrasing is what your system will face.
Review the golden dataset
Read every generated question and its expected chunks before committing. A wrong answer key produces wrong conclusions.
Match your production setup
Use the same chunking and the same top k as your RAG system. If you already have chunks, upload them as pre-chunked data.
Read the misses
Scores tell you how much, misses tell you why. Ten minutes in the Retrieval Inspector is worth more than another decimal of MRR.
Keep a baseline
Include your current model in every comparison, so you always know whether a change is an improvement.
Change one thing at a time
To compare models, keep chunking fixed. To compare chunking, keep the model fixed and use separate projects.
Avoid
Trusting small differences on small datasets
Trusting small differences on small datasets
With 20 questions, one question moves hit rate by 5 points. Treat close scores as ties, or add questions.
Only testing easy questions
Only testing easy questions
Questions that reuse the document’s exact words are easy for every model and hide real differences. Paraphrase, and include hard cases.
Changing documents after building the dataset
Changing documents after building the dataset
New chunks can answer existing questions without being listed as expected, which lowers scores unfairly. Upload all documents first, then build the dataset. Watch for the Corpus changed badge.
Mixing chunk settings for the same file type
Mixing chunk settings for the same file type
It biases retrieval toward one group of documents. Keep one set of settings per file type. Truvec warns you when you’re about to mix them.
Comparing similarity scores across models
Comparing similarity scores across models
Each model has its own scale. Compare rankings and metrics, not raw scores.
Choosing on quality alone
Choosing on quality alone
A small quality gain rarely justifies a large increase in cost or latency. See Choosing a model.
A checklist before you decide
- Documents are a representative sample, and all of them are uploaded.
- Chunking matches production, with one set of settings per file type.
- The golden dataset was reviewed, includes real questions and has 50+ questions for close calls.
- Top k matches what your system sends to the LLM.
- Your current model is included as a baseline.
- You read the misses of the top models.
- You weighed cost, latency and storage, not just quality.