1
Make sure the comparison is trustworthy
Before looking at the winner, check that:
- the golden dataset has been reviewed and includes real user questions,
- it has enough questions (50 or more for close decisions),
- top k matches what your system sends to the LLM,
- there is no Corpus changed badge on the comparison.
2
Pick your primary metric
Usually MRR, if your LLM relies mostly on the first chunks, or recall@k if answers often need several chunks. See Which metrics should I focus on?.
3
Decide what difference is meaningful
With 30 questions, one question is worth about 3 points of hit rate. A 2-point lead is likely noise. As a rule of thumb, treat models within a few points as tied unless your dataset is large. The question by question view tells you whether a lead comes from many questions or a lucky few.
4
Among the best, prefer the cheaper and faster
From the models that are best or tied for best on your primary metric, choose based on:
- cost per 1,000 queries × your expected traffic, and index cost × how often you re-index,
- latency per query, if users wait for answers in real time,
- vector storage, if you’ll index millions of chunks.
5
Check the questions that matter most
Some questions are more important than others: legal, safety or billing questions, for example. In the Retrieval Inspector, check that your candidate model handles them.
6
Record the decision
The comparison stays in Reports as a record of which models were tested and which one won. Export it as a PDF report to share the decision and its evidence with your team.
Example
On 40 questions, a 0.02 MRR difference is within noise. Model B is 6.5 times cheaper per query and needs half the storage. Choose Model B, unless the question-by-question view shows Model A consistently wins on your most important questions.
If Model A were at 0.90 MRR instead, the gap would likely be real and worth the cost for most use cases, because better retrieval directly means better answers.