The model with the highest score isn’t automatically the right choice. This method weighs quality against cost and speed, and helps you avoid choosing on noise.
1

Make sure the comparison is trustworthy

Before looking at the winner, check that:
  • the golden dataset has been reviewed and includes real user questions,
  • it has enough questions (50 or more for close decisions),
  • top k matches what your system sends to the LLM,
  • there is no Corpus changed badge on the comparison.
2

Pick your primary metric

Usually MRR, if your LLM relies mostly on the first chunks, or recall@k if answers often need several chunks. See Which metrics should I focus on?.
3

Decide what difference is meaningful

With 30 questions, one question is worth about 3 points of hit rate. A 2-point lead is likely noise. As a rule of thumb, treat models within a few points as tied unless your dataset is large. The question by question view tells you whether a lead comes from many questions or a lucky few.
4

Among the best, prefer the cheaper and faster

From the models that are best or tied for best on your primary metric, choose based on:
  • cost per 1,000 queries × your expected traffic, and index cost × how often you re-index,
  • latency per query, if users wait for answers in real time,
  • vector storage, if you’ll index millions of chunks.
The Quality vs index cost chart helps: the best trade-offs are in the top-left corner.
5

Check the questions that matter most

Some questions are more important than others: legal, safety or billing questions, for example. In the Retrieval Inspector, check that your candidate model handles them.
6

Record the decision

The comparison stays in Reports as a record of which models were tested and which one won. Export it as a PDF report to share the decision and its evidence with your team.

Example

On 40 questions, a 0.02 MRR difference is within noise. Model B is 6.5 times cheaper per query and needs half the storage. Choose Model B, unless the question-by-question view shows Model A consistently wins on your most important questions. If Model A were at 0.90 MRR instead, the gap would likely be real and worth the cost for most use cases, because better retrieval directly means better answers.
Re-run the comparison when your documents change significantly, when new models become available, or when you change how documents are chunked.