A comparison evaluates two or more embedding models on the same golden dataset, with the same top k, so their results are directly comparable. It’s the fastest way to choose a model.

Run a comparison

1

Open Comparison

2

Name it (optional)

For example “Small vs large on support docs”. Without a name, the date is used.
3

Select a golden dataset

Only committed datasets are listed.
4

Select at least two models

5

Set Top K and click Run Comparison

Each model is evaluated in turn in the background. The comparison shows how many models are done and updates automatically.
Include your current model, if you have one, as a baseline. The question is rarely “which model is best?” but “is another model better enough to justify switching?”

Reading a comparison

Select a comparison in Comparison History to open its results.

The winner

Models are ranked by MRR, then by recall. The banner shows the top model and its scores. Treat a small lead with caution: on a small golden dataset, a difference of a few points can come down to one or two questions. See Choosing a model.

Highlights

Four cards show the cheapest to index, the cheapest per 1,000 queries, the fastest queries and the smallest index. When the best model is also the cheapest or fastest, the decision is easy. When it isn’t, weigh the difference.

Results table

One row per model with hit rate, recall, precision, MRR and nDCG at your top k, latency, index cost, cost per 1,000 queries and vector storage. Models that failed are marked Failed; hover the badge to see why.

Quality vs index cost

A chart of MRR against the cost to index your corpus. Models in the top-left give the most quality for the money. A model that is far to the right but barely higher is rarely worth it.

Question by question

The Retrieval Inspector shows every model:
  • Model buttons switch between models. Each shows where that model ranked the right chunk for the current question (for example rank 1, or miss), so you can see at a glance where the models disagree.
  • The Hit rate by cut-off and Rank of the first relevant chunk charts overlay all models.
  • The Cost & footprint and Separation & latency tables have one column per model.
The most useful questions to read are the ones where models disagree: one finds the right chunk at rank 1 and another misses it. They show what each model understands better.

Comparisons and evaluations

Each model in a comparison is a regular evaluation behind the scenes, but comparison runs aren’t listed in Evaluation History, to keep it readable. Deleting a comparison deletes its evaluations.