An evaluation runs every question of a golden dataset against one embedding model and scores the results. Use it to get a baseline, or to check a model you’ve already chosen. To compare several models at once, use Compare models.

Run an evaluation

1

Open Evaluation

2

Select a golden dataset

Only committed datasets are listed. See Build a golden dataset.
3

Select an embedding model

Each model shows its provider, number of dimensions and price per million tokens.
4

Set Top K

The number of chunks retrieved per question, from 1 to 50 (5 by default). Match what your RAG system sends to the LLM.
5

Click Run Evaluation

The evaluation runs in the background and appears in Evaluation History with a Running status, which updates automatically.
The first evaluation of a model in a project embeds all your chunks, which takes longer and costs more. Later evaluations with the same model reuse those embeddings and only embed the questions.

Evaluation history

Every evaluation of the project is listed, newest first, with its model, dataset, top k, status, and for completed runs the hit rate, MRR and cost. Click an evaluation to open its results below the list. Delete an evaluation with the trash icon. A Corpus changed badge means documents were added or removed after the evaluation ran, so its results may not reflect the current documents. If an evaluation Failed, open it to read the error. See Troubleshooting for common causes.

The Retrieval Inspector

The results of a completed evaluation open in the Retrieval Inspector.

Metric cards

Hit rate, recall, MRR and nDCG at your top k, latency per query, and cost per 1,000 queries. See Metrics and Cost, speed and storage.

Charts and tables

  • Hit rate by cut-off: how quickly the right chunk shows up in the ranking.
  • Rank of the first relevant chunk: how many questions were answered at rank 1, rank 2, and so on, and how many were missed.
  • Cost & footprint: price, index cost, cost per 1,000 queries, what this run billed, corpus size in tokens, vector storage and dimensions.
  • Separation & latency: score gap between the right chunks and the best wrong one, and embedding and search times.

Question by question

This is where you learn the most. Use the All questions / Hits / Misses filter, the question list and the arrows to move through questions. For each one you see:
  • the question and its target chunks,
  • the retrieved chunks in rank order, with similarity scores, marked Ground truth or Distractor,
  • Expected but not retrieved: target chunks that didn’t make the top k,
  • whether the ground truth was found, and at which rank.
Click any chunk to read its full text.
Start with the misses. Filter on Misses and read a handful. You’ll usually see a pattern quickly: vague questions, missing expected chunks in your golden dataset, chunks that are too long, or a genuine weakness of the model. See Interpreting your results.