Every metric in Truvec answers some version of one question: did the model bring back the expected chunks, and how high did it rank them? All metrics are computed on the top k results of each question, and then averaged over all questions.

Top k

k is the number of chunks retrieved for each question. It’s a setting you choose when you run an evaluation (5 by default). Pick the number of chunks your RAG system actually sends to the LLM: if your assistant uses the 5 best chunks, evaluate with k = 5.

A worked example

Suppose the question is “How long do I have to return an item?”, the expected chunks are #14 and #15, and a model retrieves these top 5 chunks: We’ll use this example below.

The main metrics

Hit rate — did we find at least one right chunk?

1 if at least one expected chunk is in the top k, otherwise 0. Averaged over all questions, it’s the share of questions for which the LLM gets at least some of the information it needs.In the example: #14 is in the top 5, so it’s a hit (1).Read it as: “for 85% of questions, the retrieved chunks include something useful.” It’s the easiest metric to explain, but it says nothing about how many expected chunks were found or how high they were ranked.

Recall@k — how much of the answer did we find?

Expected chunks retrieved ÷ total expected chunks.In the example: 1 of the 2 expected chunks was retrieved, so recall is 0.5 (50%).Why it matters: when an answer is spread over several chunks, missing one means the LLM works with half the facts. Use it when answers often need several passages.

MRR — how high is the first right chunk?

Mean Reciprocal Rank: 1 ÷ the rank of the first expected chunk (0 if none is found).In the example: the first expected chunk is at rank 2, so the score is 0.5.Why it matters: LLMs pay more attention to the first passages they are given, and many systems only use the first 1 to 3 chunks. MRR rewards models that put the right chunk at the very top. Truvec ranks models in a comparison by MRR first, then by recall.
Normalized Discounted Cumulative Gain gives credit for every expected chunk found, with less credit the lower it is ranked, then compares the result to a perfect ranking. It goes from 0 (none found) to 1 (all expected chunks at the very top).Why it matters: it combines “did we find them all” and “were they ranked high” in a single number. It’s the most complete ranking metric, and slightly harder to explain than MRR.
Expected chunks retrieved ÷ k.In the example: 1 useful chunk out of 5 retrieved, so precision is 0.2 (20%).
Precision is naturally low when questions have few expected chunks. With one expected chunk and k = 5, the best possible precision is 20%. Use it to compare models with each other, not as an absolute grade.
The harmonic mean of precision@k and recall@k. It’s high only when both are high. It’s useful as a single summary, but it inherits the low ceiling of precision.

Diagnostic metrics

These help you understand why a model scores the way it does.
The hit rate at every cut-off from 1 to k: the share of questions with an expected chunk in the top 1, top 2, and so on. The curve shows how quickly a model finds the right chunk.A model with a low hit rate at 1 but a high hit rate at 5 finds the right chunk but not at the top. A reranker (a second model that reorders the retrieved chunks) could help, and so could sending more chunks to the LLM.
How many questions had their first expected chunk at rank 1, rank 2, and so on, plus how many were missed entirely. It reveals whether misses are “near misses” (rank 4 or 5) or complete failures.
A distractor is a retrieved chunk that isn’t expected. For each question, the score gap is the similarity of the best expected chunk minus the similarity of the best distractor, averaged over questions where both were retrieved.
  • A large positive gap means the model clearly separates the right chunk from look-alikes, so its rankings are stable.
  • A gap near zero or negative means the right chunk and wrong chunks look almost equally similar. Small changes in wording can then flip the ranking.
Raw similarity scores aren’t comparable between different models, because each model has its own scale. Compare the gap, not the raw scores.

Which metrics should I focus on?

Interpreting your results

Common result patterns and what to do about them.