Top k
k is the number of chunks retrieved for each question. It’s a setting you choose when you run an evaluation (5 by default). Pick the number of chunks your RAG system actually sends to the LLM: if your assistant uses the 5 best chunks, evaluate with k = 5.A worked example
Suppose the question is “How long do I have to return an item?”, the expected chunks are #14 and #15, and a model retrieves these top 5 chunks:
We’ll use this example below.
The main metrics
Hit rate — did we find at least one right chunk?
Hit rate — did we find at least one right chunk?
1 if at least one expected chunk is in the top k, otherwise 0. Averaged over all questions, it’s the share of questions for which the LLM gets at least some of the information it needs.In the example: #14 is in the top 5, so it’s a hit (1).Read it as: “for 85% of questions, the retrieved chunks include something useful.” It’s the easiest metric to explain, but it says nothing about how many expected chunks were found or how high they were ranked.
Recall@k — how much of the answer did we find?
Recall@k — how much of the answer did we find?
Expected chunks retrieved ÷ total expected chunks.In the example: 1 of the 2 expected chunks was retrieved, so recall is 0.5 (50%).Why it matters: when an answer is spread over several chunks, missing one means the LLM works with half the facts. Use it when answers often need several passages.
MRR — how high is the first right chunk?
MRR — how high is the first right chunk?
Mean Reciprocal Rank: 1 ÷ the rank of the first expected chunk (0 if none is found).
In the example: the first expected chunk is at rank 2, so the score is 0.5.Why it matters: LLMs pay more attention to the first passages they are given, and many systems only use the first 1 to 3 chunks. MRR rewards models that put the right chunk at the very top. Truvec ranks models in a comparison by MRR first, then by recall.
nDCG@k — how good is the whole ranking?
nDCG@k — how good is the whole ranking?
Normalized Discounted Cumulative Gain gives credit for every expected chunk found, with less credit the lower it is ranked, then compares the result to a perfect ranking. It goes from 0 (none found) to 1 (all expected chunks at the very top).Why it matters: it combines “did we find them all” and “were they ranked high” in a single number. It’s the most complete ranking metric, and slightly harder to explain than MRR.
Precision@k — how much of what we retrieved is useful?
Precision@k — how much of what we retrieved is useful?
Expected chunks retrieved ÷ k.In the example: 1 useful chunk out of 5 retrieved, so precision is 0.2 (20%).
F1@k — precision and recall combined
F1@k — precision and recall combined
The harmonic mean of precision@k and recall@k. It’s high only when both are high. It’s useful as a single summary, but it inherits the low ceiling of precision.
Diagnostic metrics
These help you understand why a model scores the way it does.Hit rate by cut-off
Hit rate by cut-off
The hit rate at every cut-off from 1 to k: the share of questions with an expected chunk in the top 1, top 2, and so on. The curve shows how quickly a model finds the right chunk.A model with a low hit rate at 1 but a high hit rate at 5 finds the right chunk but not at the top. A reranker (a second model that reorders the retrieved chunks) could help, and so could sending more chunks to the LLM.
Rank of the first relevant chunk
Rank of the first relevant chunk
How many questions had their first expected chunk at rank 1, rank 2, and so on, plus how many were missed entirely. It reveals whether misses are “near misses” (rank 4 or 5) or complete failures.
Score gap, relevant score and top distractor score
Score gap, relevant score and top distractor score
A distractor is a retrieved chunk that isn’t expected. For each question, the score gap is the similarity of the best expected chunk minus the similarity of the best distractor, averaged over questions where both were retrieved.
- A large positive gap means the model clearly separates the right chunk from look-alikes, so its rankings are stable.
- A gap near zero or negative means the right chunk and wrong chunks look almost equally similar. Small changes in wording can then flip the ranking.
Which metrics should I focus on?
Interpreting your results
Common result patterns and what to do about them.