Scores tell you how well a model retrieves. This page helps you understand why, and what to do about it. Most of the answers come from reading a handful of questions in the Retrieval Inspector, especially the misses.

What is a good score?

There is no universal threshold: scores depend on how hard your questions are and how similar your documents are to each other. As a rough guide for a golden dataset of realistic questions:
Very high scores on AI-generated questions can be misleading: generated questions often reuse the document’s words, which makes retrieval easy. Add real user questions to get a realistic picture.

Common patterns

What you see: the right chunk is usually in the top k, but rarely at rank 1. The hit-rate-by-cut-off curve starts low and climbs.What it means: the model finds the right area, but similar chunks outrank the best one.What to try:
  • Send more chunks to your LLM (a higher top k), if your context budget allows.
  • Add a reranker, a model that reorders the retrieved chunks by relevance. It typically fixes exactly this pattern.
  • Compare models: some rank more precisely than others on your data.
What you see: in the rank distribution, the Miss bar is large.What to check, in order:
  1. The golden dataset. Open a few misses. Is the expected chunk really the best answer? Did the model retrieve another chunk that also answers? If so, the dataset is wrong, not the model.
  2. Extraction. Open the expected chunk. Is the text readable, or jumbled (PDF tables, headers, footers)? Unreadable chunks can’t be matched.
  3. Chunking. Is the answer split across chunks, or buried in a very long chunk about several topics? Try different settings in a separate project.
  4. Vocabulary. Do users use terms that never appear in the documents (abbreviations, internal names, another language)? Try a larger model.
What it means: on your data, the models are equivalent, or your golden dataset is too small or too easy to tell them apart.What to do: choose on cost and speed. If the decision matters, add harder and more realistic questions and run the comparison again.
What it means: the right chunk and the best wrong chunk get nearly the same similarity. The model has trouble telling them apart, so small rewordings can flip the ranking.Common causes: near-duplicate documents (several versions of the same policy, templates), or chunks that mostly contain boilerplate text.What to try: remove outdated duplicates from your corpus, strip boilerplate, or compare models with a larger gap.
What it means: usually nothing. With one expected chunk and top k = 5, the best possible precision is 20%. Precision is only informative relative to other models. See Metrics.
What it means: probably not a worse model. The new chunks likely also answer some golden questions but aren’t listed as expected, so they count as wrong. Results created before the change show a Corpus changed badge.What to do: review the golden dataset (add the new chunks as expected where they answer), or generate a new dataset.

Reading misses efficiently

1

Filter on Misses

In the Retrieval Inspector, select Misses.
2

Read ten questions

For each, compare the retrieved chunks with the Expected but not retrieved chunks. Click chunks to read their full text.
3

Label each miss

Put each one in a bucket: golden dataset error, extraction problem, chunking problem, model weakness.
4

Fix the biggest bucket first

If most misses are dataset errors, fix the dataset before comparing models, or you’ll choose a model based on noise.