The problem RAG solves

A large language model (LLM) like ChatGPT or Claude knows a lot about the world, but it doesn’t know your internal documents, your latest product release or your company’s policies. Retrieval-augmented generation (RAG) fixes this: before answering, the system looks up the passages of your documents that are relevant to the question, and gives them to the LLM as context. If retrieval brings back the wrong passages, even the best LLM will give a wrong or vague answer, or say it doesn’t know. Most RAG answer-quality problems are retrieval problems. Truvec focuses on that step.

Chunks

Documents are too long to hand to a model whole, so they are split into smaller passages called chunks, typically a few paragraphs each. Retrieval returns chunks, not whole documents. How you split documents matters, and is covered in Chunking.

Embeddings: meaning as coordinates

An embedding model turns a piece of text into a list of numbers, called a vector or an embedding. You can think of it as coordinates on a map of meaning: texts about similar things end up close together, even when they use different words. For example, “How do I reset my password?” and “I forgot my login credentials” share almost no words, but a good embedding model places them close together. A keyword search would miss the connection.

How retrieval uses embeddings

1

Index the documents (once)

Every chunk is converted to a vector and stored in a vector database.
2

Embed the question

When a question arrives, it is converted to a vector with the same model.
3

Find the nearest chunks

The vector database returns the chunks whose vectors are closest to the question’s vector. Closeness is measured with a similarity score (cosine similarity, usually between 0 and 1 for these models, where higher means more similar).
4

Keep the top results

The system keeps the top k chunks, for example the 5 most similar, and passes them to the LLM.

Why the embedding model matters

The embedding model decides what “similar” means. Two models can rank the same chunks very differently for the same question:
  • Some understand technical or domain vocabulary better than others.
  • Larger models tend to be more accurate, but cost more per token and produce bigger vectors that take more storage.
  • A model that excels on generic web text can struggle with your jargon, abbreviations or languages.
That’s why the right choice depends on your data, and why it’s worth measuring.

What Truvec measures

Truvec evaluates the retrieval step in isolation. For each test question it knows which chunks contain the answer, and it checks whether each model brings those chunks back near the top. It doesn’t evaluate the final answer written by the LLM. That’s deliberate: if the right chunks aren’t retrieved, no LLM can answer well, so retrieval is the foundation to get right first.

Chunking

How documents are split, and why it affects results.

Golden datasets

The test questions that make an evaluation possible.