If you are building a chatbot, a search feature, or any assistant that answers questions from your documents, its answers can only be as good as the passages it finds. That finding step, called retrieval, depends heavily on the embedding model you choose. Truvec lets you measure that choice on your own data instead of guessing.

What Truvec does

You bring a sample of your documents. Truvec splits them into passages, builds a set of test questions whose correct passages are known, and checks how often each embedding model finds those passages. You get clear scores, the cost of each model, and a question-by-question view of what each model retrieved.

Measure retrieval quality

See how often the right passage is found, and how high it is ranked, for each embedding model.

Compare models side by side

Run several models on the same questions and pick the best trade-off between quality, cost and speed.

Understand every miss

Open any test question and see exactly which passages were retrieved and which expected ones were missed.

Know what it will cost

Get the cost to index your documents and to answer every thousand queries, from real token counts.

Who it is for

Truvec is for anyone who needs to choose or validate an embedding model for retrieval-augmented generation (RAG), including people who have never run an evaluation before:
  • Product and engineering teams choosing a model before building, or checking whether a cheaper model is good enough.
  • Teams with a RAG system in production who want evidence before switching models or changing how documents are split.
  • Anyone new to RAG who wants to understand why their assistant sometimes answers from the wrong passage.
New to embeddings or RAG? Start with How RAG retrieval works. It explains the ideas you need in a few minutes, without math.

How it works

1

Create a project and add documents

Upload a representative sample of your documents (PDF, Word, Markdown, text, CSV or JSON). Truvec splits them into passages called chunks.
2

Build a golden dataset

Generate test questions with AI, or write your own. Each question is linked to the chunks that contain its answer. Review them, then lock the dataset.
3

Evaluate and compare models

Run one model or compare several. Truvec asks every question and checks where the correct chunks appear in the results.
4

Read the results and decide

Compare quality, cost, speed and storage, then inspect individual questions to understand each model’s strengths and weaknesses.
Public leaderboards test models on generic data. Your documents have their own vocabulary, structure and question types: internal product names, legal clauses, support tickets, code. A model that ranks first on a leaderboard can be beaten on your data by a model that costs a fraction of the price. The only reliable way to know is to test on your own content, which is what Truvec is for.

Quickstart

Run your first evaluation in about 15 minutes.

Understand the metrics

Learn what each score means and which ones matter most.