What Truvec does
You bring a sample of your documents. Truvec splits them into passages, builds a set of test questions whose correct passages are known, and checks how often each embedding model finds those passages. You get clear scores, the cost of each model, and a question-by-question view of what each model retrieved.Measure retrieval quality
See how often the right passage is found, and how high it is ranked, for each embedding model.
Compare models side by side
Run several models on the same questions and pick the best trade-off between quality, cost and speed.
Understand every miss
Open any test question and see exactly which passages were retrieved and which expected ones were missed.
Know what it will cost
Get the cost to index your documents and to answer every thousand queries, from real token counts.
Who it is for
Truvec is for anyone who needs to choose or validate an embedding model for retrieval-augmented generation (RAG), including people who have never run an evaluation before:- Product and engineering teams choosing a model before building, or checking whether a cheaper model is good enough.
- Teams with a RAG system in production who want evidence before switching models or changing how documents are split.
- Anyone new to RAG who wants to understand why their assistant sometimes answers from the wrong passage.
How it works
1
Create a project and add documents
Upload a representative sample of your documents (PDF, Word, Markdown, text, CSV or JSON). Truvec splits them into passages called chunks.
2
Build a golden dataset
Generate test questions with AI, or write your own. Each question is linked to the chunks that contain its answer. Review them, then lock the dataset.
3
Evaluate and compare models
Run one model or compare several. Truvec asks every question and checks where the correct chunks appear in the results.
4
Read the results and decide
Compare quality, cost, speed and storage, then inspect individual questions to understand each model’s strengths and weaknesses.
Why not just pick the most popular model?
Public leaderboards test models on generic data. Your documents have their own vocabulary, structure and question types: internal product names, legal clauses, support tickets, code. A model that ranks first on a leaderboard can be beaten on your data by a model that costs a fraction of the price. The only reliable way to know is to test on your own content, which is what Truvec is for.Quickstart
Run your first evaluation in about 15 minutes.
Understand the metrics
Learn what each score means and which ones matter most.