This guide takes you from an empty account to your first model comparison. You don’t need any prior experience with embeddings. What you need:
  • A Truvec account: sign up at app.truvec.dev, or ask a teammate to invite you to their workspace. New accounts start on the Free plan, with no card needed.
  • A few documents that represent what your assistant will answer questions about: product documentation, help-center articles, policies, reports, etc. Five to twenty documents is plenty to start.
On the Free plan, a project holds up to 5 documents and 500 chunks, a golden set has up to 20 questions, and you can run one comparison of up to 3 models. That’s enough to see Truvec work on your own data; upgrade on the Billing & team page for more.
1

Create a project

Sign in, open All Projects and click New Project. Give it a name that describes the use case, for example “Support assistant” or “HR policies”.A project holds one set of documents, the test questions about them, and all the evaluations you run.
2

Upload your documents

Open the project and go to Data Upload. Click Browse files and select your documents.Keep the default chunking settings for now (Recursive, 1000 characters, 200 overlap). Truvec extracts the text and splits it into chunks. When it’s done, the chunks appear in the Ingested Chunks table.
Choose documents that people really ask about. The evaluation is only as meaningful as the sample you give it.
3

Generate a golden dataset

Go to Golden Set and click Create Dataset. Enter a name, choose how many questions to generate (20 is a good start) and click Generate Dataset.Truvec uses an AI model to write realistic questions from your chunks and records which chunks answer each one.
4

Review and commit the questions

Open the dataset and read the questions. Click a chunk number to see its text and check that it really answers the question. Edit or delete anything that looks wrong, and add questions your users actually ask.When you’re satisfied, click Commit Dataset. A committed dataset is locked, so every evaluation uses exactly the same questions.
Don’t skip the review. AI-generated questions are a starting point, and a few bad questions can distort your scores. See Golden datasets for what makes a good question.
5

Compare models

Go to Comparison, select your golden dataset, select two or more embedding models and click Run Comparison.Each model is evaluated on every question. This usually takes from a few seconds to a few minutes, depending on how many documents and questions you have. The page updates when the comparison finishes.
6

Read the results

Select the comparison in Comparison History to see:
  • a ranked table with quality, speed, cost and storage for each model,
  • highlights such as the cheapest model and the fastest one,
  • a chart of quality against cost,
  • the Retrieval Inspector, where you can step through each question and see what every model retrieved.
Not sure which number to look at? Start with MRR and Recall, then read Interpreting your results.

Next steps

Understand the metrics

What hit rate, recall, MRR and nDCG mean, with examples.

Choose a model

A practical way to turn results into a decision.

Chunking

How splitting your documents affects retrieval, and which strategy to pick.

Best practices

Avoid the most common evaluation mistakes.