Retrieval quality is only half of the decision. A model that is 2% better but five times more expensive, or twice as slow, may not be the right choice. Truvec measures the practical side of every model alongside its scores.

Cost

Embedding providers bill by the number of tokens you send them (a token is roughly three quarters of an English word). Truvec records the tokens billed by the provider and multiplies them by the model’s price per million tokens. There are two kinds of costs in a RAG system, and Truvec reports both: Truvec also shows Billed by this run: the cost of what the evaluation actually sent to the provider. The first run with a model embeds all your chunks. Later runs with the same model reuse them and only pay for the questions, so they are much cheaper.
Estimating production costs: index cost scales with the size of your whole document collection, so your sample’s index cost × (full corpus size ÷ sample size) gives a rough production estimate. Query cost scales with traffic: multiply the cost per 1,000 queries by your expected monthly questions ÷ 1,000.
Prices come from each model’s price per 1M tokens, which you can check in Settings. See Model pricing. If a model has no price set, its costs are shown as ”—”.
Embedding is usually a small part of a RAG system’s cost compared to the LLM that writes answers. Cost matters most when you index large collections, re-index often, or serve high query volumes.

Speed

In production, query embedding latency is part of every answer your users wait for. Note that latency depends on the network and on the provider’s load at the time of the run, so compare models measured in the same comparison rather than across days.

Storage

Vector storage is the space the vectors take: number of chunks × dimensions × 4 bytes. It excludes the vector database’s own index overhead, so budget extra on top. Dimensions is the length of each vector. Larger vectors can capture more nuance but take proportionally more space and memory, and make searches slightly slower. For example, a 3,072-dimension model needs twice the storage of a 1,536-dimension model for the same documents. This matters at scale: millions of chunks can mean many gigabytes of vectors.

Putting it together

The Comparison page highlights the best model for quality and the cheapest, fastest and smallest ones, and plots quality against index cost. The best trade-offs sit in the top-left corner of that chart: high quality, low cost.

Choosing a model

A practical method to weigh quality against cost and speed.