Why chunk size matters
Chunks too large
A chunk that covers several topics gets a “blurry” embedding that matches many questions a little and none of them well. It also sends more irrelevant text to the LLM.
Chunks too small
A chunk that holds a single sentence can lose its context: “It must be renewed every year” means little if you don’t know what “it” is.
Overlap
Overlap repeats the end of one chunk at the start of the next. It prevents a sentence or an idea from being cut in half at a chunk boundary. An overlap of 10 to 20% of the chunk size is common, for example 200 characters for 1,000-character chunks.Strategies available in Truvec
A note on tokens
A token is the unit models read: roughly three quarters of an English word on average, so 1,000 tokens is about 750 words. Embedding models have a maximum input length measured in tokens, and providers bill by the token.Consistency: one file type, one set of settings
Different file types can and often should use different strategies: Markdown by headers, source code by functions, prose with Recursive. That reflects how a real pipeline works. What you should avoid is splitting the same kind of document in two different ways within a project, for example half your PDFs at 1,000 characters and the other half at 200 tokens. Chunk size influences similarity scores, so retrieval becomes biased toward one group of documents and your results stop reflecting a real setup. Truvec helps you stay consistent:- The settings of your first upload become the project default, and the upload form starts from them.
- If you upload a file type that already exists in the project with different settings, Truvec warns you before uploading. You can switch to the existing settings, keep yours on purpose, or cancel.
- The Chunking settings in use table on the Data Upload page lists the settings used for each file type, and marks file types chunked in more than one way as Mixed.