The documents you upload are the corpus every model is evaluated on. Go to Data Upload in your project.

Choose a good sample

You don’t need your whole document collection. A representative sample gives reliable results faster and costs less to embed.
  • Cover every document type your system will search: help articles, PDFs, policies, FAQs, and so on.
  • Include documents people actually ask about, not only the easiest ones.
  • Keep similar documents together. Near-duplicates (several versions of the same policy, for example) make retrieval harder, which is realistic if your production data has them too.
For a first evaluation, 5 to 50 documents is usually enough.

Upload raw documents

1

Open the Raw Documents tab

It’s selected by default.
2

Choose the chunking settings

Pick a strategy and its settings in the Chunk Configuration panel. The description under the strategy explains what it does. If you’re unsure, keep Recursive with 1,000 characters and 200 overlap. See Chunking for guidance.
3

Select your files

Click Browse files and select one or more files. Truvec extracts the text, splits it into chunks and lists them in Ingested Chunks.

Supported file types

Scanned PDFs aren’t supported yet. Truvec reads the PDF’s text layer and doesn’t run OCR, so a scanned document produces no text. Convert scans to searchable PDFs first.Tables, headers and footers in PDFs are extracted as plain text, and their layout can come out jumbled. If your answers live in complex tables, check the resulting chunks.

Chunking consistency

The settings of your first upload become the project default, and the form starts from them next time. If you upload a file type that the project already contains with different settings, Truvec asks what to do before uploading:
  • Use existing settings: re-chunk this upload like the other files of the same type. This is recommended.
  • Upload with my settings: keep your settings on purpose.
  • Cancel.
Different file types can use different strategies without a warning, for example Markdown by headers and PDFs with Recursive. The Chunking settings in use table shows the settings of every file type. A type chunked in more than one way is marked Mixed. See why consistency matters.

Upload pre-chunked data

If your production pipeline already produces chunks, upload them directly so Truvec evaluates exactly what you run in production. Open the Pre-chunked Data tab, click Upload file and choose a file in one of the formats below.

Supported formats

Files must be UTF-8 encoded. Every format describes the same fields:
A JSON array of chunk objects:
chunks.json
A plain array of strings also works: ["first chunk", "second chunk"].

Validation

Truvec checks the whole file before importing anything. If a chunk is invalid, nothing is imported and the error tells you where the problem is (for example Item 3, or Line 12 for JSON Lines and CSV). The upload is rejected when:
  • the file isn’t .json, .jsonl or .csv, or isn’t UTF-8 text
  • a .json file isn’t a JSON array (use .jsonl for one chunk per line)
  • a chunk has no text, or its text is empty
  • an id is repeated
  • a CSV file has no text column, or a row has more values than the header
Other metadata fields are accepted but not stored. Pre-chunked files aren’t re-split and aren’t subject to the chunking consistency check.

After adding or removing documents

Changing the documents changes what the correct answers can be. Golden datasets and results created before the change are marked Corpus changed. Review your golden datasets after large changes. See When your documents change.

Deleting documents

To start over, click Clear above the chunks list. It removes every document of the project and their chunks. Golden dataset entries that expected those chunks are flagged as Missing chunk. Removing a single document isn’t available in the interface yet; you can do it with the API.