The Playground runs a single question against your documents with the model of your choice and shows the retrieved chunks. Nothing is scored or saved to your results. It’s for exploring.

Use it to

  • Try real user questions before adding them to a golden dataset.
  • Find chunk IDs for manual golden dataset entries: the ID of each retrieved chunk is shown.
  • Check for missing expected chunks: run a golden question and see whether other chunks also answer it.
  • Get a feel for a model before running a full comparison.

Run a query

1

Open Playground

2

Choose an embedding model and Top K

Top K can be from 1 to 20.
3

Enter a question

Type it, or under Try a Golden-Set Question, open a dataset and click one of its questions.
4

Click Run Semantic Query

The retrieved chunks are listed in rank order, with their chunk ID and similarity score.
Recent queries are kept in the page’s history so you can go back to them, until you leave the page.
The first query with a model embeds all your chunks, like the first evaluation, so it takes longer. Later queries only embed the question.

About similarity scores

Scores show how close each chunk is to the question, for that model. They’re useful to compare chunks within one result list, but not between models: each model has its own scale, and a 0.45 from one model can be as strong a match as a 0.80 from another.