Kubesense

Experiments

An experiment runs a prompt version and model configuration across every item in a dataset and aggregates the results. It is how you compare a candidate against the current setup on identical inputs, rather than inferring quality from production traffic where the inputs are never the same twice.

See Concepts for the run and run-item objects.

Before you start

An experiment needs four things in place:

  1. A dataset with items whose input keys match the variables your prompt expects.
  2. A prompt version to test.
  3. An LLM connection for the application, since KubeSense calls the model on your behalf — see Installation.
  4. Optionally an evaluator to score each response.

Without an evaluator a run still produces traces, cost, and latency; it just has no quality signal beyond what you read yourself.

Run an experiment

Start a run from Run Experiment on the Experiments tab, or from Run experiment on a dataset — the same setup either way.

Running an experiment from a dataset

A dataset's own Experiments tab lists only the runs executed against it, which is the view you want when comparing candidates on one collection.

Walkthrough

A run is defined by:

FieldMeaning
NameWhat this run is testing — it is how you will tell runs apart later
DatasetThe items to run against
PromptThe prompt version under test
Model configThe model and generation parameters to run it with
Structured output schemaOptional JSON Schema constraining the response shape
EvaluatorOptional rubric scoring each response

Change one thing per run. A run that changes the prompt and the model tells you the pair is better or worse, not which one caused it.

Name runs so the variable is visible — v4 gpt-5-mini beats test 3. A month later the name is all you have.

note: A run executes against the dataset as it exists at that moment. Adding or editing items afterwards does not change a completed run, and two runs are only comparable if the dataset did not change between them.

The experiments list

LLM Monitoring > Evaluation > Experiments lists every run for the application.

Experiments list

ColumnMeaning
NameThe run, usually naming the prompt version and dataset
StatusCOMPLETED, or in progress / failed
ModelThe model the run used, such as openai / gpt-5-mini
ProgressItems finished against total — 2/2 when everything ran
Total Cost / Average CostSpend for the run, and per item
Average LatencyMean duration per item
TokensTotal tokens consumed
ScoresAggregated score values, one per score name

Progress is the first column to read. A run showing 40/100 with good scores scored 40 items, not 100 — check it completed before comparing anything.

Run results

Open a run to see it item by item.

Experiment run detail

The heading names the run and the dataset it executed against.

ColumnMeaning
Dataset ItemThe case, linking to its definition
Run AtWhen this item was executed
TraceView Trace opens the trace this run produced for the item
LatencyDuration for this item
CostSpend for this item
Score columnsOne per score, such as # conciseness (eval), with the value for this item
Trace InputWhat was actually sent
OutputWhat came back

Each score gets its own column, so a run scored on several dimensions shows them side by side. A -- means no score was produced for that item — the evaluator failed, was sampled out, or the item errored before there was anything to judge.

View Trace is what makes a run debuggable. The aggregate tells you something changed; the trace tells you what the model actually received and returned for the case that regressed.

Compare two runs

  1. Run the current configuration to establish a baseline, if you do not already have one.
  2. Change one variable and run again against the same dataset.
  3. Compare score aggregates, cost, and latency side by side.
  4. Open the items where the two runs disagree most — the aggregate tells you whether something changed, the traces tell you what.
  5. Check cost and latency, not only quality. A candidate that scores marginally better for double the spend is usually not the one to ship.

A dataset item's cross-run view shows every run's result for a single case, which is the fastest way to find cases that regressed while the average improved.

Scoring a run

Runs produce scores the same way production traffic does:

  • Attach an LLM judge rubric to score each response automatically
  • Compare against the item's expected output where one is defined
  • Review the run's traces through an annotation queue for a candidate you are about to ship

Where to next

  • Datasets — the cases a run executes against
  • Scores — how run results are recorded
  • Prompt Hub — promote the winning version with a label