Concepts
LLM output has no single correct answer, so quality cannot be asserted once and forgotten. Evaluation is the machinery for measuring it continuously: attaching quality signals to what your application produced, and comparing candidates before they ship.
This page defines the objects. Each has its own page: Scores, LLM as a Judge, Annotations, Datasets, and Experiments.
The evaluation loop
Evaluation runs in two places, and a healthy team uses both.
Offline evaluation happens before deployment. You run a candidate prompt or model across a fixed dataset as an experiment and compare the result against the current configuration on identical inputs. It is repeatable and cheap to re-run, but it only covers cases you thought to include.
Online evaluation happens in production. Scores are attached to live traces — by a judge model, a reviewer, or your own pipeline — so you find the failures no dataset anticipated.
The two feed each other:
+---------------------- OFFLINE ----------------------+
| dataset -> experiment -> compare -> promote |
+------+--------------------------------------+-------+
^ |
add failing cases deploy
| v
+------+--------------- ONLINE ---------------+-------+
| live traces -> scores -> find regressions |
+-----------------------------------------------------+A production failure becomes a dataset item, which becomes a regression test that every future candidate must pass. Over time the dataset accumulates exactly the cases your application gets wrong — which is what makes offline evaluation worth trusting.
How the pieces fit
Score config defines what a score means: type, range, categories
|
+-- Annotation queue -> queue items -> reviewer scores -> Score (HUMAN_ANNOTATION)
+-- Eval template -> eval config -> eval job -> Score (EVAL)
+-- Dataset -> dataset items -> experiment run -> Score (EVAL)
\-- External pipeline posts to the collector -> Score (API)A score config is the vocabulary. Everything else is a way of producing scores in that vocabulary, so results from a reviewer, a judge model, and your own pipeline stay comparable.
Score
Scores are the universal object for evaluation results in KubeSense. Every method — judge, reviewer, experiment, or your own code — ends by writing one, which is what lets their results be compared on the same axes.
A score has a name, a value, and a data type:
| Data type | Value stored |
|---|---|
NUMERIC | A number, such as 0.92 |
CATEGORICAL | A label, such as helpful |
BOOLEAN | True or false |
Its source records where it came from:
| Source | Produced by |
|---|---|
API | Posted directly by an external pipeline — the default when no source is given |
EVAL | The LLM-as-a-Judge engine or an experiment run |
HUMAN_ANNOTATION | A reviewer working an annotation queue |
A score attaches to:
| Target | When |
|---|---|
| A trace | The result describes the whole request |
| An observation | The result describes one generation or step inside the request |
| An experiment run | The result describes a configuration rather than a single response — the run-level scores in Experiments |
The session a score belongs to is resolved from its trace, so session-level filtering works without setting anything extra.
A score can also carry a comment, its author, and links to the score config, annotation queue, or experiment it came from. Scores are versioned rather than overwritten: an update or delete is a new record, and readers see only the newest non-deleted version.
Score config
A score config defines a score that reviewers may apply — its name, data type, and permitted range or categories. A numeric config sets a minimum and maximum; a categorical one lists its labels and their underlying values.
Configs are what keep human annotation consistent: without one, two reviewers can record the same judgement on different scales. Configs can be archived when retired, which preserves historical scores while removing the option from new annotations.
Evaluation methods
| Method | What it does | Runs |
|---|---|---|
| LLM as a Judge | An LLM scores output against a rubric you write | Online |
| Annotation queues | Structured human review with a fixed set of score configs | Online |
| Scores via API | Your own code posts results to the collector | Online |
| Experiments | A prompt and model are run across a dataset and scored | Offline |
Which to reach for:
| Method | Best for |
|---|---|
| LLM as a Judge | Applying a repeatable rubric across production traffic at a volume no one could read |
| Annotation | Nuanced judgement, policy calls, and producing the ground truth that calibrates a judge |
| API score | A metric your application already computes — a deterministic check, a business outcome, or explicit user feedback such as a thumbs-up |
| Experiment | Comparing prompts, models, or application versions on identical inputs before shipping |
Use more than one. A judge finds likely regressions at scale; annotation tells you whether the judge is right; an experiment tells you whether a fix actually helps.
note: User feedback is not a separate object. Record a thumbs-up, a rating, or an accepted-suggestion signal as an API score on the trace, using a score config that defines what the value means. It then aggregates and filters alongside every other score.
Not currently supported
Coming from Langfuse, three things are absent:
| Not available | Nearest alternative |
|---|---|
| Code evaluators — running Python or TypeScript scoring logic inside KubeSense | Run the logic in your own pipeline and post the result as an API score |
| Ad-hoc scoring in the UI — scoring a trace directly from the trace view | Send the trace to an annotation queue; the trace view's annotate action adds it to a queue rather than scoring in place |
| Batch or retroactive evaluation of traces already ingested | Evaluators score spans as they arrive. A new evaluator applies only to traffic received after it is enabled; to evaluate historical cases, add them to a dataset and run an experiment |
LLM-as-a-Judge
Automated scoring uses three objects. If you are coming from Langfuse, the template is its evaluator (the scoring logic) and the config is its rule (the filter that decides which observations get evaluated).
Eval template — the judge prompt. It holds the prompt text, the model and model parameters to run it with, the variables it expects, and the output schema describing the scores it returns. Templates are versioned.
Eval config — binds a template to live traffic. It sets which observations to evaluate through a filter, how template variables map to fields on the span, what to name the resulting score, and a sampling rate between 0 and 1 so you can judge a fraction of matching traffic. A config can be enabled or disabled, and is blocked automatically if it fails repeatedly.
Eval job — one execution against one observation. It records status (PENDING, RUNNING, COMPLETED, FAILED, CANCELLED), the resulting score and the judge's reasoning, plus the tokens and cost the judge itself consumed — judging is a model call and has its own spend.
Evaluation runs inside the collector as spans arrive, dispatched before the write to ClickHouse so a slow provider never delays ingestion. Work is queued to a worker pool and dropped when the queue is full; GET /v1/llm/eval/metrics exposes the counters, and a non-zero dropped in steady state means the pool is undersized for your span rate.
Eval configs target individual observations. A shadow mode is available for validating a configuration without writing scores.
Dataset and dataset item
A dataset is a fixed collection of test cases used for repeatable offline comparison. It can define an input schema and an expected-output schema so items stay well-formed.
A dataset item is one case: an input, an optional expected output, and metadata. An item created from a real request keeps a link back to its source trace and observation, which is the usual way a production failure becomes a regression test.
Items are archived rather than deleted, so historical runs remain interpretable.
Experiment
An experiment — a dataset run — executes a task across every item in a dataset. The task is the thing under test: in KubeSense it is a prompt version plus a model configuration, defined on the run itself rather than in your application code. It records status, item counts, and any error, then aggregates the results: total and average cost, token usage, average latency, and score aggregates.
Scores are aggregated at two levels. Trace-level scores come from the traces the run produced; run-level scores describe the run as a whole.
An experiment item links one dataset item to the trace produced for it, along with a snapshot of the input and expected output as they were at run time. The snapshot matters: editing a dataset item later does not rewrite what an earlier run was actually tested against.
Experiments are how you compare a candidate prompt or model against the current one on identical inputs, instead of inferring quality from production traffic alone.
Annotation queue
An annotation queue is a review workflow for human scoring. It holds the items to review — whole traces, or individual spans when only one step needs judging — the score configs reviewers may apply, and the users assigned to it, and reports pending and completed counts.
Traces are added individually from a trace view or in bulk by filter. Each queue item tracks its status — PENDING, COMPLETED, or SKIPPED — along with who completed it and when. See Annotations for the review screen.
Human annotation is how you establish ground truth: it calibrates automated judges and captures the quality signals no metric expresses.
Comments
Comments are threaded notes on a trace, observation, session, or prompt, with emoji reactions. They keep investigation context attached to the object rather than in a separate tracker.
Evaluation data handling
Evaluation prompts can include trace input, output, model, user, and other mapped fields. Limit evaluator matches and avoid sending sensitive content to a third-party judge provider unless that processing is approved for your environment.
Setup order
Before adding a custom evaluator, configure the scoring definitions that your team will use to interpret its results. This follows the same general pattern as Langfuse: define what a score means, then define the judge prompt and the data it evaluates.
Complete the setup in this order:
- Create the LLM application and, when required, its LLM connection in Settings > LLM.
- Create one or more score configurations in Score Configs.
- Create the custom evaluator in Custom Evaluators.
- Configure the evaluator's model, prompt, score instructions, target filters, and variable mappings.
- Enable it for live incoming traces or use it with datasets and experiments.
Score Configs define the shape and meaning of a result. Custom Evaluators define how the judge model produces that result. Keeping these concerns separate lets the same score dimensions be used consistently for manual review, automated evaluation, and analysis. In KubeSense, annotation queues explicitly link to Score Configs, while a Custom Evaluator also stores the score name and output instructions that it emits.
A worked loop
The offline and online halves in practice:
- Identify a regression from dashboard or trace data.
- Add representative traces to a dataset.
- Annotate a sample to establish expected behavior.
- Run prompt or model experiments against the dataset.
- Deploy the candidate and monitor production scores.
- Review false positives and update the rubric or dataset.
Where to next
| Goal | Page |
|---|---|
| Define what a score means, or post one | Scores |
| Apply a rubric automatically to live traffic | LLM as a Judge |
| Have people review traces | Annotations |
| Keep a fixed set of test cases | Datasets |
| Compare a prompt or model across a dataset | Experiments |