Kubesense

Concepts

LLM output has no single correct answer, so quality cannot be asserted once and forgotten. Evaluation is the machinery for measuring it continuously: attaching quality signals to what your application produced, and comparing candidates before they ship.

This page defines the objects. Each has its own page: Scores, LLM as a Judge, Annotations, Datasets, and Experiments.

The evaluation loop

Evaluation runs in two places, and a healthy team uses both.

Offline evaluation happens before deployment. You run a candidate prompt or model across a fixed dataset as an experiment and compare the result against the current configuration on identical inputs. It is repeatable and cheap to re-run, but it only covers cases you thought to include.

Online evaluation happens in production. Scores are attached to live traces — by a judge model, a reviewer, or your own pipeline — so you find the failures no dataset anticipated.

The two feed each other:

        +---------------------- OFFLINE ----------------------+
        |  dataset  ->  experiment  ->  compare  ->  promote  |
        +------+--------------------------------------+-------+
               ^                                      |
       add failing cases                           deploy
               |                                      v
        +------+--------------- ONLINE ---------------+-------+
        |  live traces  ->  scores  ->  find regressions      |
        +-----------------------------------------------------+

A production failure becomes a dataset item, which becomes a regression test that every future candidate must pass. Over time the dataset accumulates exactly the cases your application gets wrong — which is what makes offline evaluation worth trusting.

How the pieces fit

Score config      defines what a score means: type, range, categories
   |
   +-- Annotation queue -> queue items -> reviewer scores  ->  Score (HUMAN_ANNOTATION)
   +-- Eval template -> eval config -> eval job            ->  Score (EVAL)
   +-- Dataset -> dataset items -> experiment run          ->  Score (EVAL)
   \-- External pipeline posts to the collector            ->  Score (API)

A score config is the vocabulary. Everything else is a way of producing scores in that vocabulary, so results from a reviewer, a judge model, and your own pipeline stay comparable.

Score

Scores are the universal object for evaluation results in KubeSense. Every method — judge, reviewer, experiment, or your own code — ends by writing one, which is what lets their results be compared on the same axes.

A score has a name, a value, and a data type:

Data typeValue stored
NUMERICA number, such as 0.92
CATEGORICALA label, such as helpful
BOOLEANTrue or false

Its source records where it came from:

SourceProduced by
APIPosted directly by an external pipeline — the default when no source is given
EVALThe LLM-as-a-Judge engine or an experiment run
HUMAN_ANNOTATIONA reviewer working an annotation queue

A score attaches to:

TargetWhen
A traceThe result describes the whole request
An observationThe result describes one generation or step inside the request
An experiment runThe result describes a configuration rather than a single response — the run-level scores in Experiments

The session a score belongs to is resolved from its trace, so session-level filtering works without setting anything extra.

A score can also carry a comment, its author, and links to the score config, annotation queue, or experiment it came from. Scores are versioned rather than overwritten: an update or delete is a new record, and readers see only the newest non-deleted version.

Score config

A score config defines a score that reviewers may apply — its name, data type, and permitted range or categories. A numeric config sets a minimum and maximum; a categorical one lists its labels and their underlying values.

Configs are what keep human annotation consistent: without one, two reviewers can record the same judgement on different scales. Configs can be archived when retired, which preserves historical scores while removing the option from new annotations.

Evaluation methods

MethodWhat it doesRuns
LLM as a JudgeAn LLM scores output against a rubric you writeOnline
Annotation queuesStructured human review with a fixed set of score configsOnline
Scores via APIYour own code posts results to the collectorOnline
ExperimentsA prompt and model are run across a dataset and scoredOffline

Which to reach for:

MethodBest for
LLM as a JudgeApplying a repeatable rubric across production traffic at a volume no one could read
AnnotationNuanced judgement, policy calls, and producing the ground truth that calibrates a judge
API scoreA metric your application already computes — a deterministic check, a business outcome, or explicit user feedback such as a thumbs-up
ExperimentComparing prompts, models, or application versions on identical inputs before shipping

Use more than one. A judge finds likely regressions at scale; annotation tells you whether the judge is right; an experiment tells you whether a fix actually helps.

note: User feedback is not a separate object. Record a thumbs-up, a rating, or an accepted-suggestion signal as an API score on the trace, using a score config that defines what the value means. It then aggregates and filters alongside every other score.

Not currently supported

Coming from Langfuse, three things are absent:

Not availableNearest alternative
Code evaluators — running Python or TypeScript scoring logic inside KubeSenseRun the logic in your own pipeline and post the result as an API score
Ad-hoc scoring in the UI — scoring a trace directly from the trace viewSend the trace to an annotation queue; the trace view's annotate action adds it to a queue rather than scoring in place
Batch or retroactive evaluation of traces already ingestedEvaluators score spans as they arrive. A new evaluator applies only to traffic received after it is enabled; to evaluate historical cases, add them to a dataset and run an experiment

LLM-as-a-Judge

Automated scoring uses three objects. If you are coming from Langfuse, the template is its evaluator (the scoring logic) and the config is its rule (the filter that decides which observations get evaluated).

Eval template — the judge prompt. It holds the prompt text, the model and model parameters to run it with, the variables it expects, and the output schema describing the scores it returns. Templates are versioned.

Eval config — binds a template to live traffic. It sets which observations to evaluate through a filter, how template variables map to fields on the span, what to name the resulting score, and a sampling rate between 0 and 1 so you can judge a fraction of matching traffic. A config can be enabled or disabled, and is blocked automatically if it fails repeatedly.

Eval job — one execution against one observation. It records status (PENDING, RUNNING, COMPLETED, FAILED, CANCELLED), the resulting score and the judge's reasoning, plus the tokens and cost the judge itself consumed — judging is a model call and has its own spend.

Evaluation runs inside the collector as spans arrive, dispatched before the write to ClickHouse so a slow provider never delays ingestion. Work is queued to a worker pool and dropped when the queue is full; GET /v1/llm/eval/metrics exposes the counters, and a non-zero dropped in steady state means the pool is undersized for your span rate.

Eval configs target individual observations. A shadow mode is available for validating a configuration without writing scores.

Dataset and dataset item

A dataset is a fixed collection of test cases used for repeatable offline comparison. It can define an input schema and an expected-output schema so items stay well-formed.

A dataset item is one case: an input, an optional expected output, and metadata. An item created from a real request keeps a link back to its source trace and observation, which is the usual way a production failure becomes a regression test.

Items are archived rather than deleted, so historical runs remain interpretable.

Experiment

An experiment — a dataset run — executes a task across every item in a dataset. The task is the thing under test: in KubeSense it is a prompt version plus a model configuration, defined on the run itself rather than in your application code. It records status, item counts, and any error, then aggregates the results: total and average cost, token usage, average latency, and score aggregates.

Scores are aggregated at two levels. Trace-level scores come from the traces the run produced; run-level scores describe the run as a whole.

An experiment item links one dataset item to the trace produced for it, along with a snapshot of the input and expected output as they were at run time. The snapshot matters: editing a dataset item later does not rewrite what an earlier run was actually tested against.

Experiments are how you compare a candidate prompt or model against the current one on identical inputs, instead of inferring quality from production traffic alone.

Annotation queue

An annotation queue is a review workflow for human scoring. It holds the items to review — whole traces, or individual spans when only one step needs judging — the score configs reviewers may apply, and the users assigned to it, and reports pending and completed counts.

Traces are added individually from a trace view or in bulk by filter. Each queue item tracks its status — PENDING, COMPLETED, or SKIPPED — along with who completed it and when. See Annotations for the review screen.

Human annotation is how you establish ground truth: it calibrates automated judges and captures the quality signals no metric expresses.

Comments

Comments are threaded notes on a trace, observation, session, or prompt, with emoji reactions. They keep investigation context attached to the object rather than in a separate tracker.

Evaluation data handling

Evaluation prompts can include trace input, output, model, user, and other mapped fields. Limit evaluator matches and avoid sending sensitive content to a third-party judge provider unless that processing is approved for your environment.

Setup order

Before adding a custom evaluator, configure the scoring definitions that your team will use to interpret its results. This follows the same general pattern as Langfuse: define what a score means, then define the judge prompt and the data it evaluates.

Complete the setup in this order:

  1. Create the LLM application and, when required, its LLM connection in Settings > LLM.
  2. Create one or more score configurations in Score Configs.
  3. Create the custom evaluator in Custom Evaluators.
  4. Configure the evaluator's model, prompt, score instructions, target filters, and variable mappings.
  5. Enable it for live incoming traces or use it with datasets and experiments.

Score Configs define the shape and meaning of a result. Custom Evaluators define how the judge model produces that result. Keeping these concerns separate lets the same score dimensions be used consistently for manual review, automated evaluation, and analysis. In KubeSense, annotation queues explicitly link to Score Configs, while a Custom Evaluator also stores the score name and output instructions that it emits.

A worked loop

The offline and online halves in practice:

  1. Identify a regression from dashboard or trace data.
  2. Add representative traces to a dataset.
  3. Annotate a sample to establish expected behavior.
  4. Run prompt or model experiments against the dataset.
  5. Deploy the candidate and monitor production scores.
  6. Review false positives and update the rubric or dataset.

Where to next

GoalPage
Define what a score means, or post oneScores
Apply a rubric automatically to live trafficLLM as a Judge
Have people review tracesAnnotations
Keep a fixed set of test casesDatasets
Compare a prompt or model across a datasetExperiments