Kubesense

LLM as a Judge

LLM-as-a-Judge applies a rubric to live traffic automatically. You write a judge prompt, bind it to the observations you care about, and KubeSense runs it as spans arrive and writes the result as a score with source EVAL.

It is the only evaluation method that scales to all production traffic. It is also the least trustworthy in isolation — calibrate it against Annotations before acting on its output.

See Concepts for the template, config, and job objects.

How it works

The judge is given three things: the input your application received, the output it produced, and a rubric describing what "good" means. It returns a score in the shape you asked for — a number, a category, or a boolean — together with its reasoning.

That reasoning is the point. A bare number tells you a response scored 0.7; the reasoning tells you why, which is what you need to decide whether the judge is right or the rubric is vague.

Three properties make this worth doing:

  • Scale. A rubric can be applied to every request; people cannot.
  • Nuance. Dimensions like relevance, tone, and faithfulness have no deterministic check.
  • Repeatability. A fixed rubric judges Tuesday's traffic the same way it judged Monday's, which is what makes a trend line meaningful.

The cost is that a judge is itself a model, with a model's failure modes. It is a measurement instrument that needs calibrating against human annotation before you trust it.

The evaluators list

LLM Monitoring > Evaluation > LLM as a Judge lists the evaluators configured for the selected application.

LLM as a Judge — evaluators

ColumnMeaning
NameThe evaluator configuration, such as KubeSense SRE
EvaluatorThe rubric it runs — the template, such as Correctness or Conciseness
StatusEnabled or disabled; a disabled evaluator stays configured but stops running
TargetWhat it evaluates. Spans — individual observations, not whole traces
Total CostWhat the judge itself has spent over the recent window
ResultThe most recent outcome
FilterHow many filter conditions narrow the traffic it matches
Created / UpdatedConfiguration history

The split between the Name and Evaluator columns is the split between the two objects behind them: one rubric can back several configurations, each pointed at different traffic with its own filter and sampling rate.

Total Cost is the column to watch after enabling something new. Judging is a model call per matched span, so a broad filter at full sampling can outspend the traffic it measures.

Use New Evaluator to create one.

Before you start

An evaluator needs two things that must exist first:

  1. An LLM connection for the application, so KubeSense can call the judge model. See Installation.
  2. A score config defining the dimension the judge emits, so its results share a vocabulary with human review. See Scores.

Create a custom evaluator

After the score config exists, open Settings > LLM > Custom Evaluators and select Create Evaluator. Select the application from the Project selector.

LLM Custom Evaluators settings page

Enter an evaluator name, select the judge model from an LLM connection, and write the evaluation prompt. The prompt should include the criterion, the input and output to evaluate, and clear instructions for the expected score. Use variables such as {{input}} and {{output}} to map trace data into the prompt.

Create Custom Evaluator form

The evaluator form supports:

  • Name: A stable name for the evaluation rule, such as Answer Relevance.
  • Model: The judge model from the application's configured LLM connections.
  • Evaluation Prompt: The rubric and the content being evaluated.
  • Score Reasoning Prompt: Instructions for a concise explanation of the result.
  • Score Range Prompt: Instructions that constrain the returned score to the configured range or categories.
  • Target and filters: Whether the evaluator runs on observations or dataset items, and which service, model, operation, environment, or application spans match.
  • Variable mappings: Which trace, span, generation, tool, retriever, or dataset item fields populate each prompt variable.

Create the Score Config before creating the evaluator so the evaluator's score name, type, range, and categories match the dimensions used elsewhere in the project. For annotation workflows, the Score Config must exist before it can be attached to an annotation queue. The Custom Evaluator form defines its own emitted score name and instructions, so make those values agree with the Score Config instead of creating unrelated names.

Walkthrough

Writing a judge prompt

The rubric decides whether the scores mean anything. A few things matter more than the rest:

  • Define the scale explicitly. "Rate helpfulness 0 to 1" invites drift. Say what 0 means, what 1 means, and what sits between them.
  • Give the judge only what it needs. Map the input and output it must read; extra context invites the judge to grade something you did not ask about.
  • Ask for reasoning. It lands in the score's comment and is the only way to audit a disagreement later.
  • Cover the borderline cases. Include an example of a response that just passes and one that just fails — that is where a vague rubric produces noise.
  • Use a capable judge model. A small model is cheap per call and expensive in wrong conclusions. Judge quality bounds everything downstream.

Constrain the output to the score config's type and range. A numeric config with a 0–1 range and a judge that occasionally answers "high" produces gaps in your data rather than an error.

How evaluation runs

Evaluation happens inside the collector, as spans arrive. A matching span is dispatched to a worker pool before the write to ClickHouse, so a slow judge provider can never delay ingestion — the trace lands regardless of what the judge does.

Work is queued, and the queue drops when full rather than blocking. A dropped evaluation means no score was produced for that span; the span itself is unaffected.

Evaluators target individual observations. Use the filters to match only spans that actually carry what the rubric needs — a rubric that reads input and output should not be matched against retrieval or tool spans that have neither.

Define the score type and range before writing the evaluator prompt. Make the rubric explicit, require structured output, and include examples for borderline cases. Match only observations that contain the fields required by the evaluator.

For example, an evaluator can match gen_ai_operation=chat, a particular model, and a production environment. It can then use the trace input and output to produce a quality score without evaluating retrieval or tool spans that do not contain a response.

Operating asynchronous evaluation

The evaluator runs after the incoming span has been accepted. Watch the evaluation metrics endpoint for enqueued, succeeded, failed, and dropped counts. A failed judge call affects the score, while a dropped evaluation indicates queue or worker capacity pressure; neither should be confused with loss of the original trace.

Sizing rule of thumb for the collector's eval_engine block:

queue_size ≳ workers × per_call_timeout_sec × span_rate_per_sec × config_match_rate

Lower the sampling rate on the eval config before enlarging the pool — judging every span is rarely worth the spend.

Cost

Judging is a model call and bills like one. Each job records the tokens and cost it consumed, so evaluator spend is visible separately from application spend. A broad filter at full sampling can cost more than the traffic it evaluates.

Data handling

Evaluation prompts can include trace input, output, model, user, and other mapped fields. Limit evaluator matches and avoid sending sensitive content to a third-party judge provider unless that processing is approved for your environment.

Where to next

  • Annotations — establish ground truth to calibrate the rubric
  • Scores — where judge results land
  • Experiments — apply a rubric to a dataset instead of live traffic