Kubesense

Evaluation

Evaluation attaches quality signals to LLM traffic and compares changes before you ship them.

PageWhat it covers
ConceptsHow scores, judges, datasets, and queues fit together
ScoresScore configs, data types, sources, and posting your own
LLM as a JudgeApplying a rubric automatically to live traffic
AnnotationsHuman review queues and ground truth
DatasetsFixed collections of test cases
ExperimentsRunning a prompt and model across a dataset

Every method converges on the same object: a score attached to a trace or observation.