Evaluation
Evaluation attaches quality signals to LLM traffic and compares changes before you ship them.
| Page | What it covers |
|---|---|
| Concepts | How scores, judges, datasets, and queues fit together |
| Scores | Score configs, data types, sources, and posting your own |
| LLM as a Judge | Applying a rubric automatically to live traffic |
| Annotations | Human review queues and ground truth |
| Datasets | Fixed collections of test cases |
| Experiments | Running a prompt and model across a dataset |
Every method converges on the same object: a score attached to a trace or observation.