Kubesense

Concepts

Observability is the part of KubeSense that receives your telemetry and lets you search it: applications, traces, observations, users, sessions, and the usage and cost attached to them.

This page defines each object and how they relate. Prompt Hub and Evaluation have their own concepts pages for the objects you create in the dashboard.

The data model at a glance

Telemetry is organized as a hierarchy. Everything below an application is derived from the spans your SDK exports.

Application                     the workspace for one LLM workload
├── Trace                       one request, conversation turn, or agent run
│   └── Observation             one operation inside the trace
│       ├── Generation          a model call, with usage and cost
│       ├── Tool / Agent        an action or an orchestration step
│       ├── Retriever           a vector or knowledge lookup
│       └── Span / Event        any other unit of work
├── Session                     related traces, grouped into a conversation
├── User                        who or what triggered the traces
└── Score                       a quality, policy, or business result

Grouping objects sit beside the hierarchy rather than inside it. A session groups traces, a user groups sessions, and a score attaches to a trace or a single observation.

Application

An application is the container for one LLM workload. It scopes traces, dashboards, prompts, LLM connections, datasets, evaluators, and annotation queues — two teams sharing one KubeSense installation each get their own application and never see each other's data.

Every application has a UUID. Your instrumented service must send it as the application.id resource attribute (project.id is an accepted alias) so incoming spans are filed correctly. Spans without it are still accepted, but land in the default application.

Applications also carry a retention time and the set of clusters and environments seen in their traces.

Create one in Settings > LLM > Applications. See Installation for the setup order.

Trace

A trace is one complete unit of work: a user request, a conversation turn, or an agent run. It is not a stored record of its own — KubeSense stores spans, and a trace is every span sharing the same trace_id, assembled at query time.

A trace has a root observation and any number of children linked by parent span ID. One trace commonly contains several model calls plus the retrieval, tool, and application work around them, which is why trace count is not the same as model-call count.

Trace-level context — user, session, tags, environment — lives on the observations themselves. The Langfuse SDKs propagate those values onto every span in the block so aggregation and filtering work at any level.

Observation

An observation is one operation inside a trace and the unit KubeSense actually stores: one row in kube_llm.llm_span. Every span is kept, whether or not it carries GenAI attributes, so the surrounding HTTP, database, and framework work stays available as request context.

Each observation records its identity (trace ID, span ID, parent span ID, name, kind), its timing (start, end, duration), its status, and whichever LLM fields the SDK supplied. Attributes KubeSense does not map to a column are preserved in the span's attributes map, so newer instrumentation is never silently dropped.

Observation types

The type comes from the langfuse.observation.type attribute, which the Langfuse SDKs set from as_type in Python and asType in JavaScript:

TypeWhat it represents
spanA unit of work inside a trace — the general-purpose default
generationA model request and its response
eventA discrete point-in-time occurrence
agentA step that decides application flow, usually orchestrating tools
toolA single action such as a function or API call
chainA link between steps, such as retriever output feeding a model call
retrieverA data-retrieval step that queries a knowledge source without changing state
evaluatorA step that assesses the relevance or correctness of an output
embeddingA model call that generates embeddings
guardrailA check that protects against malicious or disallowed content

Spans from OpenLIT and other OpenTelemetry instrumentation do not carry this attribute. KubeSense classifies them from the GenAI attributes instead — gen_ai.operation.name for the operation, gen_ai.type for the category.

Generation

A generation is an observation representing a model call. Beyond the common span fields it can record:

  • Provider and the requested versus returned model — these differ when a provider routes an alias to a concrete version
  • Input messages, output text, and finish reasons
  • Token usage and cost
  • Request parameters: temperature, top-p, top-k, max tokens, seed, frequency and presence penalty
  • Streaming state and time to first token
  • Response ID and system fingerprint
  • Tool calls and tool definitions

Time to first token is the delay before the model produced anything. Compared against total duration it separates provider think-time from generation time on a streamed response.

Session

See Sessions for instrumentation.

A session groups related traces into one conversation, workflow run, or agent execution. Set a stable session_id for the lifetime of the interaction and a new trace for each turn.

Because a session spans multiple traces, its duration and total cost can exceed any individual trace. Sessions are what make a multi-turn conversation readable in order rather than as scattered requests.

User

See Users for instrumentation.

A user is the person, tenant, or service account responsible for an interaction, identified by user_id. It drives the per-user cost, token, and traffic breakdowns.

One user has many sessions; one session has many traces. Observations that arrive with no user attribute are grouped under unknown rather than dropped.

Use a stable internal identifier. Email addresses, access tokens, and full profiles do not belong in this field.

Conversation

A conversation ID (gen_ai.conversation.id) is a separate thread identifier for cases where a durable multi-turn thread outlives a single session. Most applications need only users and sessions.

Enrichment attributes

These optional fields make traces filterable and comparable.

Environment

The deployment context — prod, staging, dev. It keeps test traffic out of production dashboards and is the first filter to check when a metric looks wrong.

KubeSense reads it from langfuse.environment, gen_ai.environment, env, or kubesense.env, falling back to the OpenTelemetry deployment.environment. Observations with no environment attribute are recorded as default.

Tags

Free-form labels on a trace, used to categorize by feature, endpoint, customer tier, or experiment. Tags are for slicing traffic; they are not a place for per-request values such as an ID.

Metadata and attributes

Metadata is structured context the SDK attached to the observation, together with the OpenTelemetry resource attributes and instrumentation scope. Attributes are every incoming key KubeSense did not promote to a dedicated column.

Between them, nothing the SDK exported is lost — the trace detail view exposes both.

Level and status

Level is the severity the SDK assigned to an observation: DEBUG, DEFAULT, WARNING, or ERROR. It is what the dashboard's Observations by Level panel counts, and it is set by the application, so a handled fallback can be recorded as WARNING without failing the span.

Status is separate and comes from the OpenTelemetry span status: ok, error, or unset, with a status message when one was reported. Error rate is computed from status, not level.

Version and release

service.version records which build produced the trace. Langfuse additionally emits langfuse.release from LANGFUSE_RELEASE. Both let you compare behaviour before and after a deployment.

Source

Which SDK produced the span: langfuse, openlit, or otel. It is derived from resource attributes only — see Trace source for how the label is determined and why Langfuse traffic is labelled otel by default.

Usage and cost

Token usage

KubeSense records usage in separate buckets rather than one number, because they are priced differently:

BucketMeaning
InputPrompt tokens
OutputCompletion tokens
ReasoningThinking tokens on reasoning models
Cache readPrompt tokens served from a provider cache, billed at a reduced rate
Cache creationTokens spent writing a cache entry
Input audio / output audioAudio tokens on speech-capable models
ImageImage tokens for multimodal input

When the SDK does not report a total, KubeSense derives it by summing the buckets.

Cost

Cost is recorded per bucket and as a total, in USD.

Costs reported by the SDK always win. When a span carries usage but no cost, kubecol resolves the model in the pricing registry and calculates input, output, cache, and reasoning cost from the recorded tokens. A span whose model has no pricing entry is still stored, but with zero cost — which understates spend rather than hiding the request.

All costs are estimates. They are incomplete when a provider returns no token counts, when the model name is unrecognized, or when the SDK omits cache and reasoning usage.

Model registry and pricing

A model entry in Settings > LLM > Models maps an incoming model name to prices. Each entry has:

  • A match pattern: a case-insensitive regular expression tested against the model name on the span. Entries are tried longest-model-name first and the first match wins, so anchor patterns with ^ and $.
  • A unit, such as TOKENS for text or IMAGES for image generation.
  • One or more pricing tiers, each with a priority and its own prices. Cost enrichment uses the tier marked default, or the highest-priority tier when none is marked. Tier conditions are stored but are not evaluated during ingestion today, so the default tier is the one that determines cost.
  • Prices within a tier, one per usage type: input, output, total, cache read, cache creation, audio, image.

When a tier is applied, its ID and name are recorded on the span, so a cost figure can always be traced back to the rule that produced it. Cost is calculated once at ingestion and stored, so a later price change does not restate historical spend. See Installation for how to add or edit an entry.

How data gets in

Your application creates spans through the Langfuse or OpenLIT SDK and exports them over OTLP/HTTP. kubecol maps GenAI and langfuse.* attributes onto the span model, fills in missing costs from the pricing registry, dispatches to the eval engine, and writes to ClickHouse in batches. kubeapi reads the stored spans for every LLM view.

Application + SDK -> OTLP/HTTP -> kubecol -> ClickHouse (kube_llm) -> kubeapi -> KubeSense LLM views

Both SDKs batch spans and export in the background, so instrumentation does not sit in your request path. The trade-off is that a process which exits immediately can lose its final batch: short-lived jobs, scripts, and serverless handlers must flush before exiting. See Installation for the flush call for each SDK.

Ingestion is append-only. Traces, observations, and their content are never rewritten after the fact; scores are the one object that supports update and delete, expressed as new versions.

Where to next

GoalPage
Monitor traffic, latency, and spendOverview
Inspect a request or agent runTraces
Break usage down by customerUsers
Follow a multi-turn conversationSessions
Send your first tracesInstallation