LLM Observability
KubeSense LLM Observability traces LLM applications from the SDK call to the model response. It records the request pipeline, model and provider, token usage, latency, errors, sessions, and estimated cost.
LLM applications are different from conventional services: one user request may involve several model calls, retrieval steps, tool executions, and agent decisions. KubeSense preserves that complete workflow so teams can understand not only whether a request was slow or expensive, but also which operation caused it.
What LLM Observability helps you answer
- Which models and providers are receiving the most traffic?
- Which users, sessions, or environments are driving token usage and cost?
- Are errors coming from the model call, retrieval, tools, or application code?
- How does generation latency compare with total request latency?
- Which prompts, model versions, or releases need quality review?
- Which responses should be added to a dataset or annotation queue?
KubeSense is designed for the complete loop from instrumentation to investigation and evaluation:
Instrument -> Ingest -> Explore -> Understand -> Evaluate -> ImproveThe supported ingestion path is:
LLM application
-> Langfuse SDK or OpenLIT SDK
-> OTLP/HTTP
-> kubecol LLM receiver
-> ClickHouse (kube_llm.llm_span)
-> kubeapi
-> KubeSense LLM viewsSupported SDKs
Both supported SDKs are OpenTelemetry-based and export OTLP/HTTP directly to kubecol. There is no KubeSense-specific SDK to install.
| SDK | Package | Export path | Notes |
|---|---|---|---|
| Langfuse Python | langfuse (v4) | {LANGFUSE_BASE_URL}/api/public/otel/v1/traces | Set LANGFUSE_BASE_URL to the collector base URL |
| Langfuse JS/TS | @langfuse/tracing, @langfuse/otel (v5) | {baseUrl}/api/public/otel/v1/traces | Registered as a span processor on the OpenTelemetry NodeSDK |
| OpenLIT Python | openlit | {OTEL_EXPORTER_OTLP_ENDPOINT}/v1/traces | Auto-instruments supported providers and frameworks |
| OpenLIT JS/TS | openlit | {otlpEndpoint}/v1/traces | Auto-instruments supported providers and frameworks |
| Any OTel SDK | – | /v1/traces | GenAI attributes are mapped; other spans are stored as trace context |
The pre-OpenTelemetry Langfuse clients — Python SDK v2 and the legacy langfuse npm package — send data to the Langfuse ingestion API (/api/public/ingestion) rather than OTLP. kubecol does not implement that API, so those SDK generations cannot send data to KubeSense.
KubeSense implements the OTLP ingestion surface, not the Langfuse management API. Prompt fetching, dataset APIs, and auth_check() calls made against LANGFUSE_BASE_URL do not resolve.
LLM modules
The section is organized into three modules plus setup.
Installation — configure Langfuse or OpenLIT and send traces to KubeSense.
Observability — what your application sends, and how you search it.
- Concepts - Applications, traces, observations, usage, and cost
- Overview - Traffic, latency, model usage, and spend
- Traces - Span trees, prompts, outputs, and errors
- Users - Usage and cost by person, tenant, or service account
- Sessions - Traces grouped into conversations and agent runs
Prompt Hub — author, version, and test prompts.
- Concepts - Prompts, versions, labels, tools, and schemas
- Prompts - Creating prompts and managing versions
- Playground - Testing a version against a model
Evaluation — attach quality signals and compare changes.
- Concepts - How scores, judges, datasets, and queues fit together
- Scores - Score configs, data types, sources, and posting your own
- LLM as a Judge - Applying a rubric automatically to live traffic
- Annotations - Human review queues and ground truth
- Datasets - Fixed collections of test cases
- Experiments - Running a prompt and model across a dataset
Core concepts
Six objects carry most of the model:
| Object | What it is |
|---|---|
| Application | The workspace for one LLM workload; every span carries its ID |
| Trace | One request, conversation turn, or agent run |
| Observation | One operation inside a trace — a generation, retrieval, tool call, or agent step |
| Session | Related traces grouped into a conversation or workflow run |
| User | The person, tenant, or service account behind an interaction |
| Score | A quality, policy, or business result attached to a trace or observation |
Observability concepts defines these in full, along with observation types, levels, token and cost buckets, prompts and labels, datasets, experiments, evaluators, and annotation queues.
Supported workloads
KubeSense can capture common LLM and AI pipeline operations, including:
| Workload | Examples |
|---|---|
| Model calls | Chat, completion, streaming, image generation, speech, and multimodal requests |
| AI components | Embeddings, retrieval, vector database queries, and reranking |
| Agent workflows | Agent steps, chains, tool calls, and multi-agent operations |
| Providers | OpenAI, Anthropic, Google, Mistral, and other providers represented by the SDK attributes |
| Frameworks | LangChain, LlamaIndex, and other OpenTelemetry-instrumented frameworks |
Provider and framework support depends on what Langfuse or OpenLIT instruments and exports. KubeSense accepts standard OTLP data, so custom or newer integrations can still be ingested when they follow the OpenTelemetry trace format.
What KubeSense records
LLM spans can include model requests and responses, embeddings, vector database operations, framework steps, tool calls, and agent operations. KubeSense also keeps the surrounding OTel spans so the complete request context remains available.
Common fields include:
| Group | Examples |
|---|---|
| Application | Application ID, service name, version, environment, cluster, namespace |
| Operation | Provider, operation (chat, embeddings, execute_tool), request and response model |
| Usage | Input, output, reasoning, cached, audio, and image tokens |
| Performance | Start and end time, duration, time to first token, streaming state |
| Result | Input and output content, finish reasons, response ID, status and errors |
| Context | User, session, conversation, tags, metadata, tools, and attributes |
KubeSense uses reported SDK costs when available. Otherwise, kubecol calculates cost from the model pricing registry and the recorded token counts.
Data lifecycle
- Your application creates request and child spans through Langfuse, OpenLIT, or OpenTelemetry instrumentation.
- The SDK batches and exports OTLP/HTTP telemetry to the KubeSense
kubecolreceiver. kubecoldecodes the payload, identifies the source, maps GenAI fields, and preserves unmapped attributes.- The collector enriches missing model costs from the KubeSense pricing registry.
- Spans are written in batches to ClickHouse under
kube_llmtables. kubeapireads the stored data for dashboards, trace details, users, sessions, prompts, scores, datasets, and evaluations.
The original application remains responsible for creating and exporting telemetry. KubeSense does not replace the Langfuse or OpenLIT SDK in the application.
How the integration works
The SDK creates spans around the application workflow and exports them as OTLP/HTTP. kubecol accepts the protobuf or JSON representation, preserves trace and parent span IDs, and converts each span into the KubeSense LLM span model. A single trace can therefore contain model generations together with retrieval, tool, and application work.
Incoming spans are buffered and written in batches to ClickHouse. The KubeSense API reads the stored spans and supplies the list, detail, dashboard, session, user, and evaluation views. The collector also enriches spans with model pricing when the SDK did not provide cost fields.
Recommended instrumentation model
Use one trace for one user request or agent run. Create child observations for each meaningful operation:
- Application or agent span for the complete workflow.
- Retrieval or preprocessing spans for context construction.
- Generation spans for each model call.
- Tool or vector database spans for external operations.
- Evaluation spans or scores for quality measurements.
Keep user_id, session_id, application.id, and service.name stable and consistent. This makes the same data useful in tracing, users, sessions, dashboard, and evaluation views.
Supported operation types
KubeSense accepts more than chat completions. Common operation values include chat, embeddings, execute_tool, invoke_agent, and image_generation. Provider and framework names are retained separately, so model comparisons do not depend on a single SDK vendor.
Where to start
| Goal | Start here |
|---|---|
| Send your first traces | Installation |
| Understand the terminology | Observability concepts |
| Understand traffic, latency, and spend | Overview |
| Debug a model request or agent workflow | Traces |
| Break down usage by customer | Users |
| Follow a multi-turn conversation | Sessions |
| Create and test prompts | Prompt Hub |
| Measure response quality | Evaluation concepts |
Integration boundaries
KubeSense provides the collector, storage, API, dashboard, prompt workspace, and evaluation features. Langfuse and OpenLIT provide the application instrumentation and export behavior. Configure user IDs, session IDs, prompt content, model calls, and privacy controls in the application and its selected SDK.
Prompt Hub is a dashboard-based workspace for creating, versioning, and testing prompts. It does not provide a KubeSense application SDK, runtime prompt-fetching client, or automatic prompt injection into Langfuse or OpenLIT applications.
Privacy and operational considerations
LLM telemetry can contain prompts, completions, tool arguments, retrieved documents, user identifiers, and provider metadata. Before enabling production collection:
- Decide which input and output content your organization is permitted to store.
- Disable OpenLIT message capture when content should not be exported.
- Use stable internal user and session identifiers instead of secrets or unnecessary personal data.
- Keep provider credentials in secrets and use collector authentication where required.
- Bound large prompts, completions, tool payloads, and metadata to control storage and export volume.
- Use sampling when full-fidelity collection is not needed for every request.