Kubesense

Data Streams

Data Streams gives you the Kafka side of your system: the pipeline topology (producer → topic → consumer), consumer-group lag and health, and per-topic throughput — stitched onto the same trace and metric data KubeSense already collects. It answers "is my messaging healthy, and where is it backing up?" without a separate agent.

Data Streams is gated by the ENABLE_DATA_STREAMS deployment flag and the Traces RBAC module (it is an APM sibling of the Service Map). If you don't see it in the sidebar, enable the flag and confirm your role has Traces access.

How it's built — two signals

Data Streams composes two independent signals. Each tab tells you honestly which are present, so an empty view is a setup checklist, not a dead end.

SignalSourcePowers
Kafka trace spansProducer/consumer services instrumented with OpenTelemetry, emitting messaging.system=kafka spansThe Pipeline topology (which service talks to which topic) + observed span rates and hop latency
Broker metricsThe OTel kafkametrics receiver on the sensor's otel-agent, or a Confluent Cloud integrationConsumer Groups (lag, members, consume rate) and Topics (partitions, throughput, replication health)
  • Traces alone → you see the pipeline graph, but no lag/partition health.
  • Broker metrics alone → you see topics and consumer-group lag, but nothing links them to services.
  • Both → the full picture.

Pipeline

The Pipeline tab renders the messaging topology from trace context: producers, topics, and consumers as nodes, with edges carrying the observed span rate and p95 hop latency. It is built from the trace_one_minute rollup (no raw-trace scans), so it stays fast on large maps.

  • Filter to Problems (lagging / stalled) or Idle nodes with the toolbar.
  • Click a topic or service to open a detail drawer with its rates, errors, and related entities.
  • A quiet window shows a diagnostic telling you which of the two signals is missing and how to supply it — widen the time range if the pipeline is simply idle.

Consumer Groups

Per consumer group: offset backlog (messages behind), consume rate, estimated processing time to drain, member count, and a status (healthy / lagging / stalled). Identity is (cluster, group, topic), so identically-named groups in different clusters never merge.

A — means the metric is unavailable, never a silent zero. stalled means a real backlog with a consume rate of ~0/s (it won't drain on its own).

Kafka metric source

When more than one broker source is available, a Kafka metric source selector lets you choose:

  • Automatic (recommended) — the richest adapted source reporting for your cluster (native OTel kafkametrics is preferred because it carries members, consume rate, and stalled detection).
  • OTel Kafka metrics — force the native kafkametrics schema.
  • Confluent Cloud — force the Confluent integration (see the limitations below).

Topics

Per topic: partitions, consumer-group count, total consumer lag, produce rate, and under-replicated / offline partitions. Health is only shown when the broker actually reports it — a topic with no replication signal reads unknown, never a false green.

On the Service Map

Data Streams also overlays the Service Map. Toggle Data Streams on the logical view to add the Kafka messaging layer — service → topic → service, resolving consumers directly from consume traces (no consumer-group node, matching the Pipeline graph). The overlay is non-fatal: if it can't be built for a window, the map still renders the APM services and notes that Data Streams is unavailable.

Providers

Enable the OpenTelemetry kafkametrics receiver on the sensor's otel-agent, pointed at your Kafka bootstrap, with the brokers, topics, and consumers scrapers. This is the richest source — it powers the full Consumer Groups and Topics tabs.

Use the OTel kafkametrics receiver, not the Prometheus kafka-exporter. Data Streams reads the kafka_* metric names the OTel receiver emits (kafka_consumer_group_lag_sum_ratio, kafka_topic_partitions, …); the Prometheus exporter uses a different schema and is not adapted.

Confluent Cloud

If you run a Confluent Cloud integration, Data Streams can read consumer-group lag from it directly — no in-cluster receiver required. Pick Confluent Cloud as the Kafka metric source, then the Kafka cluster to scope to.

Confluent is lag-only. Confluent's Metrics API collects consumer_lag_offsets per consumer group but does not expose per-group consume rate, member count, or a topic breakdown on lag. So for the Confluent source:

  • Offset backlog and status (caught up / lagging) are shown.
  • Consume rate, processing time, and members columns are omitted (not collected).
  • A Topics tab is not available for Confluent yet.

Setup checklist

  1. Enable the feature — set ENABLE_DATA_STREAMS=true and grant the Traces module to the role.
  2. Broker metrics — turn on the OTel kafkametrics receiver on the sensor's otel-agent (or attach a Confluent Cloud integration). Topics, lag, and consumer groups appear within a scrape interval.
  3. Trace instrumentation — instrument your producer and consumer services with OpenTelemetry so they emit messaging.system=kafka spans. This is what links topics to services on the Pipeline and Service Map.
  4. Widen the window if a pipeline looks empty — a quiet pipeline may simply be idle.