Kubesense

Managed Service Metrics

Collect metrics from GCP managed services — Cloud SQL, Memorystore Redis, Cloud Run, Cloud Functions — into KubeSense using an OpenTelemetry Collector and the Cloud Monitoring API.

Overview

GCP publishes metrics for every managed service — Cloud SQL, Memorystore Redis, Cloud Run, Cloud Functions, and more — to the Cloud Monitoring API with no agent required on the service itself. KubeSense collects these by running an OpenTelemetry Collector inside your GKE cluster that reads the Cloud Monitoring API and remote-writes into the KubeSense metrics store.

The collector never connects to your databases or caches directly. It only reads metrics GCP already exposes, so there is nothing to install on the managed service, and the only permission needed is read-only access to Cloud Monitoring.

Architecture

Cloud SQL / Memorystore Redis / Cloud Run / Cloud Functions
        │  (metrics published automatically — no agent)
        ▼
   Cloud Monitoring API
        │  pulled by the googlecloudmonitoring receiver
        │  auth: Workload Identity (no key files)
        ▼
   OpenTelemetry Collector  (runs in your GKE cluster, namespace kubesense)
        │  prometheusremotewrite exporter
        ▼
   kubesense-metrics-scraper → metrics-store → webapp / Grafana

Prerequisites

Before you begin, ensure you have:

  1. A GKE cluster with the managed services (Cloud SQL, Memorystore Redis, etc.) already running in the same GCP project
  2. KubeSense deployed in the cluster, namespace kubesense
  3. Workload Identity enabled on the GKE cluster
  4. Owner or Editor access on the GCP project (or someone with it working alongside you) to create a service account and grant IAM roles
  5. helm and kubectl access to the cluster

Step 1 — Enable the Cloud Monitoring API

In the Cloud Console: APIs & Services → Library → search Cloud Monitoring API → Enable.

This is usually already enabled if any GCP resources are running, but confirm it.

Step 2 — Create a service account for the collector

IAM & Admin → Service Accounts → Create Service Account

  • Name: kubesense-monitoring-reader (any name works)
  • Skip granting roles on this screen — that is the next step
  • Click Done

Step 3 — Grant read access to Cloud Monitoring

IAM & Admin → IAM → Grant Access

  • New principal: the service account email from Step 2 — kubesense-monitoring-reader@<PROJECT_ID>.iam.gserviceaccount.com
  • Role: Monitoring Viewer (roles/monitoring.viewer)
  • Save

This is the only permission the collector needs. It is read-only and project-wide, so it covers Cloud SQL, Redis, Cloud Run, and Cloud Functions metrics with no per-service scoping.

Step 4 — Confirm Workload Identity is enabled

Kubernetes Engine → Clusters → (your cluster) → Details → look for Workload Identity under Security.

  • If it is off: Edit → enable Workload Identity with pool <PROJECT_ID>.svc.id.goog.
  • Clusters created after ~2022 usually have it on by default.

Existing node pools must pick up Workload Identity: Node pools created before Workload Identity was enabled do not inherit it. Check each node pool's own Workload Identity setting and recreate or upgrade any that predate the change, otherwise the collector pod cannot obtain credentials.

Step 5 — Bind the Kubernetes identity to the service account

This lets the collector pod impersonate the Google service account from Step 2 with no downloaded key file.

The Kubernetes ServiceAccount name (<KSA_NAME>) is created by the OpenTelemetry Collector Helm chart — by default <helm-release-name>-opentelemetry-collector. If you have not deployed the collector yet, do Step 6 first, run kubectl get sa -n kubesense to find the exact name, then return here.

gcloud iam service-accounts add-iam-policy-binding \
  kubesense-monitoring-reader@<PROJECT_ID>.iam.gserviceaccount.com \
  --role="roles/iam.workloadIdentityUser" \
  --member="serviceAccount:<PROJECT_ID>.svc.id.goog[kubesense/<KSA_NAME>]"

The Console equivalent: IAM & Admin → Service Accounts → kubesense-monitoring-reader → Permissions → Grant Access, with principal <PROJECT_ID>.svc.id.goog[kubesense/<KSA_NAME>] and role Workload Identity User.

Step 6 — Deploy the OpenTelemetry Collector

Add the chart repository (one-time):

helm repo add open-telemetry https://open-telemetry.github.io/opentelemetry-helm-charts
helm repo update

Create a values file for the collector. The key sections are the service-account annotation that ties the pod to the GSA, the googlecloudmonitoring receiver, the deltatocumulative processor (see gotchas), and the prometheusremotewrite exporter pointing at the KubeSense metrics store:

mode: deployment

serviceAccount:
  create: true
  annotations:
    # the GSA email from Step 2
    iam.gke.io/gcp-service-account: kubesense-monitoring-reader@<PROJECT_ID>.iam.gserviceaccount.com

config:
  receivers:
    googlecloudmonitoring:
      project_id: <PROJECT_ID>
      collection_interval: 60s
      metrics_list:
        # trim to the metrics your dashboards actually use
        - metric_name: "cloudsql.googleapis.com/database/cpu/utilization"
        - metric_name: "cloudsql.googleapis.com/database/memory/utilization"
        - metric_name: "redis.googleapis.com/stats/memory/usage_ratio"
        - metric_name: "redis.googleapis.com/clients/connected"

  processors:
    # required — see the gotcha on Delta vs Cumulative metrics
    deltatocumulative: {}
    batch: {}

  exporters:
    prometheusremotewrite/kubesense:
      # in-cluster metrics-store; match your actual service DNS if different
      endpoint: http://kubesense-metrics-scraper.kubesense.svc.cluster.local:30060/api/v1/write
      tls:
        insecure: true

  service:
    pipelines:
      metrics:
        receivers: [googlecloudmonitoring]
        processors: [deltatocumulative, batch]
        exporters: [prometheusremotewrite/kubesense]

Install into the kubesense namespace:

helm install otel-collector open-telemetry/opentelemetry-collector \
  -n kubesense -f collector-values.yaml

If you did not know the ServiceAccount name when doing Step 5, get it now and complete that binding:

kubectl get sa -n kubesense

Then restart the collector so it picks up the credentials:

kubectl rollout restart deployment otel-collector-opentelemetry-collector -n kubesense

Pick your metrics deliberately: metrics_list maps one-to-one to Cloud Monitoring metric types. Include only the metrics your dashboards use — every entry is a separate API read on each collection interval, and Cloud Monitoring API calls are billed. The Cloud SQL and Redis metric names above are examples; check each dashboard's queries for the exact names it needs.

Step 7 — Verify

kubectl logs -n kubesense -l app.kubernetes.io/name=opentelemetry-collector --tail=50
  • Monitoring client successfully created — authentication worked.
  • 403 ... Unable to generate access token or Unauthenticated — the Step 5 binding is missing or has the wrong KSA name or namespace. Fix it and restart the pod.
  • No errors and periodic metric batches = data is flowing. Confirm the series appear in the KubeSense Data Explorer.

Step 8 — Import the dashboards

In the KubeSense webapp: Dashboards → Import → upload the relevant JSON (kubesense_cloudsql.json, kubesense_redis.json, kubesense_cloudrun.json, kubesense_cloudfunction.json).

Each dashboard has a $project variable — and $instance, $service, or $function where relevant. Set these after import to filter to your project ID and resource names.

Known gotchas

Histogram and Sum metrics silently disappear without deltatocumulative: GCP's distribution-type metrics — latency, CPU/memory utilization, execution time — report with Delta aggregation temporality, but the prometheusremotewrite exporter can only emit Cumulative series. It drops Delta metrics without logging an error, so the collector looks healthy while those metrics never appear.The fix is the deltatocumulative processor placed ahead of batch in the metrics pipeline, as shown in Step 6. Gauge metrics (instance count, active instances, memory usage ratio) are unaffected — only Sum and Histogram types need it.

Idle Cloud Run / Cloud Functions report nothing: Request-based metrics only produce data points when the service or function has received traffic in the last collection window. An idle service reporting no metrics is expected, not a misconfiguration — generate some traffic to confirm.

Troubleshooting

SymptomLikely causeAction
403 / Unauthenticated in collector logsWorkload Identity binding missing or wrong KSA/namespaceRecheck the Step 5 member string <PROJECT_ID>.svc.id.goog[kubesense/<KSA_NAME>], then restart the pod
Auth works but no metrics at allmetrics_list empty, or wrong project_idConfirm the receiver config and that the metrics exist in Cloud Monitoring
Gauges appear but latency/CPU/utilization missingDelta metrics dropped by the exporterAdd the deltatocumulative processor (see gotchas)
Cloud Run / Functions metrics missingService idle in the collection windowSend traffic and re-check
Metrics arrive but nothing in a dashboard$project / $instance variables not setSet the dashboard variables after import
Pod cannot get credentialsNode pool predates Workload IdentityRecreate/upgrade the node pool (Step 4)

Best practices

  • Scope metrics_list tightly — it directly drives Cloud Monitoring API cost and collector load.
  • Match collection_interval to the metric's publish rate — most GCP managed-service metrics update at 60s, so polling faster returns duplicates and spends quota.
  • Prefer Workload Identity over key files — no secret to rotate or leak, and it is the only auth path this guide uses.
  • One collector per project — the receiver reads a single project_id; for multiple projects, run one collector each or add multiple receivers.