Kubesense

Alerts API

Read, create, update and delete alert rules and notification channels with an API key.

The same alert rules you build in the UI can be driven from your own tooling — a GitOps pipeline that reconciles rules from a repo, a migration script, or a service that provisions a standard alert pack for every new cluster.

This page covers the full alerting surface: reading rules and their firing history, creating them, updating, enabling/disabling and deleting them, and the same for notification channels. For reading logs and traces see the Logs & Traces API; for metrics see the Metrics API; for dashboards see the Dashboards API.

Authentication

An API key in the X-API-Key header:

X-API-Key: <your-api-key>

Create one under Settings → API Key Management. See Authentication for details.

API keys are scoped per key: you choose which modules a key may reach when you create it. For the endpoints on this page the key needs the alerts scope (for rules) and alerts_notification_channels (for channels) — a key without the scope is rejected for that module regardless of who created it.

warning: Two things must both hold for a write call to succeed: The key is scoped to the module — otherwise the request is refused as out of scope. A key can only be given scopes its creator's own role can access, so a key never grants more than the person who made it. The role behind the key has write access to that module. Scopes select which modules a key can touch, not what it may do inside them — read vs write still comes from the role. A read-only role yields 403 on POST, PUT and DELETE, while every GET on this page still works. POST /api/alerts/rules/validate is the one exception among the POSTs: it stores nothing, so it is not write-gated and a read-only key can lint a rule before asking for it to be created.

Request conventions

ConcernHow it's passed
Content typeContent-Type: application/json
BodyJSON object (see each endpoint)
ResponseThe standard envelope — { "data": …, "error": false, "message": "…" }

Endpoints

Alert rules

MethodPathDescription
GET/api/alerts/rulesList rules — paginated, searchable, filterable.
GET/api/alerts/rules/{id}Fetch one rule.
POST/api/alerts/rulesCreate one alert rule.
PUT/api/alerts/rules/{id}Replace one alert rule.
DELETE/api/alerts/rules/{id}Delete one alert rule.
POST/api/alerts/rules/{id}/enableResume evaluation.
POST/api/alerts/rules/{id}/disablePause evaluation without deleting.
POST/api/alerts/rules/bulkCreate up to 500 rules, or dry-run validate them.
POST/api/alerts/rules/validateValidate rules without writing.
GET/api/alerts/rules/statsCounts by state across all rules.

Alert activity

MethodPathDescription
GET/api/alerts/rules/{id}/eventsFire/resolve events for one rule.
GET/api/alerts/rules/{id}/alertsCurrently active instances of one rule.
GET/api/alerts/eventsGlobal alert history across all rules.
GET/api/alerts/activeActive (or recently resolved) alert instances.

Notification channels

MethodPathDescription
GET/api/alerts/notification-channelsList channels.
GET/api/alerts/notification-channels/{id}Fetch one channel.
POST/api/alerts/notification-channelsCreate a notification channel.
PUT/api/alerts/notification-channels/{id}Update a channel (partial).
DELETE/api/alerts/notification-channels/{id}Delete a channel.
POST/api/alerts/notification-channels/{id}/testSend a test notification.

Rule {id} is the rule's UUID; channel {id} is an integer.


List alert rules

GET /api/alerts/rules

curl -s "https://<your-kubesense-host>/api/alerts/rules?limit=50&search=cpu" \
  -H "X-API-Key: $KUBESENSE_API_KEY"
ParameterDescription
limitPage size, capped at 1000. Omit it (or send 0) and every rule comes back — fine for a few dozen, not for a few thousand.
offsetRows to skip.
searchCase-insensitive substring match on the rule name only.
ex_<group> / in_<group>Facet filters, repeatable. ex_ excludes the listed values; in_ keeps only them. Groups: enabled, state, severity, query_type, created_by, notification_channel.

Rules come back newest first (created_at descending). The envelope carries a count alongside data — the total matching your filters, not the size of the page, so it is what you paginate against.

Fetch one alert rule

GET /api/alerts/rules/{id} returns the stored rule.

warning: The read shape is not the write shape. You cannot GET a rule, edit a field and PUT it straight back — the response uses the storage field names, and several are restructured:In the responseIn a POST/PUT bodyquery_configquerynotification_channel_idsnotificationChannelsroute_by_labelsrouteByLabelsthreshold_operator, threshold_value, frequency_typethreshold: { operator, value, frequency }time_windowtime_window_prometheus_formatevaluation_intervalevaluation_interval_prometheus_formatbreach_counting_windowbreach_counting_window_prometheus_formatcompared_to: "1h" (string)compared_to: { "value": 1, "unit": "hours" }enabled: true(no equivalent — see status below)Translate deliberately. If you are keeping rules in Git, store the request shape as the source of truth and treat the API response as read-only.

The response also carries evaluation state the write side has no say over — current_state, last_evaluation_time, last_state_change_time, firing_since, consecutive_breaches — plus created_at / created_by.

Create an alert rule

POST /api/alerts/rules

curl -s -X POST "https://<your-kubesense-host>/api/alerts/rules" \
  -H "X-API-Key: $KUBESENSE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "High pod CPU {{pod}}",
    "description": "Pod CPU above 80% for 5 minutes",
    "severity": "warning",
    "status": "active",
    "query": [
      {
        "label": "A",
        "selectedMode": "metrics",
        "queryMode": "code",
        "promql": "max by (pod)(rate(container_cpu_usage_seconds_total{namespace=\"prod\"}[5m]))",
        "visible": true,
        "fields": []
      }
    ],
    "metric_query_label": "A",
    "threshold": { "operator": "greater_than", "value": 0.8, "frequency": "at_least_once" },
    "threshold_frequency": "at_least_once",
    "time_window_prometheus_format": "5m",
    "evaluation_interval_prometheus_format": "1m",
    "notificationChannels": [3],
    "routeByLabels": false,
    "labels": { "team": "platform" },
    "unit": "percent_unit",
    "no_data_state": "normal"
  }'

Top-level fields

FieldTypeDescription
namestringRule name. {{field}} placeholders resolve per firing series — the field must be a group-by key (PromQL by (...) label, or a logs/traces groupBy field).
descriptionstringFree text shown on the alert.
severitystringcritical · error · warning · info.
statusstringAccepted but ignored — a created rule is always enabled. Use POST /{id}/disable to pause one.
queryarrayOne or more query objects — see below.
metric_query_labelstringLabel of the query the threshold applies to (single query → its label; formula → the formula's label).
thresholdobject{ operator, value, frequency }. Operators: greater_than, greater_than_or_equal, less_than, less_than_or_equal, equals, not_equals.
threshold_frequencystringat_least_once · more_than_once · always.
breaches_countnumberRequired with more_than_once (≥ 2).
time_window_prometheus_formatstringHow far back each evaluation looks — 30s, 5m, 1h, 1d.
evaluation_interval_prometheus_formatstringHow often the rule runs, same format.
breach_counting_window_prometheus_formatstringOptional; defaults to the time window.
condition_typestringOmit for a plain threshold. change · change_percent · new_value.
compared_toobject{ "value": 1, "unit": "hours" } — required for change/new-value conditions.
notificationChannelsarrayChannel ids (integers).
routeByLabelsboolfalse routes to the listed channels; true routes via each channel's label matchers.
labelsobjectExtra labels attached to every alert ({"team":"platform"}).
no_data_statestringnormal (resolve, default) · firing · previous.
unitstringDisplay only — ns, bytes, percent, percent_unit, ms, short.
include_samples / sample_limitbool / numberLogs rules only — attach matching log rows to notifications.

warning: Use no_data_state: "firing" only when the absence of the metric is itself the incident (a gauge or heartbeat that should always report). For a count or rate rule, a healthy window returns an empty result rather than 0, so firing would page continuously at value 0. Use normal for those.

Query objects

Each entry in query describes one data source. selectedMode is metrics, logs, traces, or formula.

Metrics — supply PromQL and put grouping in by (...):

{ "label": "A", "selectedMode": "metrics", "queryMode": "code",
  "promql": "sum by (pod)(rate(http_requests_total[5m]))",
  "visible": true, "fields": [] }

Logs / traces — aggregate with value_operation, group with groupBy, filter with rawFilters:

{ "label": "A", "selectedMode": "logs",
  "value_operation": "row_count",
  "groupBy": [{ "field": "workload", "type": "string", "is_attribute": false }],
  "rawFilters": { "level": ["ERROR"] },
  "fields": [] }

value_operation accepts row_count, unique_count, avg, sum, max, min, p99, p95, p90, p75, p50. For a numeric aggregation, name the column in fields — e.g. trace latency:

{ "label": "A", "selectedMode": "traces", "value_operation": "p95",
  "fields": [{ "field": "duration", "type": "float", "is_attribute": false }],
  "groupBy": [{ "field": "workload", "type": "string", "is_attribute": false }] }

warning: Trace duration is stored in nanoseconds, so a 500 ms threshold is 500000000. Field names are the storage names (return_code, protocol_type, pod_name), not the UI labels — see the Log & Trace Fields reference.

Formula — combine other queries by label:

{ "label": "C", "selectedMode": "formula", "expression": "(B/A)*100", "visible": true }

Set metric_query_label to the formula's label, and visible: false on the helper queries so only the computed line is charted.


Update an alert rule

PUT /api/alerts/rules/{id} takes the same body as create — every field listed above.

warning: This is a full replacement, not a patch. Any field you leave out is written as its zero value, so an update that only means to raise a threshold will quietly drop the rule's labels, notification channels and condition type if you send just {"threshold": …}. Send the whole rule every time.

curl -s -X PUT "https://<your-kubesense-host>/api/alerts/rules/$RULE_ID" \
  -H "X-API-Key: $KUBESENSE_API_KEY" \
  -H "Content-Type: application/json" \
  -d @rule.json

An unparseable UUID is 400; a UUID that matches no rule is 404.

Changing the query or the threshold takes effect on the rule's next evaluation — there is no separate reload call. A rule that is currently firing keeps its instances until the new condition resolves them.

Enable or disable a rule

curl -s -X POST "https://<your-kubesense-host>/api/alerts/rules/$RULE_ID/disable" \
  -H "X-API-Key: $KUBESENSE_API_KEY"

POST /{id}/enable and POST /{id}/disable flip evaluation without touching the rule's definition. Neither takes a body.

Disabling is the right way to silence a rule you intend to bring back — a maintenance window, a noisy rule pending a fix. It stops evaluation, so no new instances fire. Deleting is not reversible.

Delete an alert rule

DELETE /api/alerts/rules/{id}

curl -s -X DELETE "https://<your-kubesense-host>/api/alerts/rules/$RULE_ID" \
  -H "X-API-Key: $KUBESENSE_API_KEY"

Returns 200 with data: null. The deletion is recorded in the audit log against the key's user.

note: A malformed UUID returns 400, but a well-formed UUID that matches nothing currently returns 500, not 404. Treat a 500 from this endpoint as "possibly already gone" rather than a server fault, and check with a GET before retrying.


Bulk create

POST /api/alerts/rules/bulk — up to 500 rules per request, each element in the same shape as a single create.

curl -s -X POST "https://<your-kubesense-host>/api/alerts/rules/bulk" \
  -H "X-API-Key: $KUBESENSE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "dry_run": true, "rules": [ { "name": "…", "…": "…" } ] }'
FieldTypeDescription
dry_runbooltrue validates every rule and writes nothing.
rulesarray1–500 alert rule objects.

The response reports each rule individually (index, name, valid, created, plus any errors), so a partially-invalid batch tells you exactly which entries failed rather than rejecting the lot silently.

note: Run with "dry_run": true first. It's the same validation the UI's import-review screen uses, so anything it accepts will create cleanly.


Reading alert activity

Four read endpoints cover "what has this rule been doing" and "what is on fire right now". All of them are plain GETs needing only the alerts scope.

GET /api/alerts/rules/{id}/events — fire/resolve events for one rule, newest first. limit defaults to 50 and is capped at 1000.

GET /api/alerts/rules/{id}/alerts — the instances of that rule currently firing.

GET /api/alerts/rules/stats — counts by state across every rule, for a dashboard tile or a health check.

GET /api/alerts/events — global history across all rules:

ParameterDescription
rule_idScope to one rule.
fingerprintScope to one alert instance.
severity · state · event_typeExact-match filters.
searchSubstring match.
start · endWindow bounds.
limit · offsetPagination.
compact=trueTrimmed rows — smaller payload for a timeline.

GET /api/alerts/active — alert instances rather than events:

ParameterDescription
only_firingDefaults to true. Send only_firing=false to include resolved instances in the window.
fingerprintOne instance.
start · end · limit · offsetAs above.
# Everything firing right now, with its assignee and ack state.
curl -s "https://<your-kubesense-host>/api/alerts/active?limit=50" \
  -H "X-API-Key: $KUBESENSE_API_KEY"

Both list endpoints return count — the total matching the filters — next to data.


List notification channels

GET /api/alerts/notification-channels returns every channel; GET /api/alerts/notification-channels/{id} returns one.

warning: Secrets are masked on read. A JSM channel comes back with config.api_key as an empty string, never the stored value. That is deliberate — but it means a read-modify-write of a JSM channel would otherwise blank its key, so PUT treats a blank api_key as "keep the stored one". Sending a real value replaces it.

Create a notification channel

POST /api/alerts/notification-channels

curl -s -X POST "https://<your-kubesense-host>/api/alerts/notification-channels" \
  -H "X-API-Key: $KUBESENSE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "platform-slack",
    "type": "slack",
    "enabled": true,
    "config": {
      "webhook_url": "https://hooks.slack.com/services/T0/B0/xxx",
      "route_match": []
    }
  }'
FieldTypeDescription
namestringUnique channel name.
typestringslack · slack_app · email · msteams · google_chat · pagerduty · webhook · datadog_oncall · jsm.
enabledboolWhether the channel receives alerts.
configobjectType-specific settings — see Integrations for each type's fields.
template_idnumberOptional message template id.

config.route_match holds label matchers for label-based routing ([{"label":"severity","operator":"=","value":"critical"}]); leave it [] for direct assignment.

Update a notification channel

PUT /api/alerts/notification-channels/{id}

Unlike a rule update, this one is partial — send only the fields you want to change:

curl -s -X PUT "https://<your-kubesense-host>/api/alerts/notification-channels/3" \
  -H "X-API-Key: $KUBESENSE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "enabled": false }'
FieldBehaviour when omitted
name · type · configLeft as stored.
enabledLeft as stored. Send true/false to change it.
template_idLeft as stored. Send null to clear the template.

warning: config is the exception: when you send it, it replaces the stored config wholesale rather than merging. Sending {"config": {"webhook_url": "…"}} on a channel that had route_match set will drop the matchers. Read the channel first and send the full config object back with your edit applied.

Changing type is allowed — converting a slack webhook channel to slack_app, say. The config is validated against the new type, so send a config that suits it (or make sure the stored one already does).

Delete a notification channel

DELETE /api/alerts/notification-channels/{id} — {id} is the integer id.

Check what still points at a channel before removing it: rules that referenced it by id in notificationChannels lose that destination.

Test a notification channel

POST /api/alerts/notification-channels/{id}/test sends a sample alert through a saved channel. POST /api/alerts/notification-channels/test does the same for a config you have not saved yet — useful in CI to prove a webhook or token works before committing it.

Both are write-gated, and both surface the provider's own failure reason (an invalid Slack channel, a rejected token) in the message field, so log the body on a 500 rather than just the status.


Errors

StatusMeaning
400Invalid body or path — malformed JSON, a missing required field, an unknown group-by/filter field, or an id that isn't a valid UUID.
401Missing, invalid, or expired API key.
403Either the key isn't scoped to the module, or the role behind it has read-only access. Only GETs and /rules/validate work with a read-only role.
404No rule or channel with that id (PUT /rules/{id}, and every channel endpoint).
500Server error while persisting — and the response DELETE /rules/{id} gives for a rule that doesn't exist.

Errors use the same envelope as successes, with the reason in message:

{ "error": true, "message": "Alert rule not found", "data": null }