Alerts API
Read, create, update and delete alert rules and notification channels with an API key.
The same alert rules you build in the UI can be driven from your own tooling — a GitOps pipeline that reconciles rules from a repo, a migration script, or a service that provisions a standard alert pack for every new cluster.
This page covers the full alerting surface: reading rules and their firing history, creating them, updating, enabling/disabling and deleting them, and the same for notification channels. For reading logs and traces see the Logs & Traces API; for metrics see the Metrics API; for dashboards see the Dashboards API.
Authentication
An API key in the X-API-Key header:
X-API-Key: <your-api-key>Create one under Settings → API Key Management. See Authentication for details.
API keys are scoped per key: you choose which modules a key may reach when you
create it. For the endpoints on this page the key needs the alerts scope (for
rules) and alerts_notification_channels (for channels) — a key without the scope is
rejected for that module regardless of who created it.
warning: Two things must both hold for a write call to succeed: The key is scoped to the module — otherwise the request is refused as out of scope. A key can only be given scopes its creator's own role can access, so a key never grants more than the person who made it. The role behind the key has write access to that module. Scopes select which modules a key can touch, not what it may do inside them — read vs write still comes from the role. A read-only role yields 403 on POST, PUT and DELETE, while every GET on this page still works. POST /api/alerts/rules/validate is the one exception among the POSTs: it stores nothing, so it is not write-gated and a read-only key can lint a rule before asking for it to be created.
Request conventions
| Concern | How it's passed |
|---|---|
| Content type | Content-Type: application/json |
| Body | JSON object (see each endpoint) |
| Response | The standard envelope — { "data": …, "error": false, "message": "…" } |
Endpoints
Alert rules
| Method | Path | Description |
|---|---|---|
GET | /api/alerts/rules | List rules — paginated, searchable, filterable. |
GET | /api/alerts/rules/{id} | Fetch one rule. |
POST | /api/alerts/rules | Create one alert rule. |
PUT | /api/alerts/rules/{id} | Replace one alert rule. |
DELETE | /api/alerts/rules/{id} | Delete one alert rule. |
POST | /api/alerts/rules/{id}/enable | Resume evaluation. |
POST | /api/alerts/rules/{id}/disable | Pause evaluation without deleting. |
POST | /api/alerts/rules/bulk | Create up to 500 rules, or dry-run validate them. |
POST | /api/alerts/rules/validate | Validate rules without writing. |
GET | /api/alerts/rules/stats | Counts by state across all rules. |
Alert activity
| Method | Path | Description |
|---|---|---|
GET | /api/alerts/rules/{id}/events | Fire/resolve events for one rule. |
GET | /api/alerts/rules/{id}/alerts | Currently active instances of one rule. |
GET | /api/alerts/events | Global alert history across all rules. |
GET | /api/alerts/active | Active (or recently resolved) alert instances. |
Notification channels
| Method | Path | Description |
|---|---|---|
GET | /api/alerts/notification-channels | List channels. |
GET | /api/alerts/notification-channels/{id} | Fetch one channel. |
POST | /api/alerts/notification-channels | Create a notification channel. |
PUT | /api/alerts/notification-channels/{id} | Update a channel (partial). |
DELETE | /api/alerts/notification-channels/{id} | Delete a channel. |
POST | /api/alerts/notification-channels/{id}/test | Send a test notification. |
Rule {id} is the rule's UUID; channel {id} is an integer.
List alert rules
GET /api/alerts/rules
curl -s "https://<your-kubesense-host>/api/alerts/rules?limit=50&search=cpu" \
-H "X-API-Key: $KUBESENSE_API_KEY"| Parameter | Description |
|---|---|
limit | Page size, capped at 1000. Omit it (or send 0) and every rule comes back — fine for a few dozen, not for a few thousand. |
offset | Rows to skip. |
search | Case-insensitive substring match on the rule name only. |
ex_<group> / in_<group> | Facet filters, repeatable. ex_ excludes the listed values; in_ keeps only them. Groups: enabled, state, severity, query_type, created_by, notification_channel. |
Rules come back newest first (created_at descending). The envelope carries a
count alongside data — the total matching your filters, not the size of the
page, so it is what you paginate against.
Fetch one alert rule
GET /api/alerts/rules/{id} returns the stored rule.
warning: The read shape is not the write shape. You cannot GET a rule, edit a field and PUT it straight back — the response uses the storage field names, and several are restructured:In the responseIn a POST/PUT bodyquery_configquerynotification_channel_idsnotificationChannelsroute_by_labelsrouteByLabelsthreshold_operator, threshold_value, frequency_typethreshold: { operator, value, frequency }time_windowtime_window_prometheus_formatevaluation_intervalevaluation_interval_prometheus_formatbreach_counting_windowbreach_counting_window_prometheus_formatcompared_to: "1h" (string)compared_to: { "value": 1, "unit": "hours" }enabled: true(no equivalent — see status below)Translate deliberately. If you are keeping rules in Git, store the request shape as the source of truth and treat the API response as read-only.
The response also carries evaluation state the write side has no say over —
current_state, last_evaluation_time, last_state_change_time, firing_since,
consecutive_breaches — plus created_at / created_by.
Create an alert rule
POST /api/alerts/rules
curl -s -X POST "https://<your-kubesense-host>/api/alerts/rules" \
-H "X-API-Key: $KUBESENSE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "High pod CPU {{pod}}",
"description": "Pod CPU above 80% for 5 minutes",
"severity": "warning",
"status": "active",
"query": [
{
"label": "A",
"selectedMode": "metrics",
"queryMode": "code",
"promql": "max by (pod)(rate(container_cpu_usage_seconds_total{namespace=\"prod\"}[5m]))",
"visible": true,
"fields": []
}
],
"metric_query_label": "A",
"threshold": { "operator": "greater_than", "value": 0.8, "frequency": "at_least_once" },
"threshold_frequency": "at_least_once",
"time_window_prometheus_format": "5m",
"evaluation_interval_prometheus_format": "1m",
"notificationChannels": [3],
"routeByLabels": false,
"labels": { "team": "platform" },
"unit": "percent_unit",
"no_data_state": "normal"
}'Top-level fields
| Field | Type | Description |
|---|---|---|
name | string | Rule name. {{field}} placeholders resolve per firing series — the field must be a group-by key (PromQL by (...) label, or a logs/traces groupBy field). |
description | string | Free text shown on the alert. |
severity | string | critical · error · warning · info. |
status | string | Accepted but ignored — a created rule is always enabled. Use POST /{id}/disable to pause one. |
query | array | One or more query objects — see below. |
metric_query_label | string | Label of the query the threshold applies to (single query → its label; formula → the formula's label). |
threshold | object | { operator, value, frequency }. Operators: greater_than, greater_than_or_equal, less_than, less_than_or_equal, equals, not_equals. |
threshold_frequency | string | at_least_once · more_than_once · always. |
breaches_count | number | Required with more_than_once (≥ 2). |
time_window_prometheus_format | string | How far back each evaluation looks — 30s, 5m, 1h, 1d. |
evaluation_interval_prometheus_format | string | How often the rule runs, same format. |
breach_counting_window_prometheus_format | string | Optional; defaults to the time window. |
condition_type | string | Omit for a plain threshold. change · change_percent · new_value. |
compared_to | object | { "value": 1, "unit": "hours" } — required for change/new-value conditions. |
notificationChannels | array | Channel ids (integers). |
routeByLabels | bool | false routes to the listed channels; true routes via each channel's label matchers. |
labels | object | Extra labels attached to every alert ({"team":"platform"}). |
no_data_state | string | normal (resolve, default) · firing · previous. |
unit | string | Display only — ns, bytes, percent, percent_unit, ms, short. |
include_samples / sample_limit | bool / number | Logs rules only — attach matching log rows to notifications. |
warning: Use no_data_state: "firing" only when the absence of the metric is itself the incident (a gauge or heartbeat that should always report). For a count or rate rule, a healthy window returns an empty result rather than 0, so firing would page continuously at value 0. Use normal for those.
Query objects
Each entry in query describes one data source. selectedMode is metrics,
logs, traces, or formula.
Metrics — supply PromQL and put grouping in by (...):
{ "label": "A", "selectedMode": "metrics", "queryMode": "code",
"promql": "sum by (pod)(rate(http_requests_total[5m]))",
"visible": true, "fields": [] }Logs / traces — aggregate with value_operation, group with groupBy, filter
with rawFilters:
{ "label": "A", "selectedMode": "logs",
"value_operation": "row_count",
"groupBy": [{ "field": "workload", "type": "string", "is_attribute": false }],
"rawFilters": { "level": ["ERROR"] },
"fields": [] }value_operation accepts row_count, unique_count, avg, sum, max, min,
p99, p95, p90, p75, p50. For a numeric aggregation, name the column in
fields — e.g. trace latency:
{ "label": "A", "selectedMode": "traces", "value_operation": "p95",
"fields": [{ "field": "duration", "type": "float", "is_attribute": false }],
"groupBy": [{ "field": "workload", "type": "string", "is_attribute": false }] }warning: Trace duration is stored in nanoseconds, so a 500 ms threshold is 500000000. Field names are the storage names (return_code, protocol_type, pod_name), not the UI labels — see the Log & Trace Fields reference.
Formula — combine other queries by label:
{ "label": "C", "selectedMode": "formula", "expression": "(B/A)*100", "visible": true }Set metric_query_label to the formula's label, and visible: false on the helper
queries so only the computed line is charted.
Update an alert rule
PUT /api/alerts/rules/{id} takes the same body as create — every field listed
above.
warning: This is a full replacement, not a patch. Any field you leave out is written as its zero value, so an update that only means to raise a threshold will quietly drop the rule's labels, notification channels and condition type if you send just {"threshold": …}. Send the whole rule every time.
curl -s -X PUT "https://<your-kubesense-host>/api/alerts/rules/$RULE_ID" \
-H "X-API-Key: $KUBESENSE_API_KEY" \
-H "Content-Type: application/json" \
-d @rule.jsonAn unparseable UUID is 400; a UUID that matches no rule is 404.
Changing the query or the threshold takes effect on the rule's next evaluation — there is no separate reload call. A rule that is currently firing keeps its instances until the new condition resolves them.
Enable or disable a rule
curl -s -X POST "https://<your-kubesense-host>/api/alerts/rules/$RULE_ID/disable" \
-H "X-API-Key: $KUBESENSE_API_KEY"POST /{id}/enable and POST /{id}/disable flip evaluation without touching the
rule's definition. Neither takes a body.
Disabling is the right way to silence a rule you intend to bring back — a maintenance window, a noisy rule pending a fix. It stops evaluation, so no new instances fire. Deleting is not reversible.
Delete an alert rule
DELETE /api/alerts/rules/{id}
curl -s -X DELETE "https://<your-kubesense-host>/api/alerts/rules/$RULE_ID" \
-H "X-API-Key: $KUBESENSE_API_KEY"Returns 200 with data: null. The deletion is recorded in the audit log against
the key's user.
note: A malformed UUID returns 400, but a well-formed UUID that matches nothing currently returns 500, not 404. Treat a 500 from this endpoint as "possibly already gone" rather than a server fault, and check with a GET before retrying.
Bulk create
POST /api/alerts/rules/bulk — up to 500 rules per request, each element in the
same shape as a single create.
curl -s -X POST "https://<your-kubesense-host>/api/alerts/rules/bulk" \
-H "X-API-Key: $KUBESENSE_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "dry_run": true, "rules": [ { "name": "…", "…": "…" } ] }'| Field | Type | Description |
|---|---|---|
dry_run | bool | true validates every rule and writes nothing. |
rules | array | 1–500 alert rule objects. |
The response reports each rule individually (index, name, valid, created,
plus any errors), so a partially-invalid batch tells you exactly which entries
failed rather than rejecting the lot silently.
note: Run with "dry_run": true first. It's the same validation the UI's import-review screen uses, so anything it accepts will create cleanly.
Reading alert activity
Four read endpoints cover "what has this rule been doing" and "what is on fire right
now". All of them are plain GETs needing only the alerts scope.
GET /api/alerts/rules/{id}/events — fire/resolve events for one rule, newest first.
limit defaults to 50 and is capped at 1000.
GET /api/alerts/rules/{id}/alerts — the instances of that rule currently firing.
GET /api/alerts/rules/stats — counts by state across every rule, for a dashboard
tile or a health check.
GET /api/alerts/events — global history across all rules:
| Parameter | Description |
|---|---|
rule_id | Scope to one rule. |
fingerprint | Scope to one alert instance. |
severity · state · event_type | Exact-match filters. |
search | Substring match. |
start · end | Window bounds. |
limit · offset | Pagination. |
compact=true | Trimmed rows — smaller payload for a timeline. |
GET /api/alerts/active — alert instances rather than events:
| Parameter | Description |
|---|---|
only_firing | Defaults to true. Send only_firing=false to include resolved instances in the window. |
fingerprint | One instance. |
start · end · limit · offset | As above. |
# Everything firing right now, with its assignee and ack state.
curl -s "https://<your-kubesense-host>/api/alerts/active?limit=50" \
-H "X-API-Key: $KUBESENSE_API_KEY"Both list endpoints return count — the total matching the filters — next to data.
List notification channels
GET /api/alerts/notification-channels returns every channel;
GET /api/alerts/notification-channels/{id} returns one.
warning: Secrets are masked on read. A JSM channel comes back with config.api_key as an empty string, never the stored value. That is deliberate — but it means a read-modify-write of a JSM channel would otherwise blank its key, so PUT treats a blank api_key as "keep the stored one". Sending a real value replaces it.
Create a notification channel
POST /api/alerts/notification-channels
curl -s -X POST "https://<your-kubesense-host>/api/alerts/notification-channels" \
-H "X-API-Key: $KUBESENSE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "platform-slack",
"type": "slack",
"enabled": true,
"config": {
"webhook_url": "https://hooks.slack.com/services/T0/B0/xxx",
"route_match": []
}
}'| Field | Type | Description |
|---|---|---|
name | string | Unique channel name. |
type | string | slack · slack_app · email · msteams · google_chat · pagerduty · webhook · datadog_oncall · jsm. |
enabled | bool | Whether the channel receives alerts. |
config | object | Type-specific settings — see Integrations for each type's fields. |
template_id | number | Optional message template id. |
config.route_match holds label matchers for label-based routing
([{"label":"severity","operator":"=","value":"critical"}]); leave it [] for
direct assignment.
Update a notification channel
PUT /api/alerts/notification-channels/{id}
Unlike a rule update, this one is partial — send only the fields you want to change:
curl -s -X PUT "https://<your-kubesense-host>/api/alerts/notification-channels/3" \
-H "X-API-Key: $KUBESENSE_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "enabled": false }'| Field | Behaviour when omitted |
|---|---|
name · type · config | Left as stored. |
enabled | Left as stored. Send true/false to change it. |
template_id | Left as stored. Send null to clear the template. |
warning: config is the exception: when you send it, it replaces the stored config wholesale rather than merging. Sending {"config": {"webhook_url": "…"}} on a channel that had route_match set will drop the matchers. Read the channel first and send the full config object back with your edit applied.
Changing type is allowed — converting a slack webhook channel to slack_app,
say. The config is validated against the new type, so send a config that suits it
(or make sure the stored one already does).
Delete a notification channel
DELETE /api/alerts/notification-channels/{id} — {id} is the integer id.
Check what still points at a channel before removing it: rules that referenced it by
id in notificationChannels lose that destination.
Test a notification channel
POST /api/alerts/notification-channels/{id}/test sends a sample alert through a
saved channel. POST /api/alerts/notification-channels/test does the same for a
config you have not saved yet — useful in CI to prove a webhook or token works before
committing it.
Both are write-gated, and both surface the provider's own failure reason (an invalid
Slack channel, a rejected token) in the message field, so log the body on a 500
rather than just the status.
Errors
| Status | Meaning |
|---|---|
400 | Invalid body or path — malformed JSON, a missing required field, an unknown group-by/filter field, or an id that isn't a valid UUID. |
401 | Missing, invalid, or expired API key. |
403 | Either the key isn't scoped to the module, or the role behind it has read-only access. Only GETs and /rules/validate work with a read-only role. |
404 | No rule or channel with that id (PUT /rules/{id}, and every channel endpoint). |
500 | Server error while persisting — and the response DELETE /rules/{id} gives for a rule that doesn't exist. |
Errors use the same envelope as successes, with the reason in message:
{ "error": true, "message": "Alert rule not found", "data": null }