Kubesense

Model Providers

The KubeSense AI service reaches a large language model through one provider, chosen at deployment time. Every agent surface — conversations, investigations, and RCA — uses that same provider.

You bring your own model. KubeSense does not host one, and no inference is billed through KubeSense: requests go from your AI service pod directly to the provider you configure, using credentials you control.

Supported providers

AI_PROVIDERModelRuns onCredentials
anthropicClaudeAnthropicAPI key
anthropic‑vertexClaudeGoogle Vertex AIService account
anthropic‑bedrockClaudeAWS BedrockIAM role, or API key
vertex‑geminiGeminiGoogle Vertex AIService account, or express key
openaiGPTOpenAIAPI key
openai‑compatibleAnyAzure OpenAI, Groq, Together, Ollama, LiteLLMBase URL + API key

Exactly one provider is active per deployment. Set AI_PROVIDER and the credentials for that provider only; the service validates the pair at startup and refuses to boot on a configuration it cannot use, rather than failing on a user's first question.

Configuration

Two variables are common to every provider:

AI_PROVIDER=anthropic-bedrock
AI_MODEL=us.anthropic.claude-sonnet-4-5-20250929-v1:0

warning: AI_MODEL is not portable between providers. The same model carries a different identifier on each one — claude-sonnet-5 on Anthropic direct is us.anthropic.claude-sonnet-4-5-20250929-v1:0 on Bedrock. Changing AI_PROVIDER without also changing AI_MODEL fails at request time, not at startup.

Anthropic

AI_PROVIDER=anthropic
AI_MODEL=claude-sonnet-5
ANTHROPIC_API_KEY=sk-ant-...

Claude on Vertex AI

AI_PROVIDER=anthropic-vertex
AI_MODEL=claude-sonnet-5
GCP_VERTEX_PROJECT=my-gcp-project
GCP_VERTEX_LOCATION=us-east5

That is the whole configuration when the pod runs on GCP — see Vertex credentials below.

Claude on Vertex must be approved in Vertex AI Model Garden before it can be invoked — Google requires customers to explicitly accept third-party models. Approval is separate from IAM permissions.

info: Claude on Vertex is region-pinned — set GCP_VERTEX_LOCATION to a region that serves it, such as us-east5. The global endpoint works for Gemini but not for Claude.

Gemini on Vertex AI

AI_PROVIDER=vertex-gemini
AI_MODEL=gemini-3.6-flash
GCP_VERTEX_PROJECT=my-gcp-project
GCP_VERTEX_LOCATION=global

Gemini also supports Vertex express mode, which needs only an API key and no project or location:

AI_PROVIDER=vertex-gemini
AI_MODEL=gemini-3.6-flash
GCP_VERTEX_API_KEY=...

Set one or the other. Configuring an express key and a project/location is rejected at startup: the two name different projects, quotas and data boundaries, and silently preferring one would run your traffic somewhere your configuration does not say.

Vertex credentials

Both Vertex providers authenticate through Google Application Default Credentials, which finds a credential on its own. In order, ADC uses:

  1. A key file, if GOOGLE_APPLICATION_CREDENTIALS points at one
  2. Your gcloud login, for local development
  3. The service account attached to the workload — GKE Workload Identity, a GCE metadata server, or Cloud Run

info: Running on GCP? You do not need GOOGLE_APPLICATION_CREDENTIALS. On GKE with Workload Identity the third step applies: annotate the pod's Kubernetes service account with a Google service account holding the roles/aiplatform.user role, and credentials are issued automatically. Nothing is stored, nothing expires, and there is no key to rotate.

Set GOOGLE_APPLICATION_CREDENTIALS only when ADC has nothing else to find — reaching Vertex from EKS, AKS, or anywhere outside GCP:

GOOGLE_APPLICATION_CREDENTIALS=/var/secrets/gcp/sa.json

Mount the service account key as a secret and point the variable at the file path. This is the fallback, not the default: a static key is a credential you now have to store and rotate, so prefer Workload Identity wherever the workload actually runs on GCP.

Claude on AWS Bedrock

AI_PROVIDER=anthropic-bedrock
AI_MODEL=us.anthropic.claude-sonnet-4-5-20250929-v1:0
AWS_BEDROCK_REGION=us-east-1

Leave AWS_BEDROCK_API_KEY unset to authenticate through the standard AWS credential chain — that is the IRSA path on EKS, and it needs no stored secret. Annotate the pod's service account with an IAM role granting bedrock:InvokeModel and bedrock:InvokeModelWithResponseStream, and the pod's projected token is exchanged for temporary credentials automatically.

Set AWS_BEDROCK_API_KEY instead where IRSA is unavailable — on AKS or GKE, for instance, since IRSA is EKS-only.

warning: Bedrock model access is granted per account and per region, separately from IAM. Enabling a model in us-east-1 does nothing for eu-west-1. Model ids beginning us. are cross-region inference profiles, so the IAM policy must cover the inference-profile ARN as well as the foundation-model one.

OpenAI

AI_PROVIDER=openai
AI_MODEL=gpt-5.5
OPENAI_API_KEY=sk-...
OPENAI_BASE_URL=            # optional; defaults to api.openai.com

This uses the Chat Completions API. The Responses API is not offered — it requires a replayed function call to carry its paired reasoning item, which breaks a tool-calling loop on its second round trip. Reasoning text is therefore not visible on this provider; the model still reasons, and is still billed for it.

OpenAI-compatible endpoints

AI_PROVIDER=openai-compatible
AI_MODEL=<model or deployment name>
OPENAI_COMPATIBLE_BASE_URL=https://...
OPENAI_COMPATIBLE_API_KEY=...

For anything exposing the OpenAI wire format: Azure OpenAI, Groq, Together, Ollama, or a LiteLLM proxy. For Azure, AI_MODEL is your deployment name rather than a model id, and the base URL is the deployment endpoint.

This is the escape hatch. No model catalogue applies to it, so the startup model check is skipped and any model name is accepted.

Choosing where the model runs

The same KubeSense build serves every combination — the provider is configuration, not a separate image. What differs is how credentials reach the pod:

ProviderClusterCredentials
anthropic, openai, openai‑compatibleAnyAPI key in a secret
anthropic‑bedrockEKSIRSA — no stored secret
anthropic‑bedrockAKS, GKEAPI key, or AWS access keys
anthropic‑vertex, vertex‑geminiGKEWorkload Identity — no stored secret
anthropic‑vertex, vertex‑geminiAKS, EKSMounted service account key

IRSA is EKS-only, so Bedrock outside EKS means a stored credential. Azure OpenAI is reached through openai-compatible and behaves the same on any cluster, since it is an API key over HTTPS.

Verifying the configuration

The service logs its choice once at startup:

INFO: AI provider selected {"provider":"anthropic-bedrock",
      "model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","known":true}

known reports whether the model was recognised in the provider's catalogue. A false here is a warning, not an error — a model released after the AI service was built is perfectly valid and still runs. It is worth checking for a typo, or for an identifier belonging to a different provider.

A configuration error prints the variable and the problem, then exits:

Invalid AI provider configuration for AI_PROVIDER=openai:
  OPENAI_API_KEY: Invalid input: expected string, received undefined

Migrating from LiteLLM

Deployments configured with LITELLM_MODEL, LITELLM_BASE_URL and LITELLM_API_KEY keep working with no change — those variables are still accepted and resolve to the openai-compatible provider, which is the same code path they already ran.

They are deprecated, and the service logs a notice at startup naming the replacements:

# before
LITELLM_MODEL=gemini-3.6-flash
LITELLM_BASE_URL=https://litellm.example.com/v1
LITELLM_API_KEY=sk-...

# after
AI_PROVIDER=openai-compatible
AI_MODEL=gemini-3.6-flash
OPENAI_COMPATIBLE_BASE_URL=https://litellm.example.com/v1
OPENAI_COMPATIBLE_API_KEY=sk-...

Setting AI_PROVIDER alongside the legacy variables makes the new configuration win, and the ignored LITELLM_* values are named in a startup warning so the leftover is visible rather than mysterious.