Installation

KubeAgent runs as a CLI. Install it with npx for zero setup, or globally if you prefer.

Run once (no install)

$ npx kubeagent watch

Install globally

shell
npm install -g kubeagent
kubeagent --version
Prerequisites: Node.js ≥ 22, kubectl installed and pointing at your cluster, a KubeAgent account (free plan available).

Quick start

The fastest way to get started is npx kubeagent onboard — it handles login, cluster scanning, and knowledge base setup in one step.

$ npx kubeagent onboard

Or step-by-step:

  1. Authenticate: npx kubeagent login
  2. Scan cluster and build knowledge base: npx kubeagent onboard
  3. Start the monitoring loop: npx kubeagent watch
  4. Connect a notification channel (Slack, Discord, Teams, Telegram, PagerDuty, or Webhook) to receive incident and recovery alerts.
shell — first-time setup
# Onboard handles everything at once
npx kubeagent onboard
  Opening browser to complete login…
  ✓ Logged in as [email protected]
  ✓ Cluster: hetzner-prod (3 nodes, 12 namespaces)
  ✓ Detected: 18 deployments across 6 projects
  ✓ Knowledge base written to ~/.kubeagent/kb/

# Then start watching
npx kubeagent watch
  ✓ Watching cluster hetzner-prod  (Ctrl+C to stop)
  ✓ Issues resolved will be reported to Slack → #oncall

Login & auth

KubeAgent uses a browser-based login flow. Your credentials are stored locally in ~/.kubeagent/auth.json.

shell
kubeagent login          # opens browser
kubeagent login --device  # device-code flow for headless/SSH machines
kubeagent logout         # removes local credentials

Headless & SSH machines

On headless machines (SSH sessions, containers, servers without a browser), use --device. KubeAgent will display a URL and code — open the URL on any device, enter the code, and your CLI session is authenticated. Headless/SSH environments are auto-detected, so the flag is often applied automatically.

Note: KubeAgent requires a login to use AI features. There is no direct API key mode — all inference runs through the KubeAgent backend.

kubeagent status

Runs a one-shot health check and prints all detected issues. Useful in scripts and CI pipelines.

shell
kubeagent status
  Checking cluster hetzner-prod…

  [critical] pod_crash_loop: retime-api-worker-6b4d9 has restarted 42 times
  [warning]  pod_oom: dove-worker-7c8f2 was OOM-killed (last 1h)
  [info]     pod_pending: solidtime-worker pending for 2m (scheduling)

  3 issues found.
FlagDescription
--context <name>Use a specific kubectl context (overrides current-context)
--namespace <ns>Limit checks to one namespace
--jsonOutput issues as JSON (for scripting)

kubeagent onboard

Onboarding scans your cluster topology and any local code repositories you point it at. It then asks a few clarifying questions (powered by AI) and writes a knowledge base to ~/.kubeagent/kb/ that the diagnoser uses as context.

Run onboarding once when you first set up, then again whenever you add new services or significantly change your infrastructure.

shell
kubeagent onboard
  ✓ Cluster scanned: 3 nodes, 12 namespaces, 18 deployments
  ✓ Code projects found: retime-api (PHP/Laravel), dove (Go)
  ✓ Answering questions about your stack…

  What is the primary purpose of the "dove" service?
  > Transactional email delivery for all products

  Which services are most critical (cannot tolerate any downtime)?
  > dove-web, falcon, retime-api-web

  ✓ Knowledge base written to ~/.kubeagent/kb/
FlagDescription
--skip-code-scanSkip local repository scanning (faster, cluster-only KB)
--kb-dir <path>Write knowledge base to a custom directory
--context <name>Use a specific kubectl context

kubeagent watch

kubeagent watch polls your cluster every 60 seconds. When it finds an issue, it runs the AI diagnoser, proposes a fix, and — if the fix is configured as safe — applies it. Other supported writes require approval in the active terminal. In non-interactive mode they are denied.

$ npx kubeagent watch
shell — example output
  ✓ Watching cluster hetzner-prod
  [11:42:01] Detected: pod_crash_loop on retime-api-worker-6b4d9
  [11:42:03] Fetching logs… describing pod… checking events…
  [11:42:09] Root cause: OOM (memory limit 256Mi, heap spiked to 300Mi)
  [11:42:09] Proposed fix: increase memory limit to 512Mi
    [11:42:10] Requires approval → waiting in the active terminal
  [11:43:00] ✓ Approved in terminal — applying set_resources…
    [11:43:04] ✓ Done. Pod restarted cleanly.
FlagDescription
--interval <sec>Poll interval in seconds (default: 60)
--auto-fixApply safe fixes without prompting (use with caution)
--no-interactiveRun headless — approval-required actions are denied
--context <name>Use a specific kubectl context
--kb-dir <path>Use knowledge base from a custom directory

kubeagent demo

Creates an isolated kubeagent-demo namespace with one deliberately broken pod (it crash-loops with a clear error in its logs), waits for KubeAgent to detect it, sends the alert to your connected notification channels, and runs the AI diagnosis. Nothing outside the demo namespace is touched — the namespace is labeled at creation and only ever deleted when it carries that label. Auto-fix is off for the demo run, and the namespace is removed at the end (--keep retains it, kubeagent demo --cleanup removes it any time, --yes skips prompts). Requires login — the diagnosis uses a small amount of your AI credits.

kubeagent diagnose

Runs a one-shot diagnosis on whatever issues currently exist in the cluster. Same AI loop as watch but triggered manually.

shell
kubeagent diagnose
  Scanning for issues…
  2 issues found. Starting diagnosis…

  ── Issue 1: pod_crash_loop ──
  Root cause: application startup failure due to missing DB_HOST env var
  Fix: Secret "retime-env" is missing key DB_HOST — add it and rollout restart

  Verification: pod should reach Running state within 60s with 0 restarts

  ── Issue 2: pod_pending ──
  Root cause: insufficient CPU (requested 2000m, node has 1800m available)
  Fix: Scale down retime-api-worker to 3 replicas to free headroom

kubeagent query

Use query for ad-hoc questions that don't fit neatly into a diagnosis flow. The AI has full access to the same cluster tools (logs, describe, events) and the knowledge base.

shell
kubeagent query "why is dove-worker using so much memory?"
kubeagent query "show me all deployments that haven't restarted in 7 days"
kubeagent query "what would happen if I scaled retime-api-web to 1 replica?"
Tip: Queries use your knowledge base for context. The more complete your KB, the better the answers. Run kubeagent onboard after adding new services.

kubeagent scan

Scans a local directory and suggests which subdirectory maps to which deployment in your cluster. Useful for setting up the knowledge base or verifying project-to-deployment mappings.

shell
kubeagent scan ~/Code/devops
  Scanning ~/Code/devops…

  retime-api/       → retime-api-web (prod)
  dove/             → dove-web (prod)
  offka/            → offka-web (prod)
  ajimaji-api/      → ajimaji-v1 (ajimaji)

  4 matches found.
FlagDescription
--context <name>Use a specific kubectl context

kubeagent notify

Add, list, test, and remove notification channels without leaving the terminal.

shell
kubeagent notify add        # interactive setup — pick channel type and configure
kubeagent notify list       # show all configured channels
kubeagent notify test       # send a test alert to all channels
kubeagent notify remove 1   # remove channel by index

Supported channels: Slack, Discord, Microsoft Teams, Telegram, PagerDuty, and custom webhooks. See the notification sections below for detailed setup instructions.

kubeagent account

Shows your current plan, monthly token allocation, remaining balance, and extra credits. Includes a low-balance warning when you're below 20% of your monthly tokens.

shell
kubeagent account

  Token Balance
    Plan:              Pro
    Monthly tokens:    75,000 / 100,000 (25% used)
    Total remaining:   75,000
    Resets:            May 1, 2026
shell — low balance warning
kubeagent account

  Token Balance
    Plan:              Starter
    Monthly tokens:    2,000 / 10,000 (80% used)
    Total remaining:   2,000
    Resets:            May 1, 2026

    ⚠ Low balance: 20% of monthly tokens remaining.
    Upgrade or buy extra credits: app.kubeagent.net/billing

Slack setup

When KubeAgent detects an issue, it can send a message to your Slack channel. Slack is an alert destination; approval-required actions are presented in the active CLI terminal.

Connect Slack

  1. Go to app.kubeagent.net → Settings → Notifications
  2. Click Connect Slack and authorize the KubeAgent app in your workspace
  3. Select the channel to post alerts to (e.g., #oncall or #infra-alerts)
  4. Click Save

What you'll see

Slack message — incident alert
  🔴 KubeAgent — Incident detected

  Issue: pod_crash_loop on retime-api-worker-6b4d9
  Root cause: OOM kill (heap 300Mi > limit 256Mi)
  CLI: approval required for set_resources

Return to the active CLI terminal to inspect the exact arguments and approve or deny the supported action. Headless runs deny approval-required actions.

Discord setup

KubeAgent posts alerts to a Discord channel using an Incoming Webhook. No bot to install — just paste the webhook URL.

Connect Discord

  1. In Discord, open Server Settings → Integrations → Webhooks
  2. Click New Webhook, pick a channel (e.g., #infra-alerts), and copy the webhook URL
  3. Go to app.kubeagent.net → Dashboard
  4. Paste the URL in the Discord card and click Connect

KubeAgent sends a test message to verify the webhook works. If you see a confirmation in your Discord channel, you're all set.

Microsoft Teams setup

KubeAgent sends alerts to Microsoft Teams using an Incoming Webhook connector.

Connect Teams

  1. In Teams, open the target channel and click ⋯ → Connectors (or Manage channel → Connectors)
  2. Search for Incoming Webhook and click Configure
  3. Give it a name (e.g., "KubeAgent") and copy the webhook URL
  4. Go to app.kubeagent.net → Dashboard
  5. Paste the URL in the Teams card and click Connect

KubeAgent sends a test message to verify the webhook works. If you see a card in your Teams channel, you're connected.

Telegram setup

Prefer Telegram? KubeAgent has a bot you can add to any group or use in a private chat for incident and recovery notifications.

Connect Telegram

  1. Go to app.kubeagent.net → Settings → Notifications
  2. Click Connect Telegram
  3. Start a chat with @KubeAgentBot or add it to a group
  4. Send the one-time verification code shown on the settings page

Approval boundary

Telegram is an alert destination. Approval-required actions remain in the active CLI terminal.

Multiple channels: You can connect any combination of Slack, Discord, Teams, Telegram, PagerDuty, and Webhooks. Alerts fan out to all connected channels; they do not act as remote approval controls.

PagerDuty setup

KubeAgent integrates with PagerDuty using the Events API v2. Since KubeAgent is a custom integration, you need to add it manually to your PagerDuty service.

Connect PagerDuty

  1. Log in to your PagerDuty account.
  2. Navigate to Services → Service Directory.
  3. Select an existing service or click + New Service.
  4. Inside the service, click the Integrations tab.
  5. Click + Add another integration.
  6. Search for "Events API v2" and select it, then click Add.
  7. Copy the 32-character Integration Key.
  8. Paste it into app.kubeagent.net → Dashboard or run kubeagent notify add.
Note: Do not use a REST API key (20 characters). KubeAgent requires an Integration Key (32 characters) to send events.

How it works

KubeAgent triggers an incident for every detected issue. When the issue is resolved in your cluster (e.g., a pod stops crash-looping), KubeAgent automatically sends a resolve event to PagerDuty to close the incident.

Custom Webhooks

If none of the built-in integrations fit, KubeAgent can POST a JSON payload to any URL you provide — your own API, Zapier, n8n, or any webhook-compatible service.

Connect a Webhook

  1. Go to app.kubeagent.net → Dashboard
  2. In the Generic Webhook card, paste your endpoint URL
  3. Click Connect

Payload format

JSON — webhook payload
{
  "issue": "pod_crash_loop",
  "resource": "retime-api-worker-6b4d9",
  "namespace": "prod",
  "severity": "critical",
  "diagnosis": "OOM kill (heap 300Mi > limit 256Mi)",
  "proposedFix": "set memory limit to 512Mi",
  "cluster": "hetzner-prod",
  "timestamp": "2026-04-09T11:42:10Z"
}

KubeAgent sends a test POST when you first connect. Your endpoint should return a 2xx status to confirm delivery.

Run inside your cluster

The CLI on a laptop stops watching when the laptop sleeps. For 24/7 coverage, run the agent inside the cluster as a Deployment — it uses the pod's service account (no kubeconfig, no onboarding) and reports to the same dashboard and notification channels.

Helm
helm install kubeagent ./charts/kubeagent-agent \
  --namespace kubeagent --create-namespace \
  --set apiKey=<your API key from app.kubeagent.net> \
  --set clusterName=my-prod-cluster
kubectl apply
kubectl create namespace kubeagent
kubectl -n kubeagent create secret generic kubeagent-agent \
  --from-literal=api-key=<your API key>
kubectl apply -f https://kubeagent.net/install/kubeagent-agent.yaml

Configuration is env-based: KUBEAGENT_CLUSTER_NAME (display name in alerts), KUBEAGENT_INTERVAL (seconds, default 300), KUBEAGENT_AUTO_FIX (false = read-only diagnosis; approval-gated actions are always denied in agent mode and surface as notifications). The Helm chart's rbac.readOnly=true value installs a read-only ClusterRole for teams that want zero write access. One replica per cluster — two agents would double-diagnose every incident.

KubeAgent can also detect kubelet node.fs approaching eviction pressure and attribute excessive ephemeral usage to visible pods, container writable layers/logs, or local-volume and other bytes not explained by containers. CLI checks use the current kubectl identity. In-cluster collection defaults off because Kubernetes has no stats-only API-server RBAC: get nodes/proxy can reach other kubelet GET endpoints and may bypass normal admission controls. Fully trusted operators can opt in with Helm --set ephemeralStorage.enabled=true; the one-file install stays disabled. Existing checks continue when stats are unavailable.

CI/CD API

KubeAgent exposes a public API for use in GitHub Actions, GitLab CI, and other CI/CD pipelines. Authenticate with your API key (Bearer scheme).

POST /v1/checks/run

Submit a Kubernetes resource snapshot and get back health findings with an overall status (healthy, degraded, or critical).

shell — run a health check
curl -X POST https://api.kubeagent.net/v1/checks/run \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"resources": [...]}'

GET /v1/incidents

Query recent incidents for pipeline gate decisions. Use the since parameter to filter by time.

shell — query recent incidents
curl https://api.kubeagent.net/v1/incidents?since=2026-04-15T00:00:00Z \
  -H "Authorization: Bearer YOUR_API_KEY"

API Documentation

Interactive docs and the full OpenAPI spec are available at:

  • api.kubeagent.net/v1/docs — interactive API explorer
  • api.kubeagent.net/v1/openapi.json — OpenAPI 3.0 spec
GitHub Actions: A ready-to-use workflow is included — see the OpenAPI spec for the full schema and a copy-paste GitHub Actions example.

Knowledge base

The knowledge base (KB) is a set of Markdown files in ~/.kubeagent/kb/. The AI diagnoser loads them as system context before every diagnosis run.

What's in the KB

  • cluster.md — nodes, namespaces, cluster context
  • services.md — your deployments, languages, frameworks, criticality
  • notes.md — your answers from the onboarding interview
  • projects/*.md — per-project code scan summaries

Editing the KB manually

All files are plain Markdown. You can edit them directly to add context the scanner missed — known quirks, deployment runbooks, escalation contacts.

~/.kubeagent/kb/notes.md (example)
# Service notes

## puffin
- PHP 7.2 — do NOT upgrade to 7.3+, breaks Yii 1.x compat
- High memory usage during campaign sends is expected
- Scheduled jobs in "puffin" namespace are maintenance, not errors

## dove
- OOM on dove-worker during large batch sends is a known issue
- Escalate to #pakat-oncall, not #general

Re-running onboard

Run kubeagent onboard again at any time to refresh the KB. Your manual edits to notes.md are preserved.

Action safety tiers

Every action the diagnoser can take is classified into one of three tiers. This determines whether it runs automatically, requires your approval, or is never automated.

Action Tier Description
get_logs safe Read pod logs — always automatic, no approval needed
describe_resource safe Describe any resource — always automatic
get_events safe Fetch namespace events — always automatic
rollout_restart approval Trigger a rolling restart — requires user approval
restart_pod approval Delete and recreate a pod — requires user approval
set_resources approval Change CPU/memory limits — requires user approval
scale_deployment approval Change replica count — scaling to 0 always requires approval
Rollback, delete namespace, etc. unavailable Outside KubeAgent's supported action surface

Promoting actions to safe

You can mark additional actions as safe (no approval needed) via the dashboard under Settings → Action Policy, or by passing --safe-actions rollout_restart,restart_pod to any command.

Caution: Promoting mutating actions to "safe" means KubeAgent will apply them without asking. Only do this for actions you're confident are safe for your cluster.

llms.txt

KubeAgent publishes llms.txt — a standard that helps AI assistants like ChatGPT, Claude, and Perplexity understand what KubeAgent does and recommend it accurately.

  • /llms.txt — concise overview with links to key pages
  • /llms-full.txt — complete documentation inlined for deep AI context
  • /docs.md — markdown version of this docs page
llms.txt

      
What is llms.txt? A community standard for helping AI models understand your site. Place a curated markdown file at /llms.txt and AI crawlers use it to accurately describe and recommend your product. Learn more at llmstxt.org.