KubeAgent Kubernetes CLI documentation
Everything you need to get KubeAgent running — from installation to terminal approvals and incident alerts.
Written and reviewed by Hadi Farnoud · Updated July 16, 2026
Installation
KubeAgent runs as a CLI. Install it with npx for zero setup, or globally if you prefer.
Run once (no install)
Install globally
npm install -g kubeagent kubeagent --version
kubectl installed and pointing at your cluster, a KubeAgent account (free plan available).
Quick start
The fastest way to get started is npx kubeagent onboard — it handles login, cluster scanning, and knowledge base setup in one step.
Or step-by-step:
- Authenticate:
npx kubeagent login - Scan cluster and build knowledge base:
npx kubeagent onboard - Start the monitoring loop:
npx kubeagent watch - Connect a notification channel (Slack, Discord, Teams, Telegram, PagerDuty, or Webhook) to receive incident and recovery alerts.
# Onboard handles everything at once npx kubeagent onboard Opening browser to complete login… ✓ Logged in as [email protected] ✓ Cluster: hetzner-prod (3 nodes, 12 namespaces) ✓ Detected: 18 deployments across 6 projects ✓ Knowledge base written to ~/.kubeagent/kb/ # Then start watching npx kubeagent watch ✓ Watching cluster hetzner-prod (Ctrl+C to stop) ✓ Issues resolved will be reported to Slack → #oncall
Login & auth
KubeAgent uses a browser-based login flow. Your credentials are stored locally in ~/.kubeagent/auth.json.
kubeagent login # opens browser kubeagent login --device # device-code flow for headless/SSH machines kubeagent logout # removes local credentials
Headless & SSH machines
On headless machines (SSH sessions, containers, servers without a browser), use --device. KubeAgent will display a URL and code — open the URL on any device, enter the code, and your CLI session is authenticated. Headless/SSH environments are auto-detected, so the flag is often applied automatically.
kubeagent status
Runs a one-shot health check and prints all detected issues. Useful in scripts and CI pipelines.
kubeagent status Checking cluster hetzner-prod… [critical] pod_crash_loop: retime-api-worker-6b4d9 has restarted 42 times [warning] pod_oom: dove-worker-7c8f2 was OOM-killed (last 1h) [info] pod_pending: solidtime-worker pending for 2m (scheduling) 3 issues found.
| Flag | Description |
|---|---|
| --context <name> | Use a specific kubectl context (overrides current-context) |
| --namespace <ns> | Limit checks to one namespace |
| --json | Output issues as JSON (for scripting) |
kubeagent onboard
Onboarding scans your cluster topology and any local code repositories you point it at. It then asks a few clarifying questions (powered by AI) and writes a knowledge base to ~/.kubeagent/kb/ that the diagnoser uses as context.
Run onboarding once when you first set up, then again whenever you add new services or significantly change your infrastructure.
kubeagent onboard ✓ Cluster scanned: 3 nodes, 12 namespaces, 18 deployments ✓ Code projects found: retime-api (PHP/Laravel), dove (Go) ✓ Answering questions about your stack… What is the primary purpose of the "dove" service? > Transactional email delivery for all products Which services are most critical (cannot tolerate any downtime)? > dove-web, falcon, retime-api-web ✓ Knowledge base written to ~/.kubeagent/kb/
| Flag | Description |
|---|---|
| --skip-code-scan | Skip local repository scanning (faster, cluster-only KB) |
| --kb-dir <path> | Write knowledge base to a custom directory |
| --context <name> | Use a specific kubectl context |
kubeagent watch
kubeagent watch polls your cluster every 60 seconds. When it finds an issue, it runs the AI diagnoser, proposes a fix, and — if the fix is configured as safe — applies it. Other supported writes require approval in the active terminal. In non-interactive mode they are denied.
✓ Watching cluster hetzner-prod [11:42:01] Detected: pod_crash_loop on retime-api-worker-6b4d9 [11:42:03] Fetching logs… describing pod… checking events… [11:42:09] Root cause: OOM (memory limit 256Mi, heap spiked to 300Mi) [11:42:09] Proposed fix: increase memory limit to 512Mi [11:42:10] Requires approval → waiting in the active terminal [11:43:00] ✓ Approved in terminal — applying set_resources… [11:43:04] ✓ Done. Pod restarted cleanly.
| Flag | Description |
|---|---|
| --interval <sec> | Poll interval in seconds (default: 60) |
| --auto-fix | Apply safe fixes without prompting (use with caution) |
| --no-interactive | Run headless — approval-required actions are denied |
| --context <name> | Use a specific kubectl context |
| --kb-dir <path> | Use knowledge base from a custom directory |
kubeagent demo
Creates an isolated kubeagent-demo namespace with one deliberately broken pod (it crash-loops with a clear error in its logs), waits for KubeAgent to detect it, sends the alert to your connected notification channels, and runs the AI diagnosis. Nothing outside the demo namespace is touched — the namespace is labeled at creation and only ever deleted when it carries that label. Auto-fix is off for the demo run, and the namespace is removed at the end (--keep retains it, kubeagent demo --cleanup removes it any time, --yes skips prompts). Requires login — the diagnosis uses a small amount of your AI credits.
kubeagent diagnose
Runs a one-shot diagnosis on whatever issues currently exist in the cluster. Same AI loop as watch but triggered manually.
kubeagent diagnose Scanning for issues… 2 issues found. Starting diagnosis… ── Issue 1: pod_crash_loop ── Root cause: application startup failure due to missing DB_HOST env var Fix: Secret "retime-env" is missing key DB_HOST — add it and rollout restart Verification: pod should reach Running state within 60s with 0 restarts ── Issue 2: pod_pending ── Root cause: insufficient CPU (requested 2000m, node has 1800m available) Fix: Scale down retime-api-worker to 3 replicas to free headroom
kubeagent query
Use query for ad-hoc questions that don't fit neatly into a diagnosis flow. The AI has full access to the same cluster tools (logs, describe, events) and the knowledge base.
kubeagent query "why is dove-worker using so much memory?" kubeagent query "show me all deployments that haven't restarted in 7 days" kubeagent query "what would happen if I scaled retime-api-web to 1 replica?"
kubeagent onboard after adding new services.
kubeagent scan
Scans a local directory and suggests which subdirectory maps to which deployment in your cluster. Useful for setting up the knowledge base or verifying project-to-deployment mappings.
kubeagent scan ~/Code/devops Scanning ~/Code/devops… retime-api/ → retime-api-web (prod) dove/ → dove-web (prod) offka/ → offka-web (prod) ajimaji-api/ → ajimaji-v1 (ajimaji) 4 matches found.
| Flag | Description |
|---|---|
| --context <name> | Use a specific kubectl context |
kubeagent notify
Add, list, test, and remove notification channels without leaving the terminal.
kubeagent notify add # interactive setup — pick channel type and configure kubeagent notify list # show all configured channels kubeagent notify test # send a test alert to all channels kubeagent notify remove 1 # remove channel by index
Supported channels: Slack, Discord, Microsoft Teams, Telegram, PagerDuty, and custom webhooks. See the notification sections below for detailed setup instructions.
kubeagent account
Shows your current plan, monthly token allocation, remaining balance, and extra credits. Includes a low-balance warning when you're below 20% of your monthly tokens.
kubeagent account Token Balance Plan: Pro Monthly tokens: 75,000 / 100,000 (25% used) Total remaining: 75,000 Resets: May 1, 2026
kubeagent account Token Balance Plan: Starter Monthly tokens: 2,000 / 10,000 (80% used) Total remaining: 2,000 Resets: May 1, 2026 ⚠ Low balance: 20% of monthly tokens remaining. Upgrade or buy extra credits: app.kubeagent.net/billing
Slack setup
When KubeAgent detects an issue, it can send a message to your Slack channel. Slack is an alert destination; approval-required actions are presented in the active CLI terminal.
Connect Slack
- Go to app.kubeagent.net → Settings → Notifications
- Click Connect Slack and authorize the KubeAgent app in your workspace
- Select the channel to post alerts to (e.g.,
#oncallor#infra-alerts) - Click Save
What you'll see
🔴 KubeAgent — Incident detected
Issue: pod_crash_loop on retime-api-worker-6b4d9
Root cause: OOM kill (heap 300Mi > limit 256Mi)
CLI: approval required for set_resources
Return to the active CLI terminal to inspect the exact arguments and approve or deny the supported action. Headless runs deny approval-required actions.
Discord setup
KubeAgent posts alerts to a Discord channel using an Incoming Webhook. No bot to install — just paste the webhook URL.
Connect Discord
- In Discord, open Server Settings → Integrations → Webhooks
- Click New Webhook, pick a channel (e.g.,
#infra-alerts), and copy the webhook URL - Go to app.kubeagent.net → Dashboard
- Paste the URL in the Discord card and click Connect
KubeAgent sends a test message to verify the webhook works. If you see a confirmation in your Discord channel, you're all set.
Microsoft Teams setup
KubeAgent sends alerts to Microsoft Teams using an Incoming Webhook connector.
Connect Teams
- In Teams, open the target channel and click ⋯ → Connectors (or Manage channel → Connectors)
- Search for Incoming Webhook and click Configure
- Give it a name (e.g., "KubeAgent") and copy the webhook URL
- Go to app.kubeagent.net → Dashboard
- Paste the URL in the Teams card and click Connect
KubeAgent sends a test message to verify the webhook works. If you see a card in your Teams channel, you're connected.
Telegram setup
Prefer Telegram? KubeAgent has a bot you can add to any group or use in a private chat for incident and recovery notifications.
Connect Telegram
- Go to app.kubeagent.net → Settings → Notifications
- Click Connect Telegram
- Start a chat with
@KubeAgentBotor add it to a group - Send the one-time verification code shown on the settings page
Approval boundary
Telegram is an alert destination. Approval-required actions remain in the active CLI terminal.
PagerDuty setup
KubeAgent integrates with PagerDuty using the Events API v2. Since KubeAgent is a custom integration, you need to add it manually to your PagerDuty service.
Connect PagerDuty
- Log in to your PagerDuty account.
- Navigate to Services → Service Directory.
- Select an existing service or click + New Service.
- Inside the service, click the Integrations tab.
- Click + Add another integration.
- Search for "Events API v2" and select it, then click Add.
- Copy the 32-character Integration Key.
- Paste it into app.kubeagent.net → Dashboard or run
kubeagent notify add.
How it works
KubeAgent triggers an incident for every detected issue. When the issue is resolved in your cluster (e.g., a pod stops crash-looping), KubeAgent automatically sends a resolve event to PagerDuty to close the incident.
Custom Webhooks
If none of the built-in integrations fit, KubeAgent can POST a JSON payload to any URL you provide — your own API, Zapier, n8n, or any webhook-compatible service.
Connect a Webhook
- Go to app.kubeagent.net → Dashboard
- In the Generic Webhook card, paste your endpoint URL
- Click Connect
Payload format
{
"issue": "pod_crash_loop",
"resource": "retime-api-worker-6b4d9",
"namespace": "prod",
"severity": "critical",
"diagnosis": "OOM kill (heap 300Mi > limit 256Mi)",
"proposedFix": "set memory limit to 512Mi",
"cluster": "hetzner-prod",
"timestamp": "2026-04-09T11:42:10Z"
}
KubeAgent sends a test POST when you first connect. Your endpoint should return a 2xx status to confirm delivery.
Run inside your cluster
The CLI on a laptop stops watching when the laptop sleeps. For 24/7 coverage, run the agent inside the cluster as a Deployment — it uses the pod's service account (no kubeconfig, no onboarding) and reports to the same dashboard and notification channels.
helm install kubeagent ./charts/kubeagent-agent \
--namespace kubeagent --create-namespace \
--set apiKey=<your API key from app.kubeagent.net> \
--set clusterName=my-prod-cluster
kubectl create namespace kubeagent
kubectl -n kubeagent create secret generic kubeagent-agent \
--from-literal=api-key=<your API key>
kubectl apply -f https://kubeagent.net/install/kubeagent-agent.yaml
Configuration is env-based: KUBEAGENT_CLUSTER_NAME (display name in alerts), KUBEAGENT_INTERVAL (seconds, default 300), KUBEAGENT_AUTO_FIX (false = read-only diagnosis; approval-gated actions are always denied in agent mode and surface as notifications). The Helm chart's rbac.readOnly=true value installs a read-only ClusterRole for teams that want zero write access. One replica per cluster — two agents would double-diagnose every incident.
KubeAgent can also detect kubelet node.fs approaching eviction pressure and attribute excessive ephemeral usage to visible pods, container writable layers/logs, or local-volume and other bytes not explained by containers. CLI checks use the current kubectl identity. In-cluster collection defaults off because Kubernetes has no stats-only API-server RBAC: get nodes/proxy can reach other kubelet GET endpoints and may bypass normal admission controls. Fully trusted operators can opt in with Helm --set ephemeralStorage.enabled=true; the one-file install stays disabled. Existing checks continue when stats are unavailable.
CI/CD API
KubeAgent exposes a public API for use in GitHub Actions, GitLab CI, and other CI/CD pipelines. Authenticate with your API key (Bearer scheme).
POST /v1/checks/run
Submit a Kubernetes resource snapshot and get back health findings with an overall status (healthy, degraded, or critical).
curl -X POST https://api.kubeagent.net/v1/checks/run \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{"resources": [...]}'
GET /v1/incidents
Query recent incidents for pipeline gate decisions. Use the since parameter to filter by time.
curl https://api.kubeagent.net/v1/incidents?since=2026-04-15T00:00:00Z \ -H "Authorization: Bearer YOUR_API_KEY"
API Documentation
Interactive docs and the full OpenAPI spec are available at:
api.kubeagent.net/v1/docs— interactive API explorerapi.kubeagent.net/v1/openapi.json— OpenAPI 3.0 spec
Knowledge base
The knowledge base (KB) is a set of Markdown files in ~/.kubeagent/kb/. The AI diagnoser loads them as system context before every diagnosis run.
What's in the KB
cluster.md— nodes, namespaces, cluster contextservices.md— your deployments, languages, frameworks, criticalitynotes.md— your answers from the onboarding interviewprojects/*.md— per-project code scan summaries
Editing the KB manually
All files are plain Markdown. You can edit them directly to add context the scanner missed — known quirks, deployment runbooks, escalation contacts.
# Service notes ## puffin - PHP 7.2 — do NOT upgrade to 7.3+, breaks Yii 1.x compat - High memory usage during campaign sends is expected - Scheduled jobs in "puffin" namespace are maintenance, not errors ## dove - OOM on dove-worker during large batch sends is a known issue - Escalate to #pakat-oncall, not #general
Re-running onboard
Run kubeagent onboard again at any time to refresh the KB. Your manual edits to notes.md are preserved.
Action safety tiers
Every action the diagnoser can take is classified into one of three tiers. This determines whether it runs automatically, requires your approval, or is never automated.
| Action | Tier | Description |
|---|---|---|
| get_logs | safe | Read pod logs — always automatic, no approval needed |
| describe_resource | safe | Describe any resource — always automatic |
| get_events | safe | Fetch namespace events — always automatic |
| rollout_restart | approval | Trigger a rolling restart — requires user approval |
| restart_pod | approval | Delete and recreate a pod — requires user approval |
| set_resources | approval | Change CPU/memory limits — requires user approval |
| scale_deployment | approval | Change replica count — scaling to 0 always requires approval |
| Rollback, delete namespace, etc. | unavailable | Outside KubeAgent's supported action surface |
Promoting actions to safe
You can mark additional actions as safe (no approval needed) via the dashboard under Settings → Action Policy, or by passing --safe-actions rollout_restart,restart_pod to any command.
llms.txt
KubeAgent publishes llms.txt — a standard that helps AI assistants like ChatGPT, Claude, and Perplexity understand what KubeAgent does and recommend it accurately.
/llms.txt— concise overview with links to key pages/llms-full.txt— complete documentation inlined for deep AI context/docs.md— markdown version of this docs page
/llms.txt and AI crawlers use it to accurately describe and recommend your product. Learn more at llmstxt.org.