Watching · 3 clusters · 47 namespaces

Kubernetes incidents fix themselves
while you sleep.

KubeAgent runs in your cluster or locally, watches every pod around the clock, and diagnoses issues through an agentic kubectl loop. Safe fixes apply on their own. Risky ones ping you — Slack, Discord, Teams, Telegram, or PagerDuty.

Start free →
npm install -g kubeagent
✓ Monitoring & alerts free, forever ✓ 200K AI credits — never expire No credit card
1 kubeagent login
2 kubeagent onboard
3 kubeagent watch
hetzner-prod · live WATCHING

Monitor. Diagnose.
Remediate.

  1. 01 — MONITOR

    Continuous cluster polling

    Watches pods, nodes, and deployments on a continuous tick. Runs in your cluster on its own service account, or locally using your kubectl context. Detects CrashLoopBackOff, OOMKilled, ImagePullBackOff, node pressure, and more.

  2. 02 — DIAGNOSE

    KubeAgent investigates

    When an issue is found, KubeAgent fetches logs, events, and your cluster's knowledge base. It traces the root cause through an agentic loop of kubectl lookups.

  3. 03 — REMEDIATE

    Fix or ask — on your terms

    Configured safe actions can run automatically. Other supported writes pause for approval in the active terminal; unattended runs deny them. Notification channels carry incident alerts.

Built for engineers tired of being on-call.

Intelligent observability on top of your existing cluster. Works with any distribution kubectl can reach.

Read the docs →
✦ Feature

Agentic Diagnosis

KubeAgent runs kubectl tools in a loop — logs, describe, events — until it finds the root cause. Powered by Claude AI.

✦ Feature

Multi-Channel Alerts

Connect Slack, Discord, Microsoft Teams, Telegram, PagerDuty, or custom Webhooks for incident alerts and recovery updates.

✦ Feature

Safe Action Tiers

Read-only tools run automatically. Supported writes follow your safe-action policy or require terminal approval; unsupported operations are unavailable.

✦ Feature

Cluster Knowledge Base

Onboard once — KubeAgent learns your services, tech stacks, and infra. That context is injected into every diagnosis.

✦ Feature

Incident Logs

Every incident saved with root cause, fix, and outcome. Feeds back into the knowledge base so diagnoses improve over time.

✦ Feature

Zero-Agent Setup

Install with npx. No agents to deploy, no sidecars, no cluster-wide service accounts. Login works on headless machines via device-code flow.

Get alerted where you already work.

Connect once in the dashboard. KubeAgent sends incident alerts straight to your team's channel — no polling, no extra tab. Slack can also trigger an AI diagnosis with one click.

Slack

One-click AI diagnosis

Discord

Server webhooks

Microsoft Teams

Channel connector

Telegram

Bot notifications

PagerDuty

Incident escalation

Webhooks

Send anywhere else

Datadog costs 10x more and still pages you at 2am.
KubeAgent watches your clusters, diagnoses root causes,
and applies safe remediations — automatically.
Install in 5 minutes. Sleep through the night.

Other tools
  • Alert-only — tells you something broke
  • You diagnose, you fix, you get paged at 2am
  • $200–$500/mo for basic Kubernetes observability
  • Complex agent setup, cluster-wide permissions
KubeAgent
  • Diagnoses root causes with AI, applies safe fixes
  • Only pages you when a human decision is required
  • Starts free — Growth plan from $29.99/mo
  • CLI-only, zero cluster credentials leave your machine

Zero-Trust Access.
Your Cluster, Your Rules.

Unlike other monitoring tools, KubeAgent never asks for your Kubeconfig or cluster credentials. It runs locally as a CLI, using your existing terminal context.

✓ Local Context Credentials never leave your machine or network.
✓ No Direct Access We never connect to your cluster from our servers.
✓ Safe-Only Actions You control exactly what KubeAgent can and cannot do.
✓ No Incident Data Stored Selected evidence is processed for diagnosis, but request and response content is not logged or retained.
╭──────────────────────────╮
│   YOUR NETWORK           │
│   kubeconfig · kubectl   │
│                          │
│   ┌────────────────┐     │
│   │ kubeagent CLI  │     │
│   └────────┬───────┘     │
│            │             │
╰────────────┼─────────────╯
             │   OUTBOUND ONLY
             ▼
     ┌───────────────┐
     │ api.kubeagent │
     │ alert channels│
     └───────────────┘

     ✓  NO inbound cluster access
     ✓  NO stored credentials
     ✓  NO long-lived tokens

Start free.
Scale when your cluster does.

Monitoring and alerts are free forever. Every account starts with 200K AI credits for diagnosis and auto-remediation — they never expire. No credit card required. All plans include Slack, Discord, Teams, Telegram & PagerDuty integration and full dashboard access.

Free
$0 forever
200K AI credits · never expire
  • Monitoring & alerts, free forever
  • ~40 AI incident diagnoses
  • Slack, Discord, Teams & more
  • Dashboard access
Get started
Growth
$29.99/mo
3M tokens · 1–3 clusters
  • ~150 incidents/mo
  • Slack, Discord, Teams & more
  • Dashboard access
  • Extra credits $6/1M
Start Growth
Max
$199.99/mo
20M tokens · 10+ clusters
  • ~1,200 incidents/mo
  • Slack, Discord, Teams & more
  • Dashboard access
  • Extra credits $6/1M
Start Max

Questions.

More details in the docs.

Does KubeAgent have direct access to my cluster?
No. KubeAgent runs as a CLI tool in your environment (laptop, jumpbox, or CI). It uses your local kubeconfig to communicate with your cluster. Your credentials never leave your machine or network. AI requests are routed through our servers for billing, but we do not log or store the content. Nothing is retained unless you explicitly opt in for edge-case debugging.
Do I need to configure anything before using it?
Sign up, run kubeagent login, then kubeagent onboard to scan your cluster and build a knowledge base. After that, kubeagent watch starts monitoring. Takes about 5 minutes.
How do action approvals work?
When KubeAgent proposes a supported write that is not in your safe-action policy, the active CLI shows the exact command and asks for approval in the terminal. Non-interactive runs deny approval-gated actions. Slack, Discord, Teams, Telegram, PagerDuty, and webhooks receive incident alerts, not remote execution controls.
Is it safe to run in production?
KubeAgent exposes a bounded tool set. Read-only tools run automatically; supported writes follow the configured safe-action policy or require terminal approval. Unknown and unsupported operations are rejected. Use least-privilege Kubernetes credentials and test the policy in a non-production namespace first.
Which Kubernetes distributions are supported?
Any cluster accessible via kubectl — EKS, GKE, AKS, k3s, RKE2, bare-metal. If kubectl works, KubeAgent works.
Does KubeAgent collect my incident data?
No. AI requests are routed through our servers so we can track token usage for billing, but we do not log or store the content of your requests or responses. We only record anonymous usage metadata (token counts, timestamps). If you encounter an edge case and want us to help debug, you can explicitly share incident details with our support team — but that's entirely opt-in and never automatic.
What does token usage mean?
Each incident diagnosis consumes tokens from your plan. A typical incident uses ~5,000 tokens. Plan limits reflect how many incidents you can diagnose per month. Run out? Top up with extra credits at $6 per 1M tokens.