Why we built KubeAgent — and what's on this blog
Most Kubernetes monitoring tools wake you at 2 AM and leave you to diagnose. KubeAgent investigates, fixes safe issues itself, and pages you only for the rest.
If you run a Kubernetes cluster small enough that you’re personally on the pager, you already know the punchline of every monitoring product demo: it tells you something is broken, then leaves you to figure out what is broken at 2 AM. That gap — between “alert fired” and “service restored” — is where on-call burnout lives.
We built KubeAgent to close it. It runs as a CLI on your machine (or a jumpbox, or CI), polls your cluster every 60 seconds through your existing kubectl context, and when something breaks it doesn’t just send a notification. It investigates, in a loop: pulls logs, describes resources, checks recent events, cross-references the knowledge base you built during onboarding, and proposes a supported fix. Actions configured as safe can apply automatically. Other supported writes wait for approval in the active terminal; headless runs deny them. Slack, Discord, Teams, Telegram, PagerDuty, and webhooks carry incident and recovery alerts.
What you’ll find on this blog
A few things we want to write about, in roughly the order we think we’ll get to them:
- Incident walkthroughs. Real diagnoses from real clusters —
CrashLoopBackOfffrom an OOMKilled pod, an ingress that started 502-ing because a readiness probe timed out, a Postgres pool that exhausted itself under a traffic spike. We’ll show the diagnostic loop step by step. - Patterns in agentic diagnosis. Why we limit the agent to safe
kubectlverbs by default, how we structure the knowledge base so the AI has the right context without burning tokens, what kinds of root causes the model gets right vs. wrong, and how we test it without a live cluster. - The on-call economics piece. Most teams under 50 engineers can’t justify a $500/month observability bill. We’ll publish the math on what KubeAgent actually costs to run an incident through — and how that compares to a paged human.
- Field notes. Short posts when we ship something interesting: new notification channels, new safety tiers, a clever way someone in the community used
kubeagent query, OSS app CVE detection updates.
Who we’re writing for
If you’re the person who deploys to production AND wears the pager — solo founder, small DevOps team, platform team of two — this is for you. We’re not going to write listicles or “Top 10 Kubernetes Best Practices in 2026” posts. The Kubernetes ecosystem doesn’t need more of those.
Try it
If you haven’t already, the install is a one-liner:
npm install -g kubeagent
kubeagent login
kubeagent onboard
kubeagent watch
You get a free plan with 200K AI credits, no credit card. The CLI is fully local — your kubeconfig and cluster credentials never leave your machine.
Subscribe to the RSS feed if you’d like to follow along.
If you are new to the workflow, start with the Kubernetes troubleshooting guide, then read how KubeAgent separates safe automation from approval-required actions.