Anatomy of an auto-remediated CrashLoopBackOff
From the third pod restart to a stabilised memory limit — the kubectl calls, safety classification, terminal approval, and verification in order.
The most common incident KubeAgent sees in customer clusters is also the most boring: a pod gets OOMKilled, the controller restarts it, the new pod also gets OOMKilled, and the kubelet eventually flips the status to CrashLoopBackOff. Boring, but at scale it’s the single biggest source of unnecessary 2am pages. This post walks through one such incident end to end — what kubeagent watch saw, what it did, what got auto-applied, and what got escalated.
The names are anonymised; everything else (the verbs, the order, the safety classification) is exactly what the agent does.
The timeline
kubeagent watch is running on a small VM inside the same VPC as the cluster, polling on the default 60-second tick.
T+0s — detection. The watch loop calls the equivalent of kubectl get pods -n retime-api -o json. One pod (retime-api-7d8c4f9b6c-x9k2p) has restartCount: 3, the latest container state: waiting, and reason: CrashLoopBackOff. The previous termination reason was OOMKilled. This is enough to fire an incident — the watch loop hands it to the diagnoser.
T+2s — context gathering. The diagnoser runs four read-only calls in parallel:
kubectl describe pod retime-api-7d8c4f9b6c-x9k2p -n retime-api— to see the latest events, the container’s resource limits, and the controller reference.kubectl logs retime-api-7d8c4f9b6c-x9k2p -n retime-api --previous --tail=200— the logs from the killed container, not the one currently looping.kubectl get events -n retime-api --sort-by=.lastTimestamp— recent events for the namespace.- The local knowledge base under
~/.kubeagent/clusters/hetzner-prod/projects/retime-api.md— anything the operator recorded during onboarding about this service.
All four are read-only and classified as Safe, so they happen without prompting.
T+4s — diagnosis. The collected context goes to Claude with the system prompt that scopes it to “diagnose Kubernetes incidents using only the data you’ve been given.” The response:
The pod is being OOMKilled at a memory limit of 128Mi. The previous container’s heap dump shows JVM resident memory peaking at 312MB just before termination; the workload appears to be a Spring Boot service that’s stable at ~280MB under normal load. The CrashLoopBackOff is a downstream effect, not the root cause. Recommended action: raise the memory limit on the Deployment to 512Mi (with a request of 384Mi to match real working set), then
rollout restartthe Deployment to pick up the change.
That diagnosis is presented to the action planner along with the proposed fix.
T+5s — action classification. The fix has two steps:
kubectl patch deployment retime-api -n retime-api ...— editing a Deployment’s pod spec falls into the Risky tier.kubectl rollout restart deployment/retime-api -n retime-api— Safe.
Because step 1 is Risky, the active terminal presents the diagnosis summary and exact supported arguments, then asks the operator to approve or deny the action. Slack can receive the incident alert, but it is not the approval control.
T+47s — human in the loop. A platform engineer reviews the request in the active terminal and chooses Approve. KubeAgent then passes the already validated arguments to the executor.
T+48s — execution. The patch goes out via the local kubectl context. The Deployment’s spec.template.spec.containers[0].resources is updated. The Safe-tier rollout restart fires immediately after.
T+92s — verification. The watch loop’s next poll sees the new ReplicaSet has rolled out, all pods are Ready, restartCount is 0, and no new events have fired. The incident is marked resolved and an outcome message lands back in #oncall:
✓ retime-api recovered. Memory limit 128Mi → 512Mi, request 64Mi → 384Mi. 1 minute 32 seconds from detection to ready.
T+93s — record. The full incident — initial signal, the data fetched, the diagnosis, the proposed and approved actions, the outcome — is written to ~/.kubeagent/clusters/hetzner-prod/incidents/2026-05-28T03-14-...md. The next time the diagnoser sees an OOMKilled in retime-api, this incident is in its context. After three or four similar incidents, the diagnoser stops proposing the bump from 128 → 512 (it already knows the real working set) and starts proposing the right limit on the first try.
What didn’t happen
A few things the agent deliberately did not do:
- It didn’t restart the pod immediately. A pod-level restart on a CrashLoopBackOff doesn’t fix the underlying limit — it just resets the backoff timer. The agent recognised that and skipped the obvious-but-wrong fix.
- It didn’t roll back the Deployment. Rollback is sometimes the right call for a CrashLoopBackOff (if a recent deploy introduced a memory regression), but the previous termination reason was clearly resource-pressure, not a code change. The local knowledge base recorded that the last deploy to
retime-apiwas six days ago. - It didn’t ask a human to confirm the read-only data gathering. Pulling logs and events on a Pod is non-destructive; gating it behind an approval would just slow down every incident.
What’s “real” about this
Three things, since the post is making specific claims:
- The watch loop, polling interval, and read-only-data-gathering are exactly what
kubeagent watchdoes today. - The action classification (
patch→ Risky,rollout restart→ Safe) is enforced by the same allowlist described in the three-tier safety post. - The incident log + knowledge-base feedback loop is the same one populated by the
incidents/directory under~/.kubeagent/clusters/<context>/.
The specific patch values and the 92-second timeline are illustrative — your cluster’s API-server latency and the operator’s terminal response time will dominate the wall-clock — but the order of operations is exactly what runs.
Try it on your own cluster
npm install -g kubeagent
kubeagent login
kubeagent onboard # scan cluster, build knowledge base
kubeagent watch # start the loop
If you’d rather see it react to a fake OOMKill first, you can kubectl run a memory-hungry test pod with a tiny limit in a non-prod namespace — KubeAgent will pick it up on the next 60-second tick. Start free here.