the engineering log

Reducing Kubernetes MTTR with an approval-gated AI agent

How an approval-gated Kubernetes agent can gather evidence, propose a bounded change, execute through typed tools, and verify the result.

This entry covers the open-source Skyflo Kubernetes agent, a separate Apache-2.0 project. The AI engineering harness is here.

·8 min read·mttrkubernetesincident-responsesre

Why MTTR Is Still Measured in Hours

The Kubernetes ecosystem has a mature set of observability and incident-response tools: metrics, logs, tracing, alerting, deployment history, and the Kubernetes API itself.

More signals do not automatically produce a faster or safer diagnosis. Operators still have to correlate them and decide what evidence is sufficient to justify a change.

The problem isn't any individual tool. The problem is the space between tools.


The incident-response loop

A Kubernetes alert usually starts a sequence of separate tasks: identify the affected workload, inspect current state and events, check recent deployment history, form a hypothesis, propose a change, execute it, and verify the outcome. The elapsed time varies widely by incident. The recurring cost is the handoff between tools and the need to preserve enough evidence for an operator to make a safe decision.

An operations agent can reduce that manual coordination. It does not make the diagnosis correct by default, and it should not turn a model suggestion directly into a cluster mutation. The useful design is a bounded loop in which reads gather evidence, the operator reviews the proposed mutation, typed tools execute only the approved scope, and a separate phase checks the resulting state.


Where an agent can help

The opportunity is not faster typing. It is keeping evidence collection, the proposed change, operator approval, execution, and verification in one inspectable workflow. The benefit has to be measured against real incidents; it should not be inferred from a scripted walkthrough.


How AI Agents Reduce Each Phase

Phase 1: Detection; Faster Signal Correlation

Traditional flow: Alert fires → human opens 3 tools → human manually correlates signals.

Agent flow: Alert fires (or human describes the symptom) → agent queries all relevant data sources in parallel → agent presents correlated findings in a single view.

When you tell Skyflo "payment-service latency has spiked," an illustrative diagnostic sequence might include:

  • kubernetes.list_pods: finds the CrashLoopBackOff pod
  • kubernetes.get_events: finds OOMKilled events
  • kubernetes.top_pods: reads current resource usage
  • kubernetes.get_rollout_history: finds recent deployment

These are read-only MCP tool calls. Their results can be presented together so the operator can inspect the evidence without reconstructing the same sequence by hand.

The time savings come from two sources:

  1. No context switching. The agent doesn't need to "open Grafana" or "find the dashboard." It queries the data source directly.
  2. Parallel information gathering. The agent can query multiple tools simultaneously. A human processes information sequentially: check pods, then events, then metrics, then deployment history. The agent collects all of it in one pass.

Phase 2: Diagnosis; A hypothesis with evidence

Traditional flow: Human stares at correlated data → builds mental model → forms hypothesis → tests hypothesis manually.

Agent flow: Agent receives correlated data → proposes a causal hypothesis → presents the supporting evidence.

The agent should distinguish observed facts from its hypothesis. For example, an OOMKilled event and a configured memory limit are evidence; the claim that a particular deployment caused the change is a hypothesis until deployment history and application changes support it. The operator can confirm the direction or ask for more evidence.

Phase 3: Remediation; Safe Execution with Approval

Traditional flow: Human drafts command → peer reviews → human executes → human watches rollout.

Agent flow: Agent proposes specific fix with evidence → human reviews in context → agent executes via typed tool → agent watches rollout automatically.

The agent presents a specific, evidence-based proposal. For example:

code
kubernetes.patch_deployment(
  namespace="production",
  name="payment-service",
  patch=operator_reviewed_resource_change
)

The operator reviews the scope and approves or rejects it. Only the approved typed tool input proceeds to execution, and the result is attached to the operation record.

Phase 4: Verification; Automated Post-Fix Validation

Traditional flow: Human goes back to Grafana → waits for metrics update → manually checks pod health → updates incident channel.

Agent flow: The agent automatically checks pod health, resource utilization, error events, and application metrics, then reports evidence-based verification.

The agent doesn't ask "did it work?" It checks:

  • Are all pods running and ready?
  • Is memory utilization within the new limits?
  • Are there new OOMKill events?
  • Has p99 latency recovered?

Verification is a structured step in the Plan → Approve → Execute → Verify workflow rather than an optional note after execution.


An illustrative agent-assisted incident

The product path is deliberately simple: describe the symptom, inspect the gathered evidence, review the proposed mutation, approve or reject it, and read the verification result. This sequence illustrates the control model. It is not a production benchmark, and it makes no claim about how long a particular incident will take.


What changes in the workflow

The useful change is structural rather than a promise of a particular resolution time:

Collected evidence. Tool results are brought into the same operation record so the operator can inspect what the proposal is based on.

Evidence-based proposals. A proposal names the target, the requested change, and the observations that led to it. The operator can reject it or request more evidence.

Explicit verification. After execution, the workflow reads current state again and compares it with the approved intent.

Continuous operation context. If verification fails, the evidence from the attempted change remains available for the next plan or an escalation.


Measuring the impact

Measure the system against your own incident history. Useful comparisons include time spent gathering evidence, time waiting for an operator decision, failed or rejected mutation attempts, verification failures, and overall MTTR by incident class. A scripted walkthrough cannot establish these outcomes.

Also retain the operation record: what was requested, which evidence was read, what was proposed, who approved it, which typed tool input executed, and what verification observed. That record makes later review possible without treating the model's private reasoning as evidence.

For teams evaluating this approach, the use cases page shows specific operational scenarios. To see the agent architecture that enables this, read How Skyflo Works Under the Hood.


Try Skyflo

Review the approval-gated workflow in your own environment. Skyflo is open-source and self-hosted; model requests go to the provider you configure.

bash
helm repo add skyflo https://charts.skyflo.ai
helm install skyflo skyflo/skyflo