← Back
Azure Troubleshooting Agent

Defining How AI Agents Operate at Cloud Scale

An agent that troubleshoots Kubernetes, shows its evidence, and leaves the fix to the engineer.

RoleProduct Designer
Timeline2025 – 2026
ScopeAzure Copilot, enterprise
PatternHuman-in-the-loop

Problem

Diagnosing a cluster failure meant hopping across logs, dashboards, and commands.

Approach

An agent that investigates on its own, shows every step, and waits to be told to act.

Impact

In testing, everyone reached a safe next action — with the engineer still in control.

100%reached the right next action
100%recognized what needed attention
0steps run without confirmation

Speed was easy. Trust was the design problem.

The Problem

Troubleshooting was scattered across five tools.

Logs in one place, dashboards in another, and the fix in none of them.

Diagram — the manual path beside the agent-assisted path

Six manual steps, or one reviewed action. The same failure, two very different routes.

The Bet

The agent investigates. The engineer decides.

Evidence-gathering runs on its own; a consequential change never does.

"An agent you can't read is an agent you can't trust."

Automatic.

The agent collects evidence and keeps its progress visible.

Separated.

Findings stay distinct from the actions it proposes.

Confirmed.

A risky change waits for the engineer's review.

Discoverability

Meet the operator where the problem already is.

Not a generic button — one tied to the resource on screen.

Image — a contextual investigate entry on a failing resource

Everyone read these as things to act on. Recommendations, surfaced in place.

From Reasoning to a Narrative

A correct answer isn't automatically a usable one.

Every response resolves to diagnosis, solution, and a safe next step.

Diagnosis.

What is wrong, and which resource it affects.

Solution.

The change that would resolve it.

Next steps.

How to proceed safely from here.

Making the Evidence Visible

Evidence builds trust only if you can find it.

Every command, resource, and signal — surfaced, not buried.

Image — the evidence panel: commands, events, logs, and the pod manifest

Everyone needed help finding the manifest. Hidden evidence can't earn trust.

Reducing Complexity

Pick the visual the decision needs.

A snapshot, a map, a timeline — matched to the question.

Diagram — a DNS outage traced: service healthy, resolution failing, CoreDNS degraded

One failing path, end to end. Where it breaks, before the text does.

100%reached a safe next action — in prototype testing

One agent, one scenario. Then a pattern for many.

What I Learned

Trust is designed, not declared.

Show what the agent inspected and why — then leave the decision to the human.

Next Project

Azure OpenAI for Developers

Helping developers get started with AI by doing, not just reading.

Azure OpenAI for Developers — lead image