AIOps & troubleshooting

Automated root cause analysis in a single run

An AI agent that generates its own investigation flows, runs them via a controlled MCP server, and delivers root cause plus a fix proposal, the same result every time.

Single runSelf-generating skillsRead-only executionRepeatableExtensible
What you get
The whole investigation in one run, incident in, root cause out.
Read-only execution, isolated secrets, full logging.
New systems connected as skills, not projects.
1
run from incident to answer
6
systems investigated automatically
100%
read-only, nothing touched in production
Always
the same analysis, regardless of who runs it
The problem

Manual troubleshooting across multiple systems is slow and inconsistent. The same incident gets investigated differently depending on who's looking, and correlation between cluster, cloud and GitOps happens in someone's head under time pressure.

  • The investigation takes hours when the actual answer is three API calls away.
  • Two people on the same incident reach two different conclusions.
  • No documented chain from symptom to root cause, the analysis can't be repeated.
  • Knowledge stays with whoever happened to solve it last time.
How it works

From incident to root cause.

One flow, one run. Nobody hops between five tools.

01

Receive the incident

The agent (Claude Code) gets the incident description and decides which systems need investigating.

02

Generate skills

The agent writes its own code, investigation flows as skills, tailored to this specific incident.

03

Run isolated

The skill runs read-only in a Podman container via the FastMCP server, with secrets injected outside the agent's context.

04

Call the systems

Controlled API calls against AWS, EKS, ArgoCD and GitHub, never direct access from the agent.

05

Structured response

The result comes back structured: root cause, evidence and proposed fixes in a single run.

This gets connected
Kubernetes (EKS)GitOps (ArgoCD)Source code & manifestsContainers & imagesNetwork & accessEnvironment differences

The agent never reaches your infrastructure directly. All execution happens read-only via the MCP server.

The difference

Before and after.

How it looks today
  • Hours of manual troubleshooting across three consoles.
  • Two people, two conclusions, no documentation.
  • Drift between dev and prod is found by chance.
  • New system = new integration project.
How it looks with the solution
  • The whole investigation runs in one pass and gets summarized.
  • Root cause with evidence, traceable and repeatable.
  • Environment differences are compared automatically on every run.
  • New system = a new skill, wired up in days.
The impact

What it delivers in production.

  • A shorter path from incident to root cause.
  • The same analysis every time, regardless of who's on call.
  • No manual correlation between tools and domains.
  • New systems get connected as modular skills, not as new integration projects.
Example findings
Configuration mismatches between environmentsIncorrect image tags and deployment failuresArgoCD out-of-sync and configuration driftNetwork and access issuesSecurity groups blocking trafficOverlay errors between dev and prod
How it stays secure

The agent doesn't guess

It executes, every conclusion is based on data actually pulled from the systems.

No direct access

All access goes through the MCP server with clearly scoped permissions.

Isolated secrets

Secrets are injected into the container, never into the agent's context.

Modularly extensible

New target systems, cloud providers, third-party APIs, internal tools, get added as skills.

The stack

Tools we trust.

Claude CodeMCPFastMCPPodmanPython 3.11AWS EKSArgoCDGitHubECRKustomize
The difference isn't that the agent is smarter than us. It's that it does exactly the same thing every time, and never skips step three because it's boring.
Emilija Trpchevska
Emilija Trpchevska
DevOps Consultant · Stockholm
Customer case
1 run
the whole analysis
Viaplay

The pattern started as a demo case in a DevSecOps environment with EKS, ArgoCD and GitOps manifests, but is directly applicable in production. In hybrid environments like Viaplay's, where an incident almost always spans cluster, cloud and GitOps, the agent does the whole pass in one sweep and hands back root cause, evidence and next steps.

Read the full case
See it live

Book a demo of automated root cause analysis in a single run

30 minutes, digital. We show the solution in practice and what it would do in your environment, no sales pitch.