Automated root cause analysis in a single run
An AI agent that generates its own investigation flows, runs them via a controlled MCP server, and delivers root cause plus a fix proposal, the same result every time.
Manual troubleshooting across multiple systems is slow and inconsistent. The same incident gets investigated differently depending on who's looking, and correlation between cluster, cloud and GitOps happens in someone's head under time pressure.
- The investigation takes hours when the actual answer is three API calls away.
- Two people on the same incident reach two different conclusions.
- No documented chain from symptom to root cause, the analysis can't be repeated.
- Knowledge stays with whoever happened to solve it last time.
From incident to root cause.
One flow, one run. Nobody hops between five tools.
Receive the incident
The agent (Claude Code) gets the incident description and decides which systems need investigating.
Generate skills
The agent writes its own code, investigation flows as skills, tailored to this specific incident.
Run isolated
The skill runs read-only in a Podman container via the FastMCP server, with secrets injected outside the agent's context.
Call the systems
Controlled API calls against AWS, EKS, ArgoCD and GitHub, never direct access from the agent.
Structured response
The result comes back structured: root cause, evidence and proposed fixes in a single run.
The agent never reaches your infrastructure directly. All execution happens read-only via the MCP server.
Before and after.
- Hours of manual troubleshooting across three consoles.
- Two people, two conclusions, no documentation.
- Drift between dev and prod is found by chance.
- New system = new integration project.
- The whole investigation runs in one pass and gets summarized.
- Root cause with evidence, traceable and repeatable.
- Environment differences are compared automatically on every run.
- New system = a new skill, wired up in days.
What it delivers in production.
- A shorter path from incident to root cause.
- The same analysis every time, regardless of who's on call.
- No manual correlation between tools and domains.
- New systems get connected as modular skills, not as new integration projects.
The agent doesn't guess
It executes, every conclusion is based on data actually pulled from the systems.
No direct access
All access goes through the MCP server with clearly scoped permissions.
Isolated secrets
Secrets are injected into the container, never into the agent's context.
Modularly extensible
New target systems, cloud providers, third-party APIs, internal tools, get added as skills.
Tools we trust.
The difference isn't that the agent is smarter than us. It's that it does exactly the same thing every time, and never skips step three because it's boring.

The pattern started as a demo case in a DevSecOps environment with EKS, ArgoCD and GitOps manifests, but is directly applicable in production. In hybrid environments like Viaplay's, where an incident almost always spans cluster, cloud and GitOps, the agent does the whole pass in one sweep and hands back root cause, evidence and next steps.
Read the full caseBook a demo of automated root cause analysis in a single run
30 minutes, digital. We show the solution in practice and what it would do in your environment, no sales pitch.
Break the VMware lock-in
Migrate from VMware to Proxmox, Nutanix or a hyperscaler, with no service disruption.
Automate operations
Replace ticket-driven ops with infrastructure as code, GitOps and event-driven runbooks.
Build self-healing IT
Close the loop: detect, decide, act. Observability that drives automation, not dashboards no one reads.




