AI Agents Will Change MSPs. Tame Them First.

Written by Cyril Mathew | Oct 5, 2026

Over the last year, I have spent a fair amount of time thinking about where AI agents actually fit into managed services. Not in theory. In the day-to-day reality of running an MSP: supporting clients, dealing with incidents, reviewing costs, managing changes, and trying to improve how engineers spend their time.

The more I examine it, the more I believe agents will change managed services significantly. But I do not think they will replace MSPs, and I do not think the immediate future is a fully autonomous operation where agents are making production decisions on their own. For now, the better model is simpler: use agents to remove repetitive work, improve context, and help engineers make better decisions faster. That still keeps the human in the loop. And in many situations, that is exactly where the human should stay.

Because one thing becomes obvious very quickly: just because an agent can do something does not mean it should be allowed to do it.

The interesting part is often before the fix

Anyone who has worked in managed services knows this. The actual technical fix is often the easy part.

An alert comes in. Maybe a service is unhealthy. Maybe CPU has spiked. Maybe a job has failed. The engineer does not immediately take action; first, they need to understand what they are looking at. Which client is this? Which AWS account? Production or non-production? Is it a Tier 1 workload? What changed recently? Has this happened before? Is there a runbook? Is there an open change? What is the likely impact? Does the client need to approve the action?

Sometimes the actual remediation takes five minutes. The investigation leading up to it can take much longer.

Our own numbers make this plain. Across our AWS estate between June and September this year, alerting raised 18,315 incidents, 14,231 of them high urgency. The median incident was acknowledged in 22 seconds and resolved in under 13 minutes. The slowest tenth took close to three days.

That spread is the whole argument. The long tail is not slow because the fix is hard. It is slow because of everything that happens before the fix: which client, which account, whether it is production, what changed, whether there is a runbook, who needs to approve. That is the work I want an agent to compress.

Instead of an engineer spending the first 20 to 30 minutes collecting context from monitoring, tickets, documentation, previous incidents, and cloud consoles, the agent can bring that together. So the flow shifts from:

Alert → Engineer investigates → Engineer decides → Engineer acts

to something more like:

Alert → Agent gathers context → Agent investigates → Agent recommends → Engineer decides

That is already useful. The engineer still owns the decision, but they start much further ahead. For an MSP, that can have a real impact on response times and MTTR without taking unnecessary risk.

FinOps has been a good example

FinOps is another area where this became clear. Finding that a workload looks underutilized is not difficult. Deciding what to do about it is the hard part.

We might see an EC2 instance running at low utilization for weeks and think it looks oversized. But then the questions start. Is the client planning a release? Is the workload seasonal? Did the application owners intentionally leave headroom? Does this workload have a DR dependency? Is there an architecture change coming in a few months? The same complexity applies to Savings Plans and other commitments.

An agent can analyze the data, find opportunities, compare patterns, and explain why something looks inefficient. That can save a significant amount of time. But it does not always know the full business context, and that context matters. So I do not look at the FinOps agent as something that replaces the FinOps practitioner. I look at it as something that lets the practitioner start from a much better place: here is what we found, here is why it matters, here is the likely recommendation. Then the human applies judgment. That is a much better use of the person's time.

Agentic and autonomous are not the same thing

This distinction matters. An agent can be highly capable without being autonomous. It can read telemetry, query tools, review previous incidents, search runbooks, interact with ITSM, prepare remediation steps, and execute an action after approval. That is still a meaningful agentic workflow. The human approval step does not make it weak; in many cases, it is the control that makes the workflow usable.

For a production change, a DR failover, a security remediation, or a significant financial decision, a human approval point is exactly right. Not because the agent is insufficiently capable, but because somebody still needs to own the risk.

A realistic managed-services flow looks more like this:

Observe → Investigate → Recommend → Human approves → Execute → Validate → Document

Over time, for very specific and low-risk scenarios, that can evolve. A known non-production remediation may eventually move to policy-based execution. But that is not where to start.

Autonomy should be earned.

The bigger challenge is that every client is different

This is where the MSP model makes things more complicated. We are not running one environment. Every client has a different architecture, maturity level, internal IT staff, risk tolerance, compliance requirement, and change process. The same action can be perfectly acceptable for one client and completely unacceptable for another. One client may be comfortable allowing an agent to restart a non-production workload when a known set of conditions are met. Another may want human approval every time. A regulated client may have no choice.

Even within one client, different workloads need different rules. An agent may freely analyze spend, investigate a security finding, and recommend a remediation. But no one is going to allow it to trigger a production DR failover simply because the model concludes it is a good idea, not without a very clear policy and a significant amount of established trust.

So the right question is not: how autonomous is our MSP? A much better question is: for this client, this workload, and this action, how much authority should the agent have? That is much closer to how this will actually work.

What "taming" an agent actually means

When I use the word tame, I am not talking about making the agent less intelligent. I am talking about putting clear boundaries around it.

An MSP agent should know which client it is working for, which accounts are in scope, what is production, and what the SLA, RTO, RPO, approved runbooks, change process, escalation path, and known exceptions are. It should also know what it is not allowed to do. And the standard production principles still apply: least privilege, approval gates, audit trails, observability, rollback, and a clear blast radius.

That is where generic AI assistants and real MSP agents begin to separate. A model may understand AWS thoroughly. That is useful. But an MSP agent needs to understand the actual client environment, not AWS in general. This client. This workload. This policy. This situation. That context is where the real value sits.

Runbooks will need to evolve

I expect runbooks to change as agents become part of operations. Today, a runbook tells an engineer what to do. In the near future, it may also define what an agent is allowed to do: what conditions need to be true, what telemetry should be checked, which environment the action is valid for, whether it needs approval, how to validate the result, what the rollback looks like, and when the agent should stop and escalate.

That turns the runbook from simple documentation into part of the agent's operating policy. For MSPs, that feels natural. The controls already exist. The difference is that they now need to be explicit enough for an agent to follow.

One agent doing everything is the wrong model

The more use cases we examine, the less appealing a single agent doing everything becomes. Specialized agents make more sense: one for incidents, one for FinOps, one for security, one for change management, one for knowledge and operational history, one for service management. They can work together, with a supervisor agent eventually coordinating them. But even then, the important question is not how clever the orchestration is. It is whether the boundaries are clear. Multi-agent automation without clear scope can simply produce larger mistakes faster.

The MSP role changes, but it does not disappear

Engineers will spend less time on repetitive work: log gathering, dashboard hopping, manual correlation, repetitive triage, ticket updates, routine reporting. And more time on the work where experience actually matters: architecture, exception handling, automation, risk assessment, client conversations, complex incidents, governance.

The role changes, but I do not see it disappearing. If anything, the quality of the human layer becomes more important because the agent can now move much faster.

How we are approaching this at EPI-USE

At EPI-USE, we are not starting with the goal of building a fully autonomous MSP. We are starting with a more practical question: where are engineers repeatedly spending time that an agent can genuinely help with? Where can we shorten incident investigation? Where can we improve FinOps analysis? Where can we improve security triage? Where can we make service management more proactive?

The first thing we built was not a remediation agent. It was a client context agent.

Before an engineer picks up an incident or walks into a client call, the agent assembles the picture from the systems we already run: the service desk, PagerDuty, Confluence, Outlook, and the client's own commercial and scope record. It answers the same questions I listed earlier, all of it before anyone opens a console. Which client, which account, what is in scope, whether this has happened before, what is currently open.

It is also read-only by design. It cannot write to any of the systems it reads from. Writing back to the service desk, the change record, and the alerting platform is on the roadmap, and each of those will sit behind an explicit approval step when it arrives. Capability first, authority later.

For many workflows today, the agent can investigate, correlate, and recommend. The engineer or service owner still makes the decision. As a use case becomes more predictable, the controls improve, and the client becomes comfortable, selected actions can move toward policy-based execution. But we do not expect every client to move at the same speed. And we do not expect every action to become autonomous. That is not the goal.

The goal is to build client-aware agents that work inside clear guardrails, support engineers, and earn trust over time.

That is how I see the next phase of managed services. Not human versus agent. Not MSP versus AI. Human, agent, client context, and governance working together. The transition will happen gradually, by client, by workload, by action.

AI agents are going to change MSPs. But before handing them the keys to production, we need to understand where they belong, put the right boundaries around them, and let them earn the trust.

In other words: tame them first.

If you want to see how this model is taking shape in a production managed services environment, start with a 30-minute conversation with a cloud expert.

Cyril Mathew is Head of Managed Services and Senior Director at EPI-USE Services for AWS, an AWS Premier Partner and AWS Validated Managed Service Provider. He holds nine AWS certifications and specialises in cloud strategy, FinOps, agentic AI, and enterprise cloud security.