Insights & updates from our experts
The Difference Between AI You Can Trust and AI You're Just Hoping Works
Few organizations hand a new engineer root access on day one and walk away. They get scoped permissions, a mentor reviewing their first changes, and a track record they build before anyone lets them touch production unsupervised. Trust with critical systems works this way for a reason: it’s earned in stages, not assumed on arrival.
A lot of organizations skip most of that process for AI. They deploy an agent into an ops workflow, give it broad access because it’s “just automation,” and leave it to act without the staged trust-building a human would have to go through. Gartner’s May 2026 research on AI agent governance points to this pattern, warning that enterprises tend to treat AI agent governance as binary, locked down or fully trusted, rather than calibrated to what a given agent is allowed to do, and predicting that roughly 40% of enterprises will demote or decommission autonomous agents by 2027 after governance gaps surface in production, not before. AI can build trust with production over time. Few organizations build the path that gets it there.
Your rollback script hasn’t had a bad day. Your AI agent might.
A traditional runbook has one path through it: the same input produces the same output, every time. If step 4 fails, you know what state the system is in, because there was only ever one way to get there. That predictability is why deterministic automation won the right to run unsupervised: its failure modes are finite and known in advance.
AI doesn’t offer that guarantee yet, and this isn’t hypothetical. StackGen’s 2026 State of Reliability Report, the largest published analysis of company status-page records to date, found AI-related incidents have increased roughly sixfold over the past three years and now account for more than one in ten disclosed operational incidents. Among the specific failures: the report documents at least nine cases since mid-2025 of AI agents wiping data, deleting databases, or destroying live systems entirely on their own. For a marketing email draft, an agent’s variability is harmless. For a production database, it’s the gap between a clean fix and an outage on a status page. That gap is why AI can’t run alone the way your deterministic scripts do, at least not yet, and not without the same track record those tools spent twenty years building.
The access controls you already have are the ones AI needs. It’s time to extend them.
None of this calls for new infrastructure. The role-based permissions, approval workflows, and audit logs that already manage a new engineer's access sit in most environments today, unused for this specific job. Extending them to AI means defining what limited access looks like for an agent, not building a governance system from scratch.
An AI agent should get the same minimum-necessary access a new engineer would get for the same task, not broader access because it’s automated. A Cloud Security Alliance survey of IT and security professionals, fielded with Token Security in early 2026, found that 65% of organizations had experienced at least one AI agent-related security incident in the past year, and 41% of those incidents involved an agent taking an unintended action that its own permissions technically allowed. The agents misbehaved, but only because their permissions let them. The team hadn't scoped those permissions tightly enough in the first place.
Log every action an agent takes with the same rigor as a human-initiated change, and tag it as autonomous, so the audit trail holds up when something goes wrong. The organizations getting this wrong tend to make the same mistake: one uniform governance policy across every agent, regardless of what it’s allowed to touch or how much autonomy it has. Uniform rules feel safer than they are. They either lock every agent down to the point of uselessness, or leave your most sensitive agents under-governed because you wrote the policy for your least sensitive one.
Human-in-the-loop works like a ladder, not a switch. The agent climbs it one rung at a time.
Most teams treat human-in-the-loop as binary: either a person approves every action, or the agent runs unsupervised. That framing guarantees one of two bad outcomes. Either your engineers get bottlenecked reviewing everything indefinitely, or someone gets impatient and removes the review step entirely.
There’s a middle path, and it looks a lot like how you’d manage a new hire. Route every AI-proposed action, a rollback, a permission change, to a person who can accept, modify, or reject it before it touches production. Then track the outcome by task category, not just overall. If restart-and-scale proposals get accepted without modification nearly every time, that’s evidence the agent deserves less oversight on that specific task. If reviewers keep modifying or rejecting permission changes, the agent stays on a short leash there, regardless of how well it’s doing elsewhere.
Where “alone” costs you
You feel this the first time an AI agent proposes a fix for a stuck deployment at 2 a.m., and you have to decide, in the moment, whether to approve it without a second opinion. You’ll see it again in the postmortem that asks why a change happened, when the real answer is “an agent decided to,” with no record of what access it used or who was watching, and in the friction between the team that wants AI moving fast and the team that has to answer for what happens when it moves fast, alone, and wrong.
Every one of those moments gets easier when the agent wasn’t alone in the first place, when you scoped the access, logged the action, and positioned a person to catch it before it shipped.
This staged-trust model shows up in production tooling today. Xurrent assigns Sera AI Agents to teams with defined roles rather than blanket permissions, and each skill an agent runs carries its own confidence threshold: cross it, and the agent escalates to a person instead of acting. A Triage Agent working a ticket might run duplicate detection or impact assessment on its own while still kicking anything outside its threshold to a human. Xurrent logs every action, autonomous or not, under the same policy layer that governs everything else on the platform: autonomy calibrated per task, built into the software rather than granted all at once.
The goal is simple: AI doesn’t ship to production alone.
Deterministic automation earned unsupervised access over twenty years of predictable behavior. AI hasn’t put in that time. Give it the same scoped access and audit trail you’d give a new engineer. Let a person review its proposals long enough to know where it’s won more room. Speed follows once you’ve built the system underneath it.
Curious how staged autonomy works in practice? Explore Sera AI.
This idea came up on Incidentally Reliable, Episode 2 of Season 3: “Never Let AI Touch Prod Alone.” The full conversation also gets into how runbooks quietly go stale the moment they’re written, and why our guest, Zahan Parekh, still keeps a human in the loop on every one-click fix his product proposes. Listen to the full episode.
Frequently Asked Questions
.webp)
Stop switching tabs to find context. Xurrent MCP is live.
Most enterprise AI is a chatbot stuck to the side of your tools. We built something different. Two MCP servers, one for incident response and one for service operations, that let you ask your operational data anything, and in the case of ITSM, act on it, using whatever AI client your team already uses.

An AI SRE that knows your incidents
Most AI SREs are pattern matchers trained on public data. They know what a memory leak looks like in the abstract. They don't know that your payments-api has a flaky liveness probe everyone ignores, that the checkout team owns the retry policy, or that the last three "database incidents" were actually cache misconfigurations. That knowledge lives in your postmortems, your Slack channels, and the heads of two senior engineers.

How Long Should ITSM Implementation Really Take in 2026?
Most vendors will tell you ITSM implementation takes six months to a year — but modern, configuration-first platforms have rewritten the math entirely. See what real implementations look like in 2026, and why a long rollout is now a choice, not a given.















.webp)
.webp)
















