Blog

Would you let an agent touch production today?

October 8, 2026
Rohan T
Gray upward-pointing arrow icon.
Click To Explore

Table of contents

Downward-pointing chevron dropdown arrow icon in black.

I asked SREs that question all day at SREDay Bangalore. Not one of them said yes. Their reasons are the most rigorous thinking happening in this industry right now.

On 20 June we hosted SREDay at our HSR Layout office in Bangalore. Between sessions I sat down with SREs and asked each of them a version of the same question.

Would you let an AI agent make changes in your production systems today?

Different companies, different scales, from a two-person founding team to Pearson and AWS. Nobody said yes.

If you have been reading vendor decks for the last eighteen months, that sounds like a room full of laggards. It isn't. Every person who said no had already deployed AI somewhere in their stack, and every one of them had a specific, technical reason for where they had drawn the line. The reasons did not match each other.

The room and the roadmap do not agree

Gartner's Predicts 2026 for infrastructure and operations includes a strategic planning assumption that human-in-the-loop will fall from 95% of IT operations workflows in 2025 to 40% by 2028. That is a very short runway. On that timeline, more than half the workflows that have a human in them today lose that human within three years.

Meanwhile The SRE Report 2026, Catchpoint's eighth annual study of more than 400 reliability practitioners, found that AI optimism jumped to 60% while skepticism dropped to 21%. More than half of respondents plan to run agentic systems in production within twelve months.

So the industry data says the community is bullish and moving. The room in Bangalore said not yet. Both of those are true at the same time, and holding both is not a contradiction. It is calibration. People are enthusiastic about the direction and precise about the sequence.

The sequence has three rungs, and every conversation I had that day was standing on one of them.

Rung one: agents show up with no idea where they are

Saurabh Hirani is a Principal SRE at One2N, and he gave me the sharpest line of the whole day.

All these agents that do debugging for you, incident triage, all of that, they are lame when they go into the organization. And when I say lame I mean they don't have the crutch or the support of the data they should be acting on. Organizational practices around monitoring and observability are still rooted in the ancient ways of doing things, and now you're just slapping an AI agent on top of it and saying, now fix this issue.
SHSaurabh Hirani Saurabh HiraniPrincipal SRE, One2N

His advice to customers is blunt: before you bring in the army of agents, do some house cleaning.

Google Cloud's DORA team reached the same conclusion from a completely different direction. Their 2025 State of AI-assisted Software Development covered nearly 5,000 technology professionals and landed on AI as an amplifier. It magnifies whatever an organization already has. Strong platform and clear workflows get stronger. Ambiguous ownership and stale runbooks get worse, faster. Nathen Harvey, who leads DORA at Google Cloud, put it as AI creating localized pockets of productivity that are then lost in downstream chaos.

Ananda Rajagopal, co-founder and Chief Product Officer at Ciroos, described the same failure in organizational terms:

Most delays in incident resolution aren't technical; they're organizational. They happen in the labyrinth of escalation paths, Slack threads, and tribal knowledge that lives in someone's head rather than somewhere accessible.
ARAnanda Rajagopal Ananda RajagopalCo-founder and CPO, Ciroos

That is Saurabh's observation, one layer up. An agent cannot read a Slack thread that was never written, and it cannot ask the one engineer who remembers why that service has a weird retry policy.

Nandini Bhatt, an SRE II on our own team, described what fixing it looks like from inside:

"If we actually put time in and train our agents, give more context, build more. Right now we're using Claude. Build more Claude skills, give more context of the infra, of the observability, of our patterns. Then yes, it could work out great for us."

Charity Majors has been making a version of this argument all year. Her framing is that when code generation becomes cheap, the value moves to the durable artifacts of understanding: tests, evals, observability, specs, and how the system actually behaves in production. She calls 2026 a return to rigor. Her line that stuck with me is that production is a stage of development, not something that happens after development ends.

Four people, four vantage points, one conclusion. The bottleneck is not the model. It is everything the model is being asked to reason about.

Rung two: the line is reversibility

This is where the conversation got genuinely useful, and it came from Shobhit Gupta, co-founder at Segwise. His company runs plenty of AI agents in the product. None of them can alter production, and he was precise about why:

"I'll be comfortable with AI doing the changes as long as I'm comfortable with its reasoning. This is the first phase we're in at our company, where AI tells me why it thinks something happened and why it happened. If you're happy with that, solving it is not the hard part. Figuring out why something broke is the hard part."

Then he drew the map:

Obviously we'll start in certain areas. Things that are reversible. Creating a new service, increasing service capacity, changing some networking things. Those are easily reversible decisions if AI makes a mistake. Databases and the data part is the hard part. If it loses data, that's less reversible, so we'll be more cautious there.
SGShobhit Gupta Shobhit GuptaCo-founder, Segwise

Reversibility is a much better permission boundary than criticality, and the reason is mechanical. Criticality is a judgment about a system, it changes with business context, and two engineers on the same team will rank the same service differently. Reversibility is a property of the action itself. You can write it into policy, you can encode it, and you can audit it afterwards.

Gaurav, founding engineer and India site lead at oodle.ai, was the furthest along of anyone I spoke to, and he arrived at a version of the same rule:

I would probably sit somewhere in the middle, an 80/20 rule, where things that can be automated safely I would be willing to automate. You continuously monitor errors in production, raise PRs to fix them. Those are very safe operations which earlier required someone to put in manual work.
GMGaurav Maheshwari Gaurav MaheshwariFounding Engineer and India Site Lead, oodle.ai

He also gave the clearest statement of what this is actually for. SREs are supposed to split their time evenly between automation and production management. In practice the automation half kept getting eaten, because production management expands to fill whatever bandwidth exists. Agents give that half back.

Nobody in Bangalore was arguing about whether to use AI. They were arguing about which actions are cheap to undo.

Rung three: nobody can grade the agent's work

Faizana Samreen has eighteen years in the field, most of it in storage, cloud infrastructure and now observability at Pearson. She has been running the agents that are on the market, and her objection was not about capability:

What they promise on the technical features and functionality, that the autonomous SRE agents are going to replace human SREs, I don't see that happening in the near future. Right now there are a lot of players, but we're not sure how efficiently we can measure those agents unless you have use cases, and even the use cases are very few.
FSFaizana Samreen Faizana SamreenSite Reliability Engineer, Pearson

She is right, and the numbers are worse than most vendors would like.

IBM Research built ITBench to give the industry an actual scoreboard for agentic IT work. In the original paper, agents built on state-of-the-art models resolved 13.8% of 42 real SRE scenarios. In May 2026, IBM Research and Artificial Analysis published ITBench-AA covering 59 SRE tasks. Every frontier model evaluated scored below 50%.

Set that next to Gartner's own numbers. Only 17% of organizations have deployed AI agents so far, over 60% expect to within two years, and Gartner predicts more than 40% of agentic AI projects will be cancelled by the end of 2027 on cost, unclear value and inadequate risk controls. Gartner also named agent washing directly, estimating that of the thousands of vendors claiming agentic capability, roughly 130 offer anything genuinely agentic.

Dominik Zachar, founder and engineer at amigai.co, gave me the version of this I keep coming back to:

In the age where AI can speed up or fully automate response, wrong response will be bigger issue than late response.
DZDominik Zachar Dominik ZacharFounder and Engineer, amigai.co

An unmeasured agent with production access is a way to be wrong at machine speed.

The Catchpoint report has one more finding that belongs here. Median toil sits at 34% of an engineer's time, roughly flat year on year. Asked whether AI had reduced toil, 60% of directors said yes and 38% of individual contributors said yes. Twenty-two points of daylight between the people buying the tools and the people carrying the pager.

Where the value actually is right now

Here is what I did not hear once all day: autonomous remediation working in production.

Here is what I heard repeatedly. Every concrete, working, in-production example sat next to the incident rather than inside it.

Saurabh's team stopped hand-writing Terraform and YAML, which he described as the earliest and largest time saving in SRE. They also fed transcripts of customer case study recordings to a model so new joiners can talk to an agent and learn a customer's infrastructure, which cut onboarding time meaningfully in a consulting business where every customer is different.

Nandini named the two things her team automated first:

As an SRE, I hate writing postmortems as much as I enjoy debugging and going through the thrill of fixing the incident. Nobody likes to write an RCA.
NBNandini Bhatt Nandini BhattSRE II, Xurrent

Alert correlation was the other, for the same reason nobody enjoys forty pages for one underlying fault. She also gave the business case in one sentence, which is that she would rather take on planning and development work than give 40% of her day to debugging.

Gaurav's team auto-raises PRs for errors detected in production. Shobhit's builds detect a problem, reason about why it happened, and send an alert, without touching the fix.

Postmortems, correlation, infrastructure as code, onboarding, auto-raised PRs. All of it is context work. All of it makes the next incident cheaper. None of it requires anyone to hand over write access, which is exactly why it shipped first.

The legacy systems that pay for all of this

I asked Saurabh about enterprises running old tools and manual workflows, expecting the usual answer about modernization. I did not get it:

"The thing people often miss when they talk about legacy companies is that they say, oh, their coding practices are old, they don't have the right CI/CD. But they exist for a reason. They are making money. They're making more money than the shiniest thing in the market. When you respect that, you know which parts to touch and which to leave."

His rule: fix the ecosystem around the legacy system without touching the legacy system. Rebuild the CI/CD, replace the hand-rolled deployment script, leave the brittle core alone. He described using AI heavily to convert old Java 7 service discovery code for a customer, work he could not have done himself because he does not know old Java.

"Let the core be the core."

Faizana's answer to the same problem was to run both paths at once:

"We need to have the manual process in parallel. We just cannot get stuck with the legacy way of working. We will have to start using AI in parallel. We cannot let it go."

She still wakes up at 2am on Thursday nights to restart services and clear caches on a legacy application. When I asked for her weirdest root cause story, that is what she gave me. Not a story. A recurring calendar entry.

Gaurav thinks the gap between the 2am-restart world and the shift-left world closes on its own, because large enterprises run into the same bandwidth ceiling everyone else does and eventually have to automate their way out.

What the room was actually saying

Ran Tao, a cloud support engineer at Amazon Web Services, had the sharpest read on where the risk sits over the next three years:

The single biggest challenge over the next three years is change itself. Most outages already trace back to the updates we make rather than to anything external, and we are about to accelerate change dramatically by re-platforming legacy estates onto AI-driven workloads.
RTRan Tao Ran TaoCloud Support Engineer, Amazon Web Services

His prescription is staged rollouts with fast rollback, reliability budgets tied to real business tolerances, and a non-AI fallback for everything. Which is Shobhit's reversibility rule, written as a migration strategy.

So: would you let an agent touch production today?

The answer that keeps showing up, once you stop treating it as a yes-or-no question, is only if it's reversible, and only if it can show me its reasoning. That is not caution. It is the same discipline that produced error budgets and blameless postmortems, pointed at a new class of system. The industry keeps framing SRE hesitancy as a cultural problem to be managed. Sit through a day of these conversations and it reads as the only rigorous position on the floor.

The people who said no in Bangalore are also the people already running AI in their pipelines, their onboarding, their postmortems and their alert routing. They are not waiting for permission. They are waiting for evidence, and they have been extremely clear about what evidence would look like.

Give the agent your incident history, your service ownership and your on-call record. Start it on actions you can undo. Then measure it.

That order is not negotiable, and nobody at SREDay Bangalore was confused about it.

The SREs who shaped this piece

  • Saurabh Hirani, Principal SRE, One2N — LinkedIn
  • Faizana Samreen, Site Reliability Engineer, Pearson
  • Gaurav, Founding Engineer and India Site Lead, oodle.ai
  • Shobhit Gupta, Co-founder, Segwise
  • Nandini Bhatt, SRE II, Xurrent
  • Ran Tao, Cloud Support Engineer, Amazon Web Services
  • Ananda Rajagopal, Co-founder and Chief Product Officer, Ciroos
  • Dominik Zachar, Founder and Engineer, amigai.co

All quotes gathered at SREDay Bangalore on 20 June 2026 and used with permission.

Sources