AI Now Accounts for 1 in 10 Incidents, a 6x Rise in Three Years — and AI Agents Have Destroyed Live Company Systems on Their Own, StackGen Study Finds

Companies spent the past two years putting AI to work, and the cost is now arriving as downtime. One of every ten incidents companies disclosed this year comes from an AI provider or AI product, roughly a 6x rise in three years. In at least 9 documented cases since July 2025, the AI itself did the damage: an agent wiped data, deleted databases and destroyed live systems at the businesses and teams that deployed it, acting entirely on its own. The ways to fail are multiplying; the speed of recovery, the study finds, is not. Gartner® mentions, “By 2029, 90% of organizations will experience an AI-caused outage, yet continue AI SRE for its speed and scalability gains.”1

The findings come from the State of Reliability Report, released today by StackGen, a pioneer in autonomous operations management for cloud services. The study analyzed 177,960 public status-page records covering approximately 109,100 unplanned incidents, making it the largest published analysis of how online services actually fail and recover, based on companies’ own status pages. The report includes data from more than 390 companies across 13 sectors, including cloud infrastructure, payments, e-commerce, communications, security, and AI providers, from 2018 through June 2026.

The report finds AI gives a business three new ways to go down.

  • When an AI provider has an incident, every product built on it can fail too, and customers watch checkouts, apps, and logins stop working.

  • Sometimes nothing visibly breaks: the service looks fine while the AI quietly serves customers wrong answers. It is the fastest-growing failure category in the study, growing from 1 customer-facing AI quality incident in 2025 to 89 in 2026 YTD among AI-native companies.

  • In the newest case, an AI agent inside the company takes destructive action on its own. Because the agent acts with valid credentials, nothing looks wrong to any monitoring tool while it happens. The first sign is the damage itself: missing data, a system that no longer exists.

Together, the findings describe a widening gap. Incidents are rising and arriving in new forms, while median resolution times have not improved, and the most common fixes are the same ones teams have leaned on for years: waiting, restarting, rolling back. Carried forward, those trend lines leave SRE and operations teams facing more incidents than they have people to handle. The report frames it as a race: AI is adding to SRE workload faster than most teams are applying AI to take work away. The companies operating AI at the largest scale now recover faster than any other sector in the study.

Key findings:

  • AI incidents are roughly six times more common than three years ago and now top 1 in 10. Incidents disclosed by AI model and AI application companies were 1.7% of all disclosed incidents in 2023; in 2026 YTD they are 10.7%. Because the figure counts only incidents at AI companies themselves (AI-caused failures at other businesses are excluded), it is a floor, and the true share is higher.

  • AI agents have destroyed live company systems on their own at least nine times. The nine confirmed cases work out to nearly one per month since July 2025; each one is publicly documented by the operator’s own post-mortem or an independent incident database. In the most common pattern, the agent found a password or key it was never supposed to have and used it against the company’s systems, doing damage a human employee could be fired for, if done deliberately. The count includes only what companies have admitted, so the real number is higher; documented AI agent incidents of all kinds more than doubled from 2024 to 2025 and are on pace to rise again in 2026. Incidents caused by AI agents taking destructive action against production systems are now regular tech-press stories — and in July 2026, an AI agent compromising production infrastructure made international headlines.

  • More than 1 in 4 incidents the study assessed starts at a company the affected business does not control. These third party vendor incidents take roughly three times longer to fix (a 247-minute median versus 96 minutes for internally caused configuration failures). The largest single event in the dataset, the October 20, 2025 AWS outage, impacted 223 downstream companies. Core AWS services were degraded for up to a day, and some downstream recovery tails ran into days. For a business built on the same dependencies, that exposure runs straight to revenue and customers.

  • Which company you are matters about 3x more than which industry you are in. Two teams in the same sector can differ threefold in the speed of recovery. That gap results from how the team operates: observability, tooling, on-call practices, and how AI is applied to incident response.

  • The most common way companies resolve an incident is to wait for someone else to fix it. Waiting on an upstream provider is the single largest remediation category (13.6% of remediations in the report’s coded post-mortem dataset), ahead of restarting and rolling back. For the enterprise, that means recovery time on a meaningful share of incidents is set by another company’s engineers. The exposure can be managed in advance, through dependency choices and contracts, but it is unmovable once the outage starts.

  • Even as AI multiplies the ways systems fail, companies are not getting faster at fixing them. Median resolution times have been roughly flat within each industry tier since 2023, and the most common fixes (waiting, restarting, rolling back) are the same ones teams have applied for years. AI incidents are compounding while recovery speed stays flat, and on current trend lines the gap between incident growth and SRE capacity will keep widening.

  • Fast recovery at scale is achievable, and the proof comes from the companies operating AI at the largest scale. AI model providers recover fastest of any sector, with a 49-minute median this year, down from 75 minutes in 2023, and the report documents the practices behind that speed.

“Companies spent the past two years putting AI into production, and the incident record now shows the operational bill. More than one in ten disclosed incidents now comes from an AI provider or AI product. Reliability in the AI era is becoming a central priority for engineering teams: the nature of incidents is changing, and volume will only climb as AI-assisted coding pushes more change into production. The top teams resolve incidents three times faster than peers in their own industry. We published this benchmark so leaders can see where they actually stand,” said Sachin Aggarwal, CEO of StackGen.

The dataset consists entirely of observed behavior: public status-page records, with scheduled maintenance and advisory postings excluded from duration statistics. No surveys or self-reported figures are included. The report notes that status-page methodology undercounts agent-caused incidents by design, so the agent-related figures represent minimums.

The full State of Reliability 2026 report is available at https://stackgen.com/state-of-reliability-2026/. StackGen will present the findings and take questions in a live webinar today at 11:00 AM ET/8:00 AM PT; registration is available at https://www.linkedin.com/events/7473452197993734146/.

1 Gartner, AI, Autonomy, and Architects: The Future of Site Reliability Engineering”, Daniel Betts, Hassan Ennaciri, Chris Saunderson, 24 September 2025 GARTNER is a trademark of Gartner, Inc. and/or its affiliates.

About StackGen

StackGen is a pioneer in autonomous operations: AI that provisions, deploys, observes, and heals enterprise software. Its Autonomous Operations Platform moves enterprises from manual DevOps to AI-powered operations through four Aiden modules: Aiden for SRE (AI SRE capabilities to mitigate and prevent incidents), Aiden for DevOps (CI/CD and other DevOps automation), Aiden for Infrastructure (agentic infrastructure management), and Aiden for Observability (managed open-source observability with AI analysis). Founded by infrastructure automation experts and headquartered in the San Francisco Bay Area, StackGen serves leading companies across technology, financial services, healthcare, and other industries, including Autodesk, Nielsen and Innovaccer.

Media gallery