New Relic recently released its 2026 Observability Forecast, the industry’s most comprehensive report on the state of observability in the AI era. Surveying 2,575 engineering and IT leaders and practitioners across 24 countries and 12 industries, the report highlights key focus areas, challenges, and trends influencing observability investments. Notably, the data shows that despite the increased use of AI monitoring and AI-assisted observability features, high-impact IT outages are still frequent and costly.
Outage costs remain nearly flat year-over-year
The businesses surveyed lose an annualized $74 million to high-impact outages, with a mean cost of $1.85 million per hour and $30,833 for every minute systems remain down. The annualized price tag is down only slightly from the $76 million annual cost revealed last year. At the same time, high-impact outages remain a regular event, with 36% of respondents saying their companies are hit with one weekly or more. Only 15% of organisations say they never experience a high-impact outage. The average time it takes to detect high-impact outages is 41 minutes, while the average resolution time is 54 minutes. Third-party and cloud provider failures is the top cited cause of outages, followed by network failures and software change deployments.
Disruptions and manual work drain engineering capacity and cultivate “phantom velocity”
Operational friction continues to redirect technical focus away from innovation, creating a sense of “phantom velocity” where organisations move fast to deploy new code, but downtime and emergency fixes quietly cancel out those productivity gains. Engineers report spending 37% of their time addressing disruptions, up slightly from 33% last year. Additionally, 42% of organisations still learn about disruptions through inefficient channels such as manual checks and customer complaints. AIOps is deployed by only 36% of businesses surveyed today, while 35% plan to adopt it within a year.
One in four agents run blind, and half of companies monitor AI apps
For the second year in a row, AI adoption is the top cited driver of observability spend. While 83% of respondents agree that AI-generated code makes observability essential, just under half (47%) deploy AI application observability. Despite this consensus, 25% of organisations have deployed AI agents in production without any monitoring in place. An agent that writes code, changes configurations, or manages infrastructure without observability is a production incident waiting to be discovered by a customer.
Organisations monitoring agents also see greater ROI on observability
Organisations using observability to monitor AI agents in production are seeing significantly stronger returns on their observability investments. Forty-two percent of organisations actively monitoring AI agents report a 3x or greater return on observability, twice the rate of organisations that have deployed AI agents but aren’t yet monitoring them with observability tools (21%). Overall, nearly three-quarters (73%) of organisations monitoring agents in production report at least a 2x observability return. The data suggests observability is becoming an important foundation for organizations putting agentic AI into production at scale.
Observability delivers measurable ROI as the spend rises
Organisations are overwhelmingly seeing real returns on observability spend. Seventy-six percent report a positive return of at least 1x. More than a quarter (26%) report a return of 3x or higher, and a smaller group reports returns of 5x to 10x (5%). The investment is growing to match the return. Seventy-two percent of businesses expect their observability spend to increase over the next 12 months. The ROI of observability is not solely a monetary return. Since adopting observability, 70% of organisations report a faster mean time to detect (MTTD) and 69% report a faster mean time to resolve (MTTR) issues and disruptions.
OTel gains steam in the AI coding era
OpenTelemetry (OTel) is the standard that makes AI-generated code observable. Nearly three-quarters of organizations (73%) are standardised on OTel, actively migrating to it, or testing it. Only 2% have ruled it out.
“This year’s data could not be more clear. Outages are continuing to cost enterprises millions annually, and we expect that cost will remain as AI accelerates how quickly we build and ship software. AI is making it easier to create more code, applications, and services, but increase in output creates more opportunities for things to break,” said New Relic Chief Technical Strategist Nic Benders. “Yet, the data also highlights a clear path forward. Organisations monitoring their AI agents are twice as likely to yield a 3x return on their overall observability investments. By closing operational visibility gaps, organisations are best positioned to reap the benefits of AI.”