AI is in 92% of testing pipelines. So why are defects going up?

A year ago, 60% of organizations said they used AI in software testing. Today it is 92%.

That jump comes from Applause’s State of Digital Quality in Functional Testing 2026, the company’s fifth annual industry report. Only 8% of respondents say they don’t use AI for any part of testing, and 89% say it has changed how they test applications and digital experiences. By any measure, AI in the testing pipeline has stopped being an experiment.

The quality results haven’t followed. Twenty-nine percent of respondents report that the number or severity of functional defects has gone up, and 15% say both have increased. For technology leaders who approved AI tooling on the promise of shipping faster without sacrificing quality, that is an uncomfortable finding.

Faster code, same old validation

The report suggests a mismatch between how quickly AI speeds up development and how quickly quality practices have adapted. Respondents say AI has had a significant or moderate impact on development more often (79%) than on testing (72%). The gap is narrowing, but it hints at code being produced and changed faster than the methods for validating it are evolving.

Tacita Morway, Applause’s CTO, believes something specific is being lost. “Traditional automated testing answers the question: can this task be completed?” she said. “A human tester answers a harder one: could a real person work out how to do this, and get it done?”

In her view, teams are quietly dropping that second question. “Teams are giving up that second question as release cycles get faster than people can keep up with, and I think that accounts for a lot of the defect increase we’re seeing.”

That is Applause’s interpretation, and the survey shows correlation rather than proving cause. But it will sound familiar to any engineering leader who has watched a green test suite coexist with frustrated users. If AI is producing more code and more tests, the measure that matters is no longer how many tests ran but how many defects escaped to customers, and how severe they were. Organisations that still report progress in test counts may be measuring the wrong thing.

More tests are not better coverage

The survey shows AI spreading well beyond code completion. On the development side, 62% of respondents use in-editor assistants to generate or complete code, and 59% use AI for code review and documentation. In testing, the leading uses are creating test cases (65%) and writing automation scripts (62%).

Both of those are volume plays: more tests, written faster. The use case that speaks most directly to whether the right things are being tested, finding and addressing coverage gaps, is less common, at 48%. More tests do not guarantee better coverage, and the gap between those numbers may be part of the story behind the defect figures. Teams that use AI to write tests at scale, but not to interrogate what those tests leave uncovered, risk producing impressive dashboards over untested territory.

A policy is not the same as guidance

Governance appears to be lagging too. Two-thirds of organisations (66%) have documented policies on AI use in development and testing, but only 20% consider their guidelines “clear and robust.”

The distance between those two figures is where the practical work lies. Having a policy and having one that engineers can actually apply under release pressure are different things. A rule that says to use AI responsibly won’t change how a team behaves at 5 p.m. on a release day. One that spells out where AI-generated tests need human review, and who signs off, might.

Humans remain central

Despite the automation wave, almost no one in the survey argues for taking humans out of the loop. Eighty-six percent of respondents consider human involvement extremely important to functional testing, and another 13% call it somewhat important. Fifty-seven percent say humans are critical for qualitative feedback through peer review, and the same share say people are essential to designing test strategies grounded in real-world user behavior.

The challenge, as Morway describes it, is keeping that perspective while development speeds up. Adding reviewers to every release is not the answer; the aim is to put human attention where scripted checks fall short, on usability, context and judgment, and to let machines handle the rest. “Traditional automation doesn’t solve that,” she said. “Intelligent automation does — grounded in the product, the user and the risks that matter.”

The harder problem: agents that make their own decisions

Morway also flagged what may be the next big test. AI agents that make decisions and act on their own break the assumptions behind conventional testing.

“With an agent, there’s no fixed set of steps to check — the same request can take a different path every time,” she said. The failures that matter most, she added, happen when “the system is effectively making a judgment call and gets it wrong.” Her examples are inaccurate financial transactions, exposed sensitive information and dangerous guidance.

Those failures, she said, cannot be caught by confirming that a script still passes. Systems this powerful and this unpredictable need “a materially different approach to evaluation.” For organisations already deploying agents, or planning to, that means building evaluation methods before the first incident rather than after it.

The bottom line

Adopting AI in testing no longer sets anyone apart, because nearly everyone has done it. What separates organisations now is what surrounds the AI: whether they track real outcomes instead of activity, whether they check what their tests miss, whether their governance is specific enough to follow, and whether people with real-world judgment remain part of the process. Applause’s data suggests that organizations lacking those safeguards may already be seeing it in their defect rates.

AIQASoftware Testing
Comments (0)
Add Comment