Automated Test Maintenance: How Much Does It Really Cost?

A watchmaker repairing the movement of a mechanical watch at his workbench

Writing an automated test costs next to nothing these days. A language model can generate one in a matter of seconds, and the entire industry has realized this.

But the cost hasn’t gone away. It has shifted to what comes next: running the tests every day, fixing them when the application changes, and knowing how to interpret a failure. That’s where the real budget for a test suite comes into play, and it’s rarely what people consider when choosing a tool.

This guide provides order-of-magnitude estimates, a method for analyzing failures, and questions to ask before you end up with a fleet that no one can maintain.

Why is this question coming up now?

Google Cloud’s DORA 2025 report, based on a survey of nearly 5,000 professionals, reveals two figures that complement each other. Ninety percent of tech professionals use AI at work—a 14-point increase in one year. However, 30% say they have little or no trust in it, and only 24% truly trust it.

The same report paints a harsher picture for quality teams: the adoption of AI is positively correlated with delivery velocity and negatively correlated with stability. In other words, we deliver faster and experience more breakdowns. DORA describes AI as an amplifier: a strong team becomes even stronger, while a fragile team becomes even more fragile—and faster.

The 2026 edition refers to this phenomenon as the “verification tax” and models an increase in the change failure rate from 5% to 6%, with productivity gains of 35% to 40% on new code but 10% or less on complex legacy code.

A study published in June 2026 by Adaptavist—based on a survey of 2,500 knowledge workers in the United Kingdom, the United States, Canada, Germany, and Spain conducted in March—reaches the same conclusion by a different route. Forty-two percent of respondents reported spending more time verifying what the AI produces than they save by using it. Fifty-two percent regularly correct AI-generated work produced by their colleagues. The authors’ conclusion can be summed up in one sentence: AI does not eliminate the workload; it redistributes it.

These are statements, not time measurements, and the study does not focus specifically on quality teams. But it points out two things that hold true exactly as described in a test environment. Verification is becoming a full-fledged job role, just like production. And the responsibility doesn’t fall on the person who wrote the test: in a test lab, it’s almost never the author of the test scenario who reads the “red” on Tuesday morning.

The amount of code being produced is increasing, but confidence in that code is not. Verification is becoming the bottleneck, not writing code.

What makes up the maintenance cost?

An automated test suite consists of four components, and writing is not one of them.

Execution. A test that doesn't run every day doesn't protect you from anything. You need an infrastructure that executes, schedules, and manages environments and data—and that holds up when the pre-production environment is unavailable.

Debugging. An app changes. A drop-down menu changes, a workflow gains a step, a form’s label changes. Every change automatically breaks scenarios that haven’t changed.

Diagnosis. This is the most underrated step. When faced with a failed test, someone must determine which of four causes is responsible: a genuine bug, a legitimate interface change, an unstable environment, or outdated data. Three of these four causes are not bugs, and each requires a different course of action.

Test data. An expired dataset, an account locked after three failed login attempts, a shopping cart with out-of-stock items: that’s where most of the false positives come from.

How much does it cost, roughly speaking?

It is important to distinguish between two things that are often confused.

Ongoing workload. For our clients, maintaining the library takes about 5 minutes per month per scenario. This is an internal observation, not a published figure, and we present it as such.

The cost of an incident. When a scenario fails, it takes about an hour to restore it. This is a per-incident cost, not a monthly fee.

Those 5 minutes only matter when applied to your park. Based on 100 scenarios, that amounts to about 8 hours a month—or one day. Based on 500 scenarios, that’s about 42 hours—or a quarter of a full-time job—every month, just to keep the park green and credible.

This calculation is a good habit to get into before signing anything: take the target fleet size, multiply it, and ask your supplier or your team who will cover those hours.

On the other end of the spectrum, the cost of not testing is also quantifiable. Boehm and Basili, in a 2001 article in IEEE Computer, show that a defect fixed after delivery costs about 100 times more on large systems, and about 5 times more on small, non-critical projects. The trend is more important than the exact multiplier.

Reading a move—the skill that costs the most

A test suite doesn't produce "green" tests; it produces information. But you still have to read it.

Actual bug. The expected behavior is no longer present. This is the only situation that warrants opening a ticket.

Legitimate interface change. The product team intentionally modified a user flow. It's expected that the test would fail, and it needs to be updated, not fixed.

Instability. The test turns green again on the next play even though nothing has changed. That’s what we call “flaky.” It’s a park’s worst enemy, because it destroys trust: once it reaches a certain rate, no one pays attention to the reds anymore.

Outdated information. The test account has expired, the product is out of stock, or the promo code is no longer valid.

A team that can't distinguish between these four cases spends its time manually reclassifying them and eventually disables the noisiest scenarios. That's how a park dies: not all at once, but through gradual neglect.

What actually helps: visual evidence of each failure, step-by-step details, automatic replay before triggering an alert, and grouping incidents by cause rather than issuing an alert for each scenario.

Who supports the park over the long term?

The Adaptavist statistic cited above takes on its full meaning here. When one out of every two people corrects a colleague’s AI-generated work, the question is no longer who writes the tests, but who will review them in two years.

That’s the question no one asks when choosing a tool—and the one that makes all the difference three years later.

A test suite written in code belongs to those who know how to read it. When the person who built it switches teams or leaves the company, files remain that no one dares to touch. The test suite continues to run for a few months, then the errors pile up, and eventually it’s shut down.

Three questions to ask before you get there:

  • How many people today know how to edit an existing script on their own?
  • If one of them leaves tomorrow, how long will it take for another one to take over its responsibilities?
  • What can still be reused if you switch tools or service providers?

This last question is the most important one, and it has a technical answer: reversibility. A platform whose scenarios can be exported to a standard format like Playwright remains your property, even if you leave. This is also what allows an automation QA team that writes code and business teams that do not write code to work together on the same platform.

Take over an existing fleet

Taking over a game world that you didn't create requires following a specific order; otherwise, you'll spend weeks fixing scenarios that were already obsolete.

  1. Take stock of what’s actually working. A script that’s been disabled for six months isn’t an asset—it’s a liability.
  2. Measure the failure rate and the rework rate. This is your true level of instability, and it determines everything else.
  3. Sort by business criticality, not alphabetically. Prioritize the paths that drive revenue first.
  4. Fix the top of the list; remove the bottom. A pool of 80 reliable scenarios is better than a pool of 300 that are suspect.
  5. Document the test data before making any changes to the test scenarios themselves.

What AI Changes—and What It Doesn't Change

AI writes tests—and does so quite well. A study published in June 2026 on real bugs in Python even shows that AI-generated tests detected 69% of defects, compared to 17.2% for tests written by humans, with nearly identical code coverage. There’s an important caveat: the AI-generated tests had access to the bug fix, whereas the human-written tests did not. That said, it’s worth noting in passing that code coverage says nothing about the ability to detect a defect.

Where caution is needed is regarding duration. A March 2026 study involving 22,374 program variants and eight language models shows that more than 99% of the generated tests that fail on a modified version of the code passed on the original version while executing the modified section. The authors refer to this as residual alignment with the initial behavior: a red result indicates that something has changed; it does not indicate whether the change is correct. The study focuses on test generation, not on automatic repair tools, but it illustrates the scale of the problem: generating tests is easy; keeping up with changes is difficult.

A useful counterpoint, at last. A randomized METR experiment in July 2025 found that experienced open-source developers were 19% slower when using AI, even though they estimated afterward that they had gained 20%. The sample size is small—sixteen developers working on code they know by heart—and the authors explicitly state that the results should not be generalized. But the gap between perceived and measured performance is worth keeping in mind when drawing up a budget.

Reducing Maintenance: Practical Steps

  • Test the visible behavior, not the technical structure. A scenario based on what the user sees will survive a code rewrite; a scenario tied to technical identifiers will not survive a change in the framework.
  • Reuse building blocks. A connection, a payment tunnel, and cookie consent can all be reused. Fixing one building block fixes all the scenarios that use it.
  • Isolate test data. Use dedicated, renewable accounts that are never shared with manual acceptance testing.
  • Replay before alerting. An automatic replay before the alert eliminates much of the noise without changing the scenarios.
  • Keep the evidence. A video of each failure and a step-by-step breakdown eliminate the back-and-forth of qualification, which is a complete waste of time.
  • Keep it reversible. Exporting to Playwright ensures that the work you’ve put in isn’t lost.

How much does it cost to maintain your infrastructure?At Mr Suricate, we start with your scenarios, your deployment schedule, and the state of your environments to give you a rough estimate of the scope of your project and show you how we handle test maintenance.
Schedule an appointment

Frequently Asked Questions

How much does it cost to maintain an automated test suite?

Calculate routine maintenance in minutes per month per scenario, and the cost of fixing a broken scenario in hours. For our clients, routine maintenance takes about 5 minutes per month per scenario. For a total of 100 scenarios, that amounts to about one day per month.

How can I prevent my tests from failing every time there's an update?

By basing them on observable behavior rather than on the page’s technical structure, by sharing reusable building blocks, and by isolating test data. A test that describes what the user does survives a redesign; a test that describes the code does not.

How can you tell if a failed test is a real bug or just an instability?

A scenario that returns to the "green" status on the next run, without any changes having been made, is an instability, not a bug. This distinction is made clear through an automatic rerun before an alert is triggered, visual confirmation of the failure, and a step-by-step breakdown.

Who maintains the tests when the person who wrote them leaves the company?

No one, if the platform is only understandable to its creator. This is the main risk of a platform that is entirely coded and owned by a single person. Two solutions: a scenario format that is understandable to non-developers, and reversibility that ensures the work remains usable elsewhere.

How do I take over an existing test suite?

Take inventory of what’s actually running, measure the instability rate, rank by business criticality, fix the items at the top of the list, and eliminate those at the bottom. A small, reliable fleet is better than a large one whose results no one trusts.

Do we still need a testing tool if AI writes the tests?

Writing code is no longer the problem. Day-to-day operations, troubleshooting when the application malfunctions, and diagnosing failures remain the challenges—and these are what drain the budget. Today, a testing tool is used to maintain a system over the long term, not to generate test scenarios.

Image by François-Xavier Le Gal

François-Xavier Le Gal

François-Xavier Le Gal is Deputy CEO of Mr Suricate, a French provider of a no-code SaaS solution for automated testing and monitoring. He helps companies ensure the reliability of their digital experiences and manage software quality, including functional, non-regression, performance, accessibility, and compliance testing. On the Mr Suricate blog, he shares insights, methodologies, and real-world feedback on automated testing, QA, and digital performance.

Find him on LinkedIn

See also

Switch from manual testing to automated testing without writing any code

In 30 minutes, we'll show you how to cover your critical test cases, detect regressions before your users do, and maintain your test scenarios over time.