Building an in-house tool has never been more accessible. A developer aided by an LLM can deliver a first version of a tool in a timeframe that would have seemed unrealistic three years ago, and knowing how to do it is hardly an issue anymore.
That leaves one less comfortable question—the one every CIO eventually asks: Who will carry forward what we’ve just built for the next five years?
What Has Really Declined
The cost of producing code—and the decline is real. Up to 30% of Microsoft’s code is written by AI, according to Satya Nadella in April 2025. The cost of inference fell from $20 per million tokens in November 2022 to $0.07 in October 2024, with performance remaining constant—a 280-fold decrease, according toStanford HAI’s 2025 AI Index.
To dispute these figures would derail the conversation right from the first paragraph. Writing code has become inexpensive, and there’s no going back.
Code development has become a minor budget item. That is no longer where the cost of software is determined.
What Hasn't Declined
Software doesn’t just cost money on the day it’s written. It continues to cost money every month afterward: changes requested by the business, version updates for dependencies, infrastructure, documentation, and training the person who takes over when the original developer leaves. None of these costs decrease just because the code was produced more quickly.
The mechanism even works the other way around. When production costs are lower, output increases, and each additional line of code adds to the legacy that must be maintained. An application portfolio built twice as fast isn’t half as expensive to maintain. Above all, it’s larger.
And what went up
Verification. The DORA 2026 report gives it a name—the “verification tax”—and projects an increase in the change failure rate from 5% to 6%. It estimates the gains at 35% to 40% for a new project, but at 10% or less for a complex existing one, according to the summary published by InfoQ in May 2026.
The 2025 edition of the same report, based on nearly 5,000 respondents, already showed this trend. 90% of tech professionals use AI at work. Thirty percent say they trust it little or not at all, and only 24% truly trust it. Its adoption is positively correlated with delivery velocity and negatively correlated with stability. DORA describes it as an amplifier: a strong team becomes even better, while a fragile team becomes more fragile—and faster.
One final finding puts a damper on overly optimistic estimates. In July 2025, METR found that experienced open-source developers spent 19% more time using AI, even though they estimated afterward that they had gained 20%. The trial was randomized, involving 16 developers and 246 real-world tasks in their own code repositories. The sample size is small; these are experts working on code they know by heart, using tools from early 2025, and the authors themselves caution against generalizing the results. METR attempted to replicate the study in early 2026 but abandoned the effort because it could not assemble a control group: too many developers refused to work without AI. What remains useful, however, is the discrepancy between the perceived gain and the measured gain.
AI hasn't eliminated the burden; it has simply shifted it to the person doing the verification.
Automated Testing: A Textbook Example
Writing your first Playwright test costs next to nothing, and with an LLM, it costs even less. It’s better to say it before the reader thinks it.
Writing a test and maintaining a test suite are two distinct tasks. The latter involves ensuring the long-term sustainability of daily execution, stability, data, environments, browsers, failure diagnosis, reporting, the integration pipeline, and the skills of those who carry out all of this. Instability alone is an engineering challenge, even in the most mature organizations: a test that passes and then fails without any changes having been made wastes time every time a decision must be made as to whether it is giving false results.
AI-generated tests do not address this second challenge. A study published in March 2026 examined 22,374 program variants and eight language models. The tests are generated based on the original version of the program, and they pass. Then the code evolves and its behavior changes. The tests turn red, and more than 99% of these red tests pass on the original version while executing the modified section. The authors refer to this as residual alignment with the initial behavior. In other words, a failing test indicates that something has changed; it does not indicate whether the change is correct, and someone must decide on a test-by-test basis (Haroon, Khan, Gulzar, arXiv 2603.23443). The study focuses on test generation, not on assisted test correction: applying it to fleet maintenance is an analogy, not a demonstration.
For our clients, maintaining their test suite takes about 5 minutes per month per scenario. This is an internal observation, not a published figure. We provide the unit rather than a total; it’s up to you to multiply it by the size of your test suite. What this effort entails is detailed in our article on the cost of maintaining automated tests.
Automating a test has become easy. Maintaining a reliable suite of tests over the long term is not.
Three Questions to Ask Yourself Internally
Expertise doesn't tell you where to invest. What matters is whether your team wants to dedicate its expertise to that area for the next five years. Mastering a subject doesn't mean you have to build and maintain all the technology related to it.
- How many people actually know how to maintain this tool today?
- How much of their time is spent maintaining it rather than creating value?
- If your net worth doubles, do you have to double your efforts to maintain it?
If the answer to the first question is "one person," the issue becomes a risk to business continuity rather than a technical decision.
What This Means for Us
When a scenario fails, the incident is escalated along with a video of the failure, and you can ask the AI for a suggested fix. It starts with the intent of the test—what the scenario is supposed to prove—not just the element that caused the failure. Your team then validates it.
The division of roles is simple. You determine what needs to be tested, why, what behavior is correct, and which risks are critical. This is what we call the test intent. Mr Suricate Playwright handles execution, maintenance, monitoring, and scaling. And since each scenario can be exported to Playwright, none of this ties you down.
How much does it really cost to maintain your test environment?
See how Mr Suricate a scenario library is maintained over time, and what your team has at their disposal.
Frequently Asked Questions
Writing, yes, often—especially with an LLM. The equation changes when you factor in what comes next: day-to-day operation, stability, data, troubleshooting, and taking over the system when the developer leaves. It’s this total cost of ownership that needs to be compared, not the cost of writing.
It lowers construction costs—that’s undeniable. It does not reduce ownership costs, and it adds a verification burden that the DORA 2026 report refers to as the “verification tax.” The calculation still needs to be done, but it should cover the tool’s entire lifecycle, not just its construction.
This is what software costs over its entire lifecycle, beyond its initial development: maintenance, updates, infrastructure, documentation, training teams, and—now—verifying the output of AI. For an internal tool, this is the line item that lasts the longest.
For our clients, maintaining their fleet takes about 5 minutes per month per scenario. This is based on our internal observations. Multiply that by the size of your fleet to get a rough estimate, then compare it to the time your team currently spends on this task.
With Mr Suricate, scripts can be exported to Playwright. What has been built can still be used outside the platform.
Keep your expertise in-house: what needs to be tested, why, and what constitutes correct behavior. Execution, maintenance, and monitoring can be handled by a platform, provided that the system remains exportable and your team retains control over what is considered correct.


