On July 20, we published an article here about the harness—the layer that executes the code generated by an AI, observes it, and derives evidence from it. It left one question unanswered, and it’s the more troubling of the two: evidence of what, exactly?
A passing test proves one thing, and one thing only. The code executed all the way through without triggering anything the test can recognize. What the test can recognize has a name: the oracle. It’s the part nobody talks about, and it’s the one that determines whether a test suite is worth anything.
The oracle is the criterion, not the scenario
In software testing, the oracle is what determines whether something is correct or incorrect. Not the test case, not the selector, not the runner. It’s the criterion.
This is not a term we made up. William Howden introduced it in 1978. The seminal survey, published by Barr, Harman, McMinn, Shahbaz, and Yoo in *IEEE Transactions on Software Engineering* in 2015, defines it as a procedure that distinguishes correct from incorrect behavior of the system under test. And the ISTQB Advanced Level Test Analyst syllabus devotes an entire section to it, under the name “test oracle problem.” Those who have earned the certification have studied it. Few have revisited it since AI began writing tests.
When a scenario verifies that a product listing shows 79.90 euros, the execution is straightforward. The key is knowing that 79.90 is the correct price for this product, for this customer, on that day, with the current promotion. When an API returns a 403, the key is knowing whether that 403 is a regression or the expected behavior for this profile.
This knowledge isn't found anywhere in the code. It comes from domain expertise. That's why it resists automation, and that's why a test harness without a solid oracle remains nothing more than a money-making machine.
A common example: A selector changes after a redesign, the test fails, and a self-healing tool locates the element and brings the test back to a passing state. No one asks why the selector changed. If the redesign also moved the validation to a different flow, the test continues to pass along a path that no user takes anymore.
Two recent studies, the same blind spot
Researchers from Virginia Tech and Carnegie Mellon evaluated eight language models on 22,374 program variants. On the original programs, the results were good: 79% line coverage, and test suites that passed. The researchers then modified the programs’ behavior and asked the models to write new tests for the modified versions. One-third of these tests failed on the code they were supposed to test. And more than 99% of those that failed passed on the original version, even while executing the modified section.
Accuracy matters: the model had the new code right in front of it. Yet it wrote tests for the old behavior. The authors refer to a residual alignment with the original program: the model reproduces what it has already seen instead of interpreting what the code actually does now. A generated test can therefore execute the correct line of code but assert the wrong result. Only someone who knows what is correct can decide.
The second study, published in June, compares AI-generated tests with tests written by developers, using real bugs in Python. The AI tests detected 69% of them, while the human-written tests detected 17.2%.
The key difference lies elsewhere. Code coverage is virtually the same in both cases: 88.5% of lines for the AI versus 84.8% for humans. What differs is the density of verification: an average of 5.5 assertions versus 3.2, and five edge cases covered versus three. The human tests, written in advance without knowledge of the defect, executed the same amount of code but verified less of it.
There’s one caveat to note right away, because it’s mentioned in the study and changes how we interpret the results: the AI-generated tests had access to the bug fix at the time they were generated, whereas the developers wrote theirs before they knew about the bug. The honest conclusion isn’t that AI writes better tests. It’s that the context available at the time a test is written determines what that test will be able to detect.
Writing a test and evaluating a result are not the same task
The objection comes naturally: if tests written by humans are less reliable than generated tests, why entrust validation to a human?
Because both studies measure test generation, not judgment. Piling up assertions, considering edge cases, leaving nothing out—that’s comprehensiveness, and a machine is better at it than a developer in a hurry. But in the study on Python bugs, it was the fix that told the AI what was correct. In the study on the 22,374 variants, no one told the AI what the new behavior should be, and its failing tests confirmed the old behavior in more than 99% of cases. In both cases, no one asked the machine to decide what was correct. Someone already knew.
In a real test suite, when a test fails, that person doesn’t exist yet. Someone has to decide whether 79.90 is still the correct price, whether that 403 error is expected, or whether that button was moved on purpose. That’s the oracle, and that’s where humans are irreplaceable. Not because they write better tests, but because they possess the knowledge of what is correct. The machine provides comprehensiveness. Humans provide the criteria.
Why the cover Became the Standard Anyway
The coverage is unbeatable—it calculates itself. No meetings, no business-side arbitration—just a percentage in the commission. The “oracle” calls for the opposite: someone who knows the product and can make the call.
However, there is a measure that would serve as a useful replacement for code coverage: the mutation score. This involves intentionally introducing defects into the code and seeing how many the test suite catches. Meta has published findings on its large-scale, AI-assisted use of this metric on its own codebases. It’s significantly more expensive to compute than a percentage of lines, and that’s precisely why few teams do it.
Measuring what can be calculated automatically rather than what really matters becomes costly precisely when the volume of tests generated skyrockets. DORA surveyed nearly 5,000 professionals in 2025: 90% use AI at work, 30% say they trust it little or not at all, and only 24% truly trust it. The next edition put a name to this gap—the “verification tax”—with a change failure rate rising from 5% to 6% and productivity gains dropping to 10% or less as soon as complex legacy systems are involved.
How This Affects a Testing Strategy
A red test isn't a problem to be eliminated. It has four possible causes: an application bug, a legitimate interface change, a change in the business workflow, or environmental instability. Two of these require an update to the test scenario, while the other two require that the test be left exactly as it is. Diagnosing the issue is less expensive than fixing it blindly, and it prevents you from silencing an alert that was actually correct.
The domain knowledge that powers the oracle is an asset, not documentation. At Mr Suricate, it’s called “application memory”: what the platform knows about your application, your user flows, and what’s normal for your business. The underlying model is replaceable—and it will be replaced many times. This memory, however, is not.
To learn more: How test maintenance works at Mr Suricate.
That leaves the question of cost. For our clients, maintaining the library takes about five minutes per month per scenario. For a library of 500 scenarios, that amounts to 42 hours per month—a quarter of a full-time position—to maintain a library that, without this effort, would quietly deteriorate.
Our Position
A pass is only valuable if someone can explain what it proves. This is true for a handwritten test, and it’s twice as true for a test generated or corrected by AI.
A sturdy harness is still necessary. But it’s only as good as the guidance you give it. The rest is just mechanics.
If you want to know what your test cases are actually checking right now, that’s a half-hour conversation.
Sources
- Haroon, Khan, Gulzar (Virginia Tech, Carnegie Mellon), "Evaluating LLM-Based Test Generation Under Software Evolution," arXiv preprint, March 2026. arxiv.org/abs/2603.23443
- Vathana, Bhatt, Patel, Eisty, LLM vs. Human Unit Tests: Fault Detection on Real Python Bugs, arXiv preprint, June 2026. arxiv.org/abs/2606.08588
- DORA 2025: Adoption and Trust, Nearly 5,000 Respondents. cloud.google.com
- DORA 2026, audit fee, InfoQ summary. infoq.com
- Barr, Harman, McMinn, Shahbaz, Yoo, “The Oracle Problem in Software Testing: A Survey,” IEEE Transactions on Software Engineering, 2015. ieeexplore.ieee.org
- ISTQB, glossary, test oracle. glossary.istqb.org
- CFTL, ISTQB CTAL-TA v4.0 syllabus in French, Section 1.3.4: Determining Test Oracles. cftl.fr
- Meta Engineering, " LLMs Are the Key to Mutation Testing and Better Compliance," September 2025. engineering.fb.com



