AI tests hit 100% coverage and missed 8 of 43 seeded bugs
By Nihar Ranjan Das · Fri Oct 09 2026 · 11 min read · 0 views
View as a Web StorySoftware#AI testing#mutation testing#test automation#Pest#Playwright#Laravel

AI tools write tests that look complete and still miss boundary bugs. I asked Claude to write unit tests for a 26-line order-total function. The suite passed with 100% line coverage and 96% branch coverage. Mutation testing then planted 43 small bugs, and 8 of them went unnoticed.
The fix took one more round. I fed the surviving bugs back to the model, and it added five boundary tests. Branch coverage reached 100% and the mutation score rose from 81.4% to 93%. The rest of this post shows the numbers, the code, and the loop you can run on your own project.
What do AI testing tools actually do in 2026?
AI testing is the use of language models to write, run, or repair software tests. In 2026 it splits into three jobs, and each job has a different failure mode.
| Job | Example tools | What it produces | Typical failure |
|---|---|---|---|
| Unit test generation | Claude Code, Copilot, Cursor | Test files from source code | Happy paths pass, boundaries untested |
| UI test agents | Playwright Test Agents | Plans, tests, repairs | Healer skips or "fixes" a real failure |
| Self-healing platforms | mabl, Testim, Applitools | Maintained UI suites | Vendor claims, hard to verify |
Playwright Test Agents are three built-in agents that work through the agentic loop to produce test coverage, according to the Playwright documentation. The planner explores the app and writes a Markdown test plan. The generator turns that plan into test files. The healer reruns failing tests and proposes patches.
The Playwright docs state one limit plainly. The healer returns either a passing test or "a skipped test if the healer believes that functionality is broken." A skipped test is not a green test. Check your skip count after every healer run.
This post focuses on the first job, unit test generation. It is the most common starting point, and it is the easiest to measure.
Why does 100% coverage still miss bugs?
Line coverage only proves that a line executed during a test run. It says nothing about whether any assertion would fail if that line were wrong. A test can run every line and assert almost nothing.
Mutation testing is a technique that changes your code in small ways and reruns your tests against each change. The Stryker documentation defines it as introducing changes to your code, then running your unit tests against the changed code. If no test fails, the change "survived."
Mutation score is the share of mutants that your tests kill. A higher score means your tests detect more of the changes. Google applies the same idea at scale. A paper by Petrović, Ivanković, Fraser and Just, Practical Mutation Testing at Scale, evaluated it with more than 24,000 developers on more than 1,000 projects. The authors show mutants only on changed code during review.
Consider a check such as age >= 18. A mutant that changes it to age > 18 survives unless a test uses exactly 18. That is the pattern AI-written tests fall into, as the next section shows.
Advertisement
What was the experiment?
I wrote a small orderTotal function in JavaScript. It sums line items, applies a tiered discount, handles two coupons, adds tax for US orders, and waives shipping above $75. It has the kind of rules that real checkout code has.
The rules are:
- A subtotal of $100 or more earns 10% off. A subtotal of $500 or more earns 15% off.
- Coupon
WELCOME5takes $5 off. - Coupon
HALFtakes 50% off when the subtotal is above $50. - Discount can never exceed the subtotal.
- US orders pay 8% tax by default. Other countries pay none.
- Shipping is free when the taxable amount is $75 or more. Otherwise it costs $6.99 in the US and $14.99 elsewhere.
I gave Claude the function source and asked for a complete unit test suite using the Node.js built-in test runner. It returned 12 tests in one pass. Three had wrong expected values, because the model forgot shipping in its arithmetic. I checked each by hand, confirmed the function was right, and corrected only those three numbers.
I then ran two measurements. Node's built-in coverage flag gave line and branch coverage. A small script I wrote applied 43 mutations, one at a time, and reran the suite for each. The mutations swap operators such as >= for >, nudge constants such as 75 to 76, and delete statements.
What did the numbers show?
The first suite scored 100% on lines and 96.15% on branches. It also let 8 of 43 mutants survive, a mutation score of 81.4%.

Node.js documents its coverage flag in the test runner reference. The coverage report was the most reassuring part of the run. It showed full line coverage and nothing in the "uncovered lines" column. The mutation run told a different story.
Which bugs survived?
Eight mutants survived the first suite. Two are harmless and six are real gaps.
| Line | Change | Verdict |
|---|---|---|
| 14 | subtotal >= 100 became >= 99 |
Real: no test at $99.50 |
| 19 | subtotal > 50 became >= 50 |
Real: no test at exactly $50 |
| 19 | 50 became 51 |
Real: needs a price like $50.50 |
| 22 | discount > subtotal became >= |
Equivalent: same result either way |
| 20 | Math.max(discount, half) became half |
Equivalent: half always wins |
| 22 | clamp line deleted | Real: negative totals possible |
| 25 | taxable >= 75 became > 75 |
Real: no test at exactly $75 |
| 25 | 75 became 76 |
Real: boundary untested |
An equivalent mutant is a change that does not alter behavior, so no test can kill it. Line 22's > versus >= is one. The Math.max call is another, because the HALF coupon gives 50% off and the tier discount never exceeds 15%. Remove both from the count and the first suite killed 35 of 41 real mutants, or 85.4%.
The deleted clamp matters most. Without it, a $3 order with the WELCOME5 coupon produces a negative taxable amount. The customer would pay $4.83 instead of $6.99 and receive a negative tax line. The AI-written suite never tried a coupon larger than the order.
How do you fix it in one round?

Give the survivors back to the model and ask for tests that kill them. This is the something-new step, and it is cheap. Mutation tools print each survivor with its file, line, and diff. That output is a precise to-do list for a language model.
Use a prompt like this:
These mutants survived your test suite. For each one, write a test that
fails when the mutation is applied and passes on the original code.
Skip any mutant you can prove is equivalent, and say why.
line 14: subtotal >= 100 -> subtotal >= 99
line 19: subtotal > 50 -> subtotal >= 50
line 22: delete "if (discount > subtotal) discount = subtotal;"
line 25: taxable >= 75 -> taxable > 75
In my run, Claude wrote five tests from that list. They cover a $99.50 order, a taxable amount of exactly $75, a HALF coupon at exactly $50, a HALF coupon at $60, and a $3 order with WELCOME5. One of its first-draft expected values was wrong again, which I fixed by hand.
After the second round, branch coverage reached 100% and 40 of 43 mutants died. The mutation score rose from 81.4% to 93%. Excluding the two equivalent mutants, the score is 97.6%. One real survivor remains: a HALF coupon at a price like $50.50. That one needs a test at a non-integer boundary.
How do you review AI-written tests before you trust them?
Review the assertions, not the line count. A generated test is only as good as the values it compares. Use this checklist on every AI-written test file.
- Check the expected values by hand. Three of the first 12 expected values were wrong in this experiment, and the model had miscounted shipping each time. A wrong expectation that happens to match the code is worse, because the test passes and protects nothing.
- Look for snapshot-of-current-behavior tests. Some generators run the code, copy the output into the assertion, and call it a test. Such a test locks in today's bugs.
- List the boundaries in the source. Every
>=,<, and===in the function needs a test on each side of its constant. Count them, then count the tests. - Find the guard clauses. Early returns, clamps, and default values are easy to skip, as the deleted clamp showed.
- Check for mock overuse. If a test mocks the function it claims to test, it asserts nothing about real behavior.
- Run the suite against a broken copy. Change one operator yourself. If everything stays green, the suite is weaker than its coverage suggests.
Item 6 is mutation testing done by hand. It takes one minute and shows you the problem before you install any tool.
How do you run mutation testing on your own project?
Pick the tool for your language. Each one needs about ten minutes to set up.
JavaScript and TypeScript. StrykerJS is the standard tool.
npm install --save-dev @stryker-mutator/core
npx stryker init
npx stryker run
PHP and Laravel. Pest ships mutation testing in its core. The Pest documentation says it needs Xdebug 3.0 or newer, or PCOV, and a covers() call in each test file.
./vendor/bin/pest --mutate --parallel
./vendor/bin/pest --mutate --min=80
The --min flag fails the run when the score drops below your threshold. That makes it usable as a CI gate. Start at the score you have today, then raise it by a few points each month.
A boundary test in Pest. Datasets let you cover both sides of a boundary in one test. This example checks the free-shipping line from the experiment. It assumes an OrderTotal service with a shippingFor() method.
covers(OrderTotal::class);
it('charges shipping only below the free-shipping line', function (float $taxable, float $shipping) {
expect((new OrderTotal)->shippingFor($taxable, 'US'))->toBe($shipping);
})->with([
'just below the line' => [74.99, 6.99],
'exactly on the line' => [75.00, 0.0],
'just above the line' => [75.01, 0.0],
]);
The middle row is the one that kills the >= to > mutant. Generated suites usually include the first and third rows and skip the second.
Python. mutmut and cosmic-ray do the same job. The workflow is identical.
Two habits keep the cost down. Run mutation testing only on changed files in pull requests, as Google does. Run the full suite nightly.
Do UI test agents need mutation testing too?
Yes, but the technique changes. You cannot cheaply mutate a browser flow, so break the app on purpose instead. Rename a button label, change a price, or hide a field, then run the generated Playwright suite. A good suite fails on each change. A weak one stays green because it only checks that the page loads.
This also gives you a way to judge the healer. Break a selector on purpose and watch what the healer does. A healer that patches the selector and keeps the assertion is working as intended. A healer that deletes or loosens the assertion to get a pass is hiding the bug. Review its diff every time, and treat a skipped test as a failure to investigate.
Should you buy an AI testing platform instead?
Not before you know your own baseline. Mutation score is a platform-neutral metric, and it works on any suite, whether a person or a model wrote it. Measure first, then decide. The same discipline applies to your review bots, so read our guide to ai code review tools next.
Use this table as a starting point.
| Your situation | What to do |
|---|---|
| Unit tests written by AI, no mutation data | Run Stryker or Pest --mutate this week |
| Brittle UI tests breaking on every release | Try Playwright Test Agents before a paid platform |
| Many visual regressions | A visual tool such as Applitools can earn its fee |
| Vendor promises a maintenance cut | Ask for a trial on your repo, and track hours before and after |
Treat vendor figures with care. Claims such as large maintenance reductions come from the vendor's own customers, not from an independent test. The Playwright healer is free and open, so it makes a fair baseline for any paid comparison.
What are the limits of this experiment?
This is one function, one model, and one run. The mutants came from a script I wrote, not from Stryker. The 12 first-pass tests came from Claude in a single prompt, with three expected values corrected by hand. Another model, or another prompt, would produce different gaps.
So treat the 81.4% as an example, not a benchmark. The pattern is what matters. Every survivor sat on a boundary or on a defensive guard. Those are the places where generated tests are weakest, because the model copies the happy path from the source and rarely probes the edges.
You can reproduce this in an afternoon. Take your own most-changed module, generate tests, and run the mutation tool. If your score is above 90%, your process works. If it is below 80%, you have found the real cost of trusting coverage.
Advertisement
FAQ
Is 100% code coverage enough for AI-generated tests?
No. Coverage shows which lines ran, not which bugs a test would catch. In this experiment, an AI-written suite with 100% line coverage still let 8 of 43 seeded bugs pass. Pair coverage with a mutation score for a truer picture.
What is a good mutation score?
Many teams aim for 80% or higher on core business logic. Do not chase 100%, because equivalent mutants cannot be killed. Set a floor with a tool flag such as Pest's `--min`, then raise it slowly.
Can AI fix surviving mutants automatically?
Mostly yes. Paste each survivor's line and diff into the prompt and ask for a killing test. In this experiment, five new tests killed five of the six real survivors. Review every generated expected value, since models often miscalculate them.
Does mutation testing slow down CI?
It can, because the suite runs once per mutant. Limit runs to changed files on pull requests and run the full suite nightly. Pest's `--parallel` flag and Stryker's incremental mode both reduce the cost.
Do Playwright Test Agents replace manual test writing?
They cut the cost of first drafts and repairs. They do not remove review. The healer can end with a skipped test when it believes the feature is broken, so watch your skip count.
Comments
Loading…
Sign in to join the conversation.
Related posts

Do you need a vector database? Tested at 1M vectors
Exact search over 100,000 vectors took about 4 ms on a laptop. See where a vector database starts to pay off, with measured latency, recall, memory and costs.
Fri Oct 09 2026 · 11 min read · 0 views

Relational Databases: Top 10 Ranked and How They Work
Oracle still leads the October 2026 DB-Engines list of relational databases, but PostgreSQL is the only top-four system that gained ground this year. PostgreSQL scored 688.75, up 45.56 points from
Fri Oct 09 2026 · 9 min read · 1 views

What Is PostgreSQL? A Plain-English Guide With Real Tests
PostgreSQL is a free, open-source relational database that stores data in tables and answers questions written in SQL. It began in 1986 as the POSTGRES project at the University of California,
Fri Oct 09 2026 · 8 min read · 0 views