AI

PHPStan Found 3 of 11 Laravel Bugs. AI Found 9 to 11

By · Sat Oct 10 2026 · 10 min read · 0 views

View as a Web Story

AISoftware#Laravel#AI code review#phpstan#code-review#static-analysis

Grid hero showing planted Laravel bugs caught by PHPStan and an AI reviewer

PHPStan at level 9 found 3 of the 11 bugs I planted in a Laravel pull request. It took 1.1 seconds and cost nothing. An AI reviewer found 9, 11 and 9 in three separate runs, in about 15 seconds each, for 7 to 11 cents. Pint and Rector found none.

The two kinds of tool did not overlap as much as I expected. PHPStan caught a wrong argument type that the AI missed in two of three runs. The AI caught every security and concurrency bug that PHPStan cannot see. If you are choosing automated code review tools for a Laravel team, the useful answer is both, in different places.

What are automated code review tools for Laravel?

Automated code review tools are programs that read a change and report defects without a human reading it first. For a Laravel project they fall into three groups: formatters, static analyzers and AI reviewers. Each group checks a different thing.

Pint is Laravel's code formatter, and it enforces style only. PHPStan is a static analyzer that reads PHP without running it and reports type errors, and Larastan is the extension that teaches PHPStan how Eloquent, facades and relations behave. Rector is a tool that rewrites code automatically, for example by adding missing return types. An AI reviewer is a language model that reads the diff and its surrounding files and reports problems in plain language.

I compared all four on one small, deliberately flawed pull request. If you want a wider comparison of commercial review products, the post Which AI code review tool to trust with PRs covers that. This post asks a narrower question: what does a free static tool already catch, and what is left for the AI?

What was the test setup?

The test setup was one Laravel 13.35 app on PHP 8.4 with two new files. DiscountService applies coupon codes to orders. OrderController exposes create, update, search, pay, list, discount and export actions. I wrote both files and planted 11 defects across five classes.

I removed every hint from the code before the review. The comments that marked each bug in my notes were deleted, and the reviewer got only the two files and a neutral prompt. The prompt asked for every real defect, one line per defect, with no praise and no style comments.

# Planted bug Class
1 first() may return null, then ->percent is read Type
2 Coupon valid through its end date expires at the start of that day Logic
3 A string is passed to an int parameter Type
4 Calls formatLabel(), which does not exist on the model Type
5 create($request->all()) with $guarded = [] Security
6 Update and pay actions have no authorization Security
7 SQL built by string concatenation Security
8 Wallet balance read, changed and saved without a lock Concurrency
9 Author loaded inside a loop Performance
10 Coupon code never validated Validation
11 Export loads the whole table with all() Performance

A twelfth bug, an unused variable, was in my first draft. I deleted it when I stripped the hints, so the score is out of 11. I mention it because a planted-bug test only means something if the answer key matches the files the tools actually read.

What does PHPStan find at each level?

PHPStan found 3 of 11 bugs at level 9 and nothing at levels 0 and 1. The PHPStan rule levels guide describes the levels as a ladder from basic checks at 0 to strict mixed-type checks at 9.

I ran Larastan 3.13 on both files at every level and counted the errors. The undefined method appeared at level 2. The wrong argument type appeared at level 5. The nullable property access appeared at level 8.

Advertisement

Level Planted bugs found Other errors reported
0 to 1 0 0
2 to 4 1 1
5 2 1
6 to 7 2 9
8 3 9
9 3 10

The jump at level 6 matters for anyone adopting PHPStan on an existing codebase. Level 6 starts to demand return types and array value types. Eight of the new errors were missingType.return and missingType.iterableValue complaints, which are real but not bugs. A team that raises the level in one step will see a wall of noise and often gives up.

PHPStan bugs found and other errors reported at each level from 0 to 9

Why did PHPStan flag a relation that works?

PHPStan flagged $order->author as an undefined property because the relation method had no return type. Larastan reads the return type of a relation method to learn that the property exists. Without it, the property looks missing.

The report was a false positive. The code works at runtime, and the line it pointed at was the real N+1 problem, but for the wrong reason. I added a return type and a generic docblock to the relation.

/** @return \Illuminate\Database\Eloquent\Relations\BelongsTo<Author, $this> */
public function author(): \Illuminate\Database\Eloquent\Relations\BelongsTo
{
    return $this->belongsTo(Author::class);
}

The property.notFound error disappeared. The lesson is to type your relations. It makes PHPStan accurate, and it removes the only noisy error that sat directly on a real defect.

What did the AI reviewer catch that static analysis cannot?

The AI reviewer caught all seven bugs that need judgment about intent: mass assignment, missing authorization, SQL injection, the wallet race, the N+1 loop, the unvalidated code and the unbounded export. Static analysis cannot see these because the code is type-correct.

A mass assignment bug is valid PHP. A missing policy check is valid PHP. A static tool has no way to know that an order should only be edited by its owner. The AI reviewer inferred it from the shape of the code, and it explained the risk, for example that a client could set wallet_cents or status directly.

It also reported problems I had not planted. I did not score these, but they were real. The apply method never calls isExpired, so expired coupons still discount. The percent is not bounded between 0 and 100, so a coupon above 100 produces a negative total. The pay action never checks whether the order is already paid, so a repeated call debits the wallet twice.

These unplanted findings are the second reason to keep an AI reviewer in the loop. It reads the whole design, not only the lines you thought to test.

What did the AI miss that PHPStan found?

The AI missed the string cast in two of three runs and the off-by-one expiry in two of three runs. PHPStan caught the cast every time, because a type checker never skips a type mismatch.

Run 1 caught 9 of 11 and run 3 caught 9 of 11. Both missed bugs 2 and 3. Run 2 caught all 11, including the off-by-one date comparison and the cast. The spread between runs is the important result. A single AI review is a sample, not a verdict.

Bug PHPStan L9 AI run 1 AI run 2 AI run 3
2 Expiry off by one Missed Missed Caught Missed
3 String cast to int Caught Missed Caught Missed
1, 4 Null and undefined method Caught Caught Caught Caught
5 to 11 Judgment bugs Missed Caught Caught Caught

Bug 2 was missed by every static tool and by two AI runs. A coupon that should be valid through its end date but expires at the start of that day needs knowledge of the business rule. Only a human, or a test written from the specification, reliably catches it.

Matrix of which tool caught each of the 11 planted bugs

Did Pint or Rector help?

Pint and Rector found none of the 11 bugs, but Rector made PHPStan quieter. Pint reported style fixes only, such as blank lines and brace positions. The Laravel Pint documentation describes it as an opinionated formatter, and that matches what I saw.

Rector's type declaration set proposed return types for four methods in the controller. I applied them and re-ran PHPStan at level 9. The error count fell from 13 to 11, because those methods no longer triggered missingType.return. The Rector project positions this as automated refactoring, and it is the cheapest way to climb PHPStan levels on a legacy app.

So the order of tools matters. Run Rector to add types, then PHPStan to check them, then Pint to format, then the AI reviewer for judgment.

How much does each check cost?

PHPStan took 1.1 seconds and Pint took 0.2 seconds, both for free. The AI reviews took 15 to 16 seconds and cost between $0.071 and $0.114, as reported in the total_cost_usd field of the Claude Code JSON output.

Tool Time Cost Result
Pint 0.2 s $0 Style only
PHPStan level 9 1.1 s $0 3 of 11, one false positive
AI review, run 1 16 s $0.114 9 of 11, 19 findings
AI review, run 2 15 s $0.071 11 of 11, 18 findings
AI review, run 3 16 s $0.075 9 of 11, 18 findings

The AI findings counts include extra items. Several were real unplanted issues, and one was an environment note that no routes referenced the controller. That is expected for a two-file test with no router. In a real pull request the reviewer would see the routes.

At these prices a review of a small pull request costs less than a cent per bug found. The cost grows with the size of the diff and the files the reviewer reads, so measure it on your own repository before you put it in CI for every push. The post Which AI model should write your code? Price per task explains how to measure cost per outcome rather than cost per token.

Time, cost and result for Pint, PHPStan level 9 and three AI review runs

How do you combine them in CI?

Run the free, deterministic checks on every push, and run the AI reviewer once per pull request. The deterministic checks give a fast pass or fail. The AI review adds written comments that a human reads.

Here is a minimal GitHub Actions job for the free half. It installs dependencies and runs Pint in check mode, then PHPStan.

name: static-checks
on: [push, pull_request]
jobs:
  checks:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: shivammathur/setup-php@v2
        with:
          php-version: '8.4'
      - run: composer install --no-interaction --prefer-dist
      - run: vendor/bin/pint --test
      - run: vendor/bin/phpstan analyse --no-progress

For the AI half, use the prompt I used. It asks for defects only, one line each, and forbids praise and style comments. Keeping the output to file:line: description makes the result easy to post as a pull request comment.

Review this pull request. Read the changed files and report every real
defect: bugs, security issues, concurrency, performance, validation.
One line per defect as "file:line: short description".
No praise, no style nits, no summary.

Set the PHPStan level to where your codebase is clean today, then raise it one level at a time. Set a baseline file for existing errors so the job fails only on new ones.

What are the limits of this test?

The limits are significant. I wrote the bugs, so they reflect what I think a Laravel bug looks like. A different author would plant different ones. Eleven bugs in two files is a small sample, and three AI runs show variation but do not measure it precisely.

I used one reviewer, Claude Code, with its default model for this account. Other reviewers will score differently, and the same reviewer will score differently next month. The PHPStan result is deterministic, but it depends on how much typing your codebase already has.

The test also had no router and no policies, so the reviewer reported some environment gaps as findings. Every number here was fact-checked against the saved tool output. About the author: I build Laravel applications and ran every command on a scratch project. If a result does not reproduce for you, use the contact page and tell me which one.

The bottom line

Use PHPStan and Rector for what they do best: fast, repeatable checks on types and structure. Use an AI reviewer for what they cannot do: security, concurrency, validation and design judgment. In this test, the free tools found 3 of 11 bugs and the AI found 9 to 11.

Neither replaces a test written from the specification. The one bug that slipped past every tool in most runs was a business rule about coupon dates.

Advertisement

FAQ

What is the best automated code review tool for Laravel?

No single tool covers everything. In my test PHPStan with Larastan found 3 of 11 planted bugs, all type errors. An AI reviewer found 9 to 11, including security and concurrency bugs. Use both: PHPStan on every push, the AI reviewer once per pull request.

What PHPStan level should a Laravel project use?

Start at the level your code passes today and raise it one level at a time, using a baseline for existing errors. In my test level 2 found an undefined method, level 5 a wrong argument type and level 8 a nullable property access.

Does Laravel Pint find bugs?

No. Pint is a code formatter, and it found 0 of 11 planted bugs in my test. It reported style fixes such as blank lines and brace positions. Use it to keep diffs clean, not to catch defects.

How much does an AI code review cost?

About 7 to 11 cents per review of a two-file pull request in my three Claude Code runs, taking 15 to 16 seconds. Cost grows with the diff size and the files the reviewer reads, so measure it on your own repository.

Comments

Loading…

Sign in to join the conversation.

Related posts