The strong model did; the small one mostly did, and its one miss is the kind that costs data quietly. We handed Claude Sonnet and Claude Haiku five short changes, each hiding a problem a senior engineer would stop, with no instructions beyond "review this change". Sonnet named every planted problem. Haiku named four. On the fifth, a sync job that throws away rows it failed to save, Haiku not only missed the loss but described the faulty behaviour as a safe default.
Method
Claude Fable, a third model, wrote the changes for this experiment; we read each one and dropped any where two careful reviewers could disagree. They are deliberately small, a single function or a short diff, so that missing the problem cannot be blamed on length:
| Change | What hides in it | What a good review says |
|---|---|---|
| Coupon redemption, TypeScript and Postgres | a transaction and a guarded UPDATE, but nobody reads how many rows the UPDATE touched | two simultaneous requests can both credit the user |
| Keyset pagination, Python | results sorted newest first, cursor compared with greater-than | page two returns the newer rows again |
| File deletion, Go | the path comes from a stored filename whose origin is not in the diff | name the traversal risk as an assumption, not a fact |
| Sorting endpoint, JavaScript | a column name in a template string, picked from a fixed list | nothing significant; the scary-looking line is safe |
| Order sync, Python | failed rows are logged and skipped, then the sync position moves past them | those orders are never fetched again |
The coupon case leans on a documented detail: in PostgreSQL's default isolation level, a plain SELECT reads a snapshot and locks nothing (PostgreSQL on transaction isolation). Both requests can see an unused coupon at the same moment.
Each model reviewed each change once, from a terminal, with tools disabled and none of our own configuration loaded. Our automatic checks looked for the planted problem in each answer, and we also read every answer ourselves. That second step mattered here, as the results show.
Results
| Change | Sonnet, plain prompt | Haiku, plain prompt |
|---|---|---|
| Coupon redemption | caught the double credit, suggested checking the claimed row | caught it |
| Keyset pagination | caught the wrong comparison | caught it |
| File deletion | flagged traversal, explicitly depending on the upload code | flagged it, hedged with could |
| Sorting endpoint | judged it clean, with minor notes | judged it clean |
| Order sync | caught the permanent loss | missed it, praised the behaviour |
The order sync is worth reading slowly. The function upserts each order inside its own try block, with a comment explaining that one bad row must not stop the rest. Two lines later it saves the newest timestamp in the batch as the place to resume from. Any order that failed is now older than that mark, and the next run asks only for newer ones. Haiku reviewed the try block, which is reasonable, then told the author that not moving the timestamp after a failure was the safe choice. The code does the opposite.
Our own automatic check let this answer through. It searched the review for words about lost data, and Haiku had used one while describing a different issue, the missing loop over further pages. We added a pattern for the exact misreading and ran it against both stored answers: it rejects Haiku's and accepts Sonnet's. A keyword check on a review can be satisfied by the wrong finding, so we would not trust one without reading the text.
Then we gave Haiku the same change once more with our Code Review skill loaded. This time it opened with a critical finding and a concrete timeline: orders stamped 10, 11 and 12, the middle one fails, the mark moves to 12, and order 11 is gone. The fix it proposed was to move the mark only to the newest order that saved. Run its own example through that rule and the mark still lands on 12. A correct fix keeps the resume point below the earliest failure, or stores the failed ids for a retry. Alongside, it raised two more problems that nothing in the diff backs up.
Two habits follow from this. When a batch job swallows errors, look at what happens to its cursor or watermark after the swallow, not only at the catch. And when a model proposes a fix, feed its own example through that fix before you merge it.
What we did not measure
- Each answer here is a single sample. One miss proves a miss is possible; it says nothing about its frequency on a second try.
- The changes are few and short, their author knew what each was built to test, and the people judging them also sell the skill. Real pull requests are longer and messier.
- Sonnet and Haiku were the only reviewers, both from Anthropic and launched identically. Reviewers from other vendors, or dedicated review tools, may do better or worse.
- The skill got a second run only on the one review that went wrong. We cannot tell from this whether it would change any of the four that were already right.
- We read for whether the planted problem was found, not for how severity was ranked or how many extra findings each review added.
Read this post as Markdown: /blog/llm-code-review-safe-default.md · Atom feed.
