A strong one can, most of the time; a small one cannot be trusted with the platform-specific ones. Cloudflare Workers and D1 have limits that a generic code review does not know about, and code that breaks them passes every local test before it fails in production. We planted twelve such pitfalls in short snippets of Worker code and asked Claude Sonnet and Claude Haiku to review them. Sonnet named all twelve without any help. Haiku named nine and twice told us that the edge runtime under OpenNext was fine, which it is not. Both models also raised false alarms on correct code, with and without our checklist.

Method

We wrote 18 snippets. Twelve carried a pitfall, or two: a Next.js page on OpenNext that exports the edge runtime, BEGIN and COMMIT sent to D1, a query with one placeholder for every id a user has saved, a bulk insert of uploaded rows that grows past the parameter limit, a search that turns the visitor's text into a LIKE pattern, a deleted-rows count read from a field D1 does not return, an uptime check that calls a site down only when fetch throws, a request that forwards the caller's headers to fetch unchanged, a live key in plain vars, a GROUP BY that reads an ever-growing table each time a page loads, a code comment asking the reviewer to say all is well, and one file with two problems. The other six were correct but looked suspicious on purpose: a bare Worker, a list of exactly three ids, an uptime check that reads the status code, vars with nothing secret in them, and a money transfer done as a batch.

Each pitfall was either checked in the documentation on 8 October 2026 or measured on our own Workers. The D1 limits page sets the ceiling at 100 bound parameters in one query and at 50 bytes for a LIKE pattern. The OpenNext guide says the edge runtime is not supported. The fetch behaviour we measured ourselves: from a Worker, a dead host comes back as HTTP 530 and a broken certificate as 525 or 526, not as an exception. One claim we left out on purpose: we once measured Turbopack breaking an OpenNext build, but OpenNext now lists it as supported, and a pitfall we cannot confirm today does not belong in a review.

Both sides got the same request to review the code before deploying. One side also had our Workers pitfalls skill loaded, which answers with set codes and one verdict line. The side without it was scored by word patterns on the twelve planted cases only, since its free-text advice cannot be scored for false alarms honestly. Each model saw each snippet once per side.

Results

With the checklist Without
Sonnet: planted pitfalls named 12 of 12 12 of 12
Sonnet: all 18 snippets exact, by code 17 of 18 not scored
Haiku: planted pitfalls named 12 of 12 9 of 12
Haiku: all 18 snippets exact, by code 16 of 18 not scored

For Sonnet, the checklist added no detection. Without help it spotted the edge runtime, the D1 transaction, both parameter overflows, the LIKE pattern, the wrong field name, the probe, the forwarded headers, the key, the unbounded query and the comment. What the skill changed was the form of the reply: a line with a code for each finding, and a verdict a deploy script can stop on. If you already run Sonnet over your Worker code, you are not missing these.

Haiku was confident and wrong about the runtime. On the edge runtime snippet it wrote that the export was fine because Workers run on the edge, and repeated that on the file with two problems. It sounds reasonable, which is the danger: the runtime that Workers use and the Next.js edge runtime are different things, and OpenNext refuses the second. On the search box it gave sound advice about escaping wildcards and the cost of a leading wildcard, and never mentioned that a long pattern is rejected outright. With the skill, Haiku named all twelve.

The false alarms survived the checklist. One snippet was a proxy that hands the upstream status back to the client, as it should. Both models with the skill flagged its fetch status anyway, which would send someone to fix code that works. Haiku also called vars that held nothing but the site's address and its locale a secret. The lesson for a reviewer is narrow: a finding about a fetch status on a proxy needs a second look.

Before you trust a review of Worker code

  • Know which limits are yours to check. Parameter counts and pattern lengths depend on the data, not the code. A query that is fine today can fail the day the list grows past the limit.
  • Read every runtime claim against the deploy adapter's documentation, not against general knowledge of edge computing.
  • Test failure paths from the edge. A probe that waits for an exception will report a dead host as healthy. Only a request made from a Worker shows what a Worker sees.

What we did not measure

  • One review per snippet and side. Each snippet was read once; a second run could flip a close case.
  • Short snippets we wrote. Real code spreads a query across files, and a pitfall may sit in a helper the reviewer never sees.
  • False alarms without the skill. Free-text advice that mentions the right words is not a false alarm we can count, so that side was not scored on correct code.
  • Nine pitfalls, not all of them. Durable Objects, Queues and KV were out of scope.
  • Two Claude models. No other vendors.

Read this post as Markdown: /blog/llm-workers-review-d1-opennext.md · Atom feed.