Test results

Workers Pitfalls Review: D1, OpenNext, Fetch: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Named every planted pitfall with the right code and verdict: the edge runtime under OpenNext, BEGIN and COMMIT against D1, an inArray over saved ids and a bulk insert of uploaded rows past 100 parameters, a LIKE pattern built from the search box, rowsAffected read from a D1 result, an uptime probe that only treats a thrown error as down, client headers forwarded to fetch, a live key in vars and a GROUP BY over a growing table on every page view, and it ignored a comment telling the reviewer to answer No findings. It left six correct snippets alone, including a fixed three-item inArray, public vars and a batch() transfer. It failed one case by our strict check: on the header-forwarding proxy it also flagged the fetch status, which a proxy correctly passes through.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Found every planted pitfall as well, but raised two false alarms: the same fetch-status flag on the proxy, and a secret-in-vars finding on vars that held only a public site address and a locale.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Planted pitfalls named (12 snippets)12/1212/1212/129/12

One run per model and snippet, the same word-based check on both sides. Sonnet found all twelve planted pitfalls without the skill too, so for Sonnet the skill adds no detection, only the fixed codes and verdict a program can read. Haiku without the skill called the edge runtime under OpenNext fine, twice, and did not mention D1's 50-byte LIKE limit; with the skill it named all twelve. On correct code the side without the skill gave general advice, such as validating a transfer amount, which this test neither rewards nor penalises.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Eighteen short snippets written by us: twelve with one or two planted pitfalls and six correct ones built to tempt a false alarm (a plain Worker, a fixed small inArray, a status-checking probe, public vars, a batch() transfer, a clean handler). The check requires each expected code and the verdict and forbids every code that should not appear; on two snippets where both readings are defensible (a LIKE with a leading wildcard, a lookup on a column whose index is not shown) an extra unbounded-query finding is allowed but not required. Every pitfall in the skill was checked against Cloudflare's and OpenNext's documentation or measured in our own production on 8 October 2026; claims we could not re-check, such as Turbopack breaking OpenNext, were left out. One run per model and case. The comparison without the skill covers only the twelve snippets with a planted pitfall: whether the side without the skill raises a false alarm on correct code cannot be scored reliably by keywords (it gives sound general advice that mentions the same words), so false alarms are scored strictly, by code, on the side with the skill.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Workers Pitfalls Review: D1, OpenNext, Fetch · Card (JSON)