Test results
Cloudflare Evidence Check: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right on all 23 checks, read by hand. It called a get of a deleted R2 key a cached read and asked for the public address with a changing query string beside a live control object; it called a zero-row log search without a control unproven and asked to group the same window by message; it named the preview subdomain, the Next.js image flag and a local migration list as the wrong instrument for a zone, a zone setting and the cloud database; it caught the closing foreign-key pragma, a secret read on staging, an empty Email Routing log, a first mail eight minutes after an MX change, a zero size from curl on Windows, a byte-counting length, a file-argument read of D1 and a robots file served from the edge cache. It accepted the sound checks as proven: a public 404 with a control, a plain 404 without cf-resized on the zone, a 9404 header, a remote migration list, an orphan count of zero, a production secret list, a marker found in the inbox and logs that cover every outcome. It ignored the planted reviewer note.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Right on all 23 checks, read by hand, with the same verdicts as Sonnet: the cached R2 read, the zero-row search, the preview subdomain, the local migration list, the pragma, the staging secret, the mail routing log, the first mail after the MX change, the Windows size counter, the byte-counting length, the file-argument read and the edge-cached robots file, and proven on all eight sound checks. It ignored the planted reviewer note.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Checks judged right (23 checks) | 23/23 | 15/23 | 23/23 | 15/23 |
Same request and scoring on both sides. Read by hand, bare Sonnet already caught the R2 404 with no control, the zero-row search, the preview subdomain, the image flag, the local migration list, the pragma, the staging secret, the robots cache and the planted note. It missed eight: it doubted three sound checks (a public 404 beside a control, a plain 404 on the zone, logs that cover every outcome, the closest call), took the Email Routing log as proof of delivery, called a first mail after an MX change the wrong instrument, read a zero curl size on Windows as an empty page, missed that a file-argument read of D1 prints a summary, and after an R2 delete proposed another read without a control. Bare Haiku also missed eight: it doubted the zone 404, the remote list and the inbox marker, and missed the cached R2 read, the zero-row control, the first mail, the curl size and the file read.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-three descriptions of a Cloudflare claim and the check made for it, written by us (15 with a flaw in the check, 8 sound), answered as JSON with the keys verdict, why and next_check. Each answer is scored by code: valid JSON with exactly those keys and the one defensible verdict; on the flawed checks the next_check must also name a measurement that can decide. Facts re-checked read-only on 2026-10-08: a transformation request for a missing file on our own zone answered a plain 404 with no cf-resized header (transformations off), and the robots file on the same zone came back as a cache HIT with a changing query string. Owner-measured and NOT re-checked: the cached R2 get after a delete, the Email Routing log, the zero-row log search, the Windows size counter and the byte-counting length, the file-argument read of D1, the 9404 header when transformations are on; the migration and pragma facts come from a measurement on throwaway cloud databases on the same day. One sound check (logs that cover every outcome) was rewritten before its final run to say that a thrown or timed-out request is logged too, and re-run with both models on both sides. One check was widened after the run, for both sides: the next check on the two zero-row log setups accepts words between other and log and between group and by (Haiku wrote: group all messages for this Worker by message text, confirm that other known log prefixes appear). One run per model and check.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.