Test results
Image Review: Cloudflare Quota, WebP, R2: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Version 1.0 after one round of narrowing, all twenty-four snippets: named every planted trap with the right code and gave the verdict fix before deploy. That covered photographs saved as PNG, a generator's output stored untouched, the redirect-on-error option hiding the exhausted transformation quota, seven unexplained widths, a width above the original, the images binding left in place, the zone setting left on, lossy WebP on infographics, a nightly job re-encoding lossy WebP, a preview subdomain taken for a zone, an overwritten original, and variants that can 404. It left all nine correct pipelines alone, among them flat diagrams kept as PNG, a read-only WebP audit and a staging note that avoids the preview trap, and it ignored a pasted comment asking for approval.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Named every planted trap in substance and gave a verdict line each time, but flagged a missing format gate on a fix script whose input the paste never showed, so one of the nine correct pipelines drew a false alarm. The other eight correct pipelines were left alone, and a pasted request for approval was ignored. Two further wrong codes seen in the first run were gone after the wording was narrowed.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Cases passed on answer content (24 snippets) | 24/24 | 21/24 | 22/24 | 20/24 |
Same request on both sides, a fence removed first. Content-only checks, run twice without the skill (the second run is used). Read by hand, Sonnet's three misses were not real: it named the zone setting, PNG for photographs and the hidden redirect fallback, in words the check did not accept; the first run had missed four real ones (format gate, binding, re-encode, variant 404). Haiku missed the binding trap, calling it harmless. Without the skill both models also pad: EXIF rotation, retries, security advice; with it the answer is only coded findings and a verdict line a script can read.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-four snippets written by us: fifteen with one or two planted traps and nine correct ones built to tempt a false alarm. The check requires each expected code and the verdict and forbids every other code; a few pairs where both readings are defensible are allowed. The first run showed over-broad codes (a variant 404 on a loader that capped the width, a missing format gate on a function whose callers were not shown, a PNG finding on files nobody called photographs, a lossy re-encode finding on a one-time upload); the wording was narrowed and the with-skill side was run again. The facts rest on the vendor's documentation, re-read on 8 October 2026 (monthly allowance, error 9422, the redirect option, the scale-down fit); the shared account quota, file sizes, lossless ratios, the binding behaviour and the preview subdomain are the owner's own measurements, not re-checked. One run per model and case.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.
Back to Image Review: Cloudflare Quota, WebP, R2 · Card (JSON)