Test results

Screenshot Evidence Check: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on all 23 captures, read by hand: it called the 420-pixel headless window a crop and not a phone defect, the blank hero image in a very tall capture, the pale logo that no window-size flag can test, the document-width reading beside a clipping container, the empty reveal sections, the local Lighthouse score that fell while observed paint got faster, the automation session that had stopped running scripts, the zero download size from curl on Windows, the production-only look at a media-host setting and the hydration date. It also called the sound captures proven (element measurement, hosted audit, reduced-motion capture, touch emulation with its in-page check, a normal-height hero, a steady idle Lighthouse run, a control page that worked, a staging lightbox whose image host differed from the production fallback) and did not invent doubt.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on 22 of 23 captures, but it missed one: on the local-versus-hosted accessibility score its reply was not valid JSON (a comma where a colon belongs), and its verdict there was not-proven where the check expects contradicted. Everything else matched, including the cropped 420-pixel window, the tall capture, the zero curl size and the sound captures, the staging lightbox among them, which it called proven.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Captures judged right (23 captures)23/2320/2322/2318/23

Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already knew every flawed capture: the cropped narrow window, the tall capture, touch emulation, the clipping container, the noisy Lighthouse score, the zero size on Windows. What it lacked was trust in good evidence: it doubted three sound captures (a reduced-motion capture showing all six cards, an emulated touch session whose in-page check returned true, a control page that worked). Haiku without the skill missed five: it called the 420-pixel crop not-proven, took a production-only look for a tool artifact, doubted a sound hero image blamed the tool where the control worked, and doubted the staging lightbox although its host could only have come from the variable.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-three descriptions of how a page was captured or measured, written by us (15 with a flaw in the evidence, 8 sound), answered as JSON with four keys. Each answer is scored by code: valid JSON with exactly the four keys and a fixed verdict; on the flawed ones the next check must also name a measurement that can decide. Facts re-checked in a local headless Chrome 154 on 2026-10-08: a window of 420 and 460 laid the page out at 500 and the 420 picture is a crop; hover: none is false in a headless window; scrollWidth equals clientWidth beside an overflow-x: clip child. Owner-measured and NOT re-checked: the blank images in a 9000 px tall capture, the Lighthouse spread, hosted versus local scores, the automation session that stops running scripts, the zero download size on Windows, the staging, date and client-bundle facts. Checks widened after the run, for both sides: the verdict on five setups accepts a second defensible verdict (tall capture, Lighthouse faster, automation session, 390-pixel breakpoint, everyday browser speed); the next check on the grid case accepts a capture of the commit before the grid change or a one-column view; the next check on the zero-size case accepts reading the size in the storage itself. One setup (staging lightbox) did not say what the fallback value is, so Haiku could defensibly doubt it; the input was fixed and re-run with both models on both sides, and it is counted. One run per model and case.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Screenshot Evidence Check · Card (JSON)