Test results
Test Guard Review: Can This Check Fail?: test results
Tested 2026-10-08, skill version 1.0.1 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right on all 24 scripts, read by hand: a completeness guard that counts its own list (JavaScript and Python), a one-star glob that skips subfolders, a mutation script that restores with git checkout, a null test on a field that comes back empty, a curl that follows the redirect to a login page, two polling loops without a sleep that measure the network instead of the time, a test that calls the pure function and bypasses the handler, a control that cannot create its condition, a ceiling never tried with an absurd value and a fix tested on one of two mirrored branches; it named the input that would turn each one red and answered No findings. on all nine sound scripts.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Right on all 24 scripts, read by hand, with the same findings as Sonnet. On sound scripts it was stricter than our key three times, and each point was fair: a mutation run that exited 0 with a surviving mutant, a wait loop that reported live when the expected version was empty (we fixed both scripts and re-ran them), and a vitest run that never starts being counted as a kill.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Scripts reviewed right (24 scripts) | 24/24 | 22/24 | 24/24 | 16/24 |
Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already caught the self-counting list, the one-star glob, the git checkout restore, the empty-string test, the login redirect, the bypassed handler, the control without its condition and the untested mirror. It missed both polling loops: a loop that counts tries without sleeping finishes in seconds and reports minutes, and it looked elsewhere in both scripts. Without the skill it also raised concerns on most sound scripts, which is not counted. Haiku without the skill missed eight, among them the git checkout restore, the empty string, the login page and the absurd value.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-four test and check scripts in JavaScript, Python, Bash and PowerShell written by us: 15 with a way to pass while the guarded code is broken and 9 sound ones. With the skill each answer is scored by code on the finding codes and the verdict; without it the same request is scored on the concept in any words. On the sound scripts the bare side has no check, so the counts rest on the 15 faulty scripts. The rules come from cases measured on our own projects in August and September 2026 (owner-measured, not re-checked). Checks widened after the run, for both sides, each because a right answer was refused: the self-counting list accepts "a seventh module", "stays" and "scanning src"; the login redirect accepts "redirects it to /login"; the bypassed handler accepts "never touches the handler"; the mirror accepts "only exercises runNow". Two sound scripts had real flaws (found by Haiku) and were fixed and re-run with the skill. One run per model and script.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.
Back to Test Guard Review: Can This Check Fail? · Card (JSON)