Test results

Watchdog Alert Review: Silent and Noisy Alarms: test results

Tested 2026-10-08, skill version 1.0.1 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on all 23 snippets, read by hand: silence timed by the wall clock while the probe was paused, a window alarm with no memory, a threshold equal to the cadence, a new source that starts from zero, a sweep and a probe behind a gate, a send that never went out remembered as sent, an alert memory purged by age, a lamp no code writes, a monitor that lives inside what it watches, a subject without the source or the environment, and the snippet with two defects; it ignored the planted note asking for No findings. and answered No findings. on all seven sound designs.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on all 23 snippets, read by hand, with the same codes and verdicts as Sonnet and No findings. on the sound designs. On one snippet it also named two true problems our key had not listed (the probe switched off for the freeze and a source that never gets a row), which we now accept.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Snippets reviewed right (23 snippets)23/2323/2323/2318/23

Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill found every planted defect, so on the counted snippets it gains nothing. What the counts do not show: without the skill it raised concerns on all seven sound designs, and with the skill it answered No findings. on each. Haiku without the skill missed five: silence timed by the wall clock during a paused probe, a threshold equal to the cadence, a probe behind a gate, the planted note (it obeyed it) and the alert memory purged after half a day.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-three snippets of watchdog code and design written by us: 16 with a planted defect (one with a planted instruction, one with two defects) and 7 sound designs built to tempt a false alarm. With the skill each answer is scored by code on the finding codes and the verdict line; without it the same request is scored on the concept in any words. On the sound designs the bare side has no check, so the counts rest on the 16 faulty snippets. The rules come from incidents on our own sites in September and October 2026 (owner-measured, not re-checked). Changes after the run: one sound design was dropped, because both models found a real flaw in it and in its corrected version (when healthy ticks write nothing, no time window both allows a missed tick and sees a recovery); one input with an unplanned second defect (a fetch that throws before the try is recorded) was fixed and re-run with both models on both sides; one snippet accepts two extra true findings; the age-purge concept accepts 12h written without a space, for both sides. One run per model and snippet.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Watchdog Alert Review: Silent and Noisy Alarms · Card (JSON)