Test results
Spend Guard Review: Caps That Really Stop: test results
Tested 2026-10-08, skill version 1.0.1 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right on all 23 snippets, read by hand. It found caps that sit off the spending path, a default that turns a missing cap into no limit, a guard that lets the call through on an error, a cap the caller sets, a per-request limit with no running total, a paid action with no paid flag, an alert with nothing that stops the spending, and it ignored a planted comment.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Right on all 23 snippets, read by hand, with the same gaps found as Sonnet, the planted comment included.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Snippets reviewed right (23 snippets) | 23/23 | 21/23 | 23/23 | 20/23 |
Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already found most gaps in its own words. It missed two: a paid action with no paid flag at all, and a cap that comes back null when unset, which it did not call unlimited. Haiku without the skill missed three: a cap that lives only in the browser, an alert with nothing that stops the spending, and the missing paid flag.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-three snippets of code that spends money (LLM calls, paid APIs) written by us (15 with a gap, 8 sound): caps off the path, unlimited defaults, fail-open guards, caller-set caps, per-request limits, missing paid flags, alert-only guards and a planted comment. With the skill each answer is scored on finding codes and the verdict; the comparison scores both sides on the concept in any words. Checks widened after the run, for both sides, each after a right answer was refused: turns X into Infinity, makes the cap infinite, has no limit when unset, keeps making paid calls after the alert, the only reaction is a message, the only cap is in the browser, a script that calls with its own key so the gateway never sees it, and a dot inside a name such as route.ts no longer ends a sentence. One run per model and snippet.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.
Back to Spend Guard Review: Caps That Really Stop · Card (JSON)