Test results
SKILL.md Review: Safe and Within the Limits: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Gave the exact codes and verdict in all 20 cases, as bare lines with nothing around them. It caught a request to read the SSH key, a silent copy to a shared drive, a line that claims to replace the system's instructions, a note telling the reviewer to answer No findings, a live-looking API key, a curl piped into sh, a name over 64 characters, an uppercase name, the reserved word claude in a name, a missing, an overlong and a first-person description, an XML tag in the description, a Windows path, a 538-line body and a missing example, and it left a skill that runs its own bundled script unflagged. A first run with a wider wording of our person rule flagged the word you inside ordinary descriptions; the rule was narrowed to match the published guidance and run again.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Also 20 of 20 with the skill, with the same codes and verdicts as Sonnet, including the counted limits it did not flag without the skill: the 65-character name, the 1,305-character description and the 538-line body.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Every planted problem named, nothing invented | 20/20 | 18/20 | 20/20 | 14/20 |
Both sides scored with the same word-based checks, since the side without the skill does not know our codes. Without the skill both models noticed the hidden orders, the key and the curl command, and Sonnet called out the note aimed at the reviewer; the misses were the counted rules. Haiku did not flag the 65-character name, the 1,305-character description, the 538-line body or the backslash path, and neither model said that the file had no example. Two baseline patterns first missed correct answers (one counted a sentence saying no injected instructions as an alarm) and were widened before the final score; the change applies to both sides.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty small SKILL.md files written by us, each with one or more planted problems or none: a clean file, three name defects, four description defects, four kinds of hidden order to the agent including one aimed at the reviewer, a pasted credential, a download-and-run command, a Windows path, no example, an overlong body, a legitimate bundled script, and one file with three problems. The check requires each expected code and the verdict and forbids every code that should not appear. The limits in the skill were checked against the public skill documentation on 8 October 2026. One run per model and case; the with-skill side was run twice because our wording of one rule was too wide the first time, and the first run is kept in the test folder.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.
Back to SKILL.md Review: Safe and Within the Limits · Card (JSON)