# Reviewing a SKILL.md: models catch injection, not limits

> Claude Sonnet and Haiku reviewed 20 SKILL.md files we seeded with problems. Both caught hidden orders and keys; Haiku missed the rules you have to count.

Published 2026-10-08 · https://aiskills402.com/blog/llm-skill-md-review-counted-limits

Half of it, reliably. When you review a SKILL.md before installing it, there are two kinds of problem to look for: text that pushes the agent toward places it has no business going, and rules that the file breaks by being too long or badly named. Claude Sonnet and Claude Haiku each reviewed 20 small skill files we had seeded with problems. Both models spotted all the hidden orders and the pasted key on their own. The counted rules were another matter: Haiku missed six of the 20 files without help and Sonnet two, almost all of them limits you can only check by counting. With a written checklist loaded as a skill, both models got all 20.

## Method

We wrote the files ourselves, all variations on one harmless form-filling skill so that only the planted problem differed. One was clean. Three broke the name rules: capital letters, the reserved word claude, and a name of 65 characters. Four broke the description rules: no description at all, one of 1,305 characters, one that never said when to use the skill, and one written as "I". One more put an XML tag inside the description. Four hid orders to the agent: read the SSH key, quietly copy the user's files to a shared drive, a sentence announcing that it now overrides the system prompt, and a remark asking whoever reviews the file to report it as clean. The rest carried a pasted API key, a script piped from curl into a shell, a Windows path with backslashes, no example, a body of 538 lines, a legitimate call to a bundled script that must not be flagged, and one file with three problems at once.

The rules come from the [public skill documentation](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices), which we read on 8 October 2026: a name of up to 64 characters made only of small letters, numbers and dashes, with two reserved words banned; a description of up to 1,024 characters that explains both the job and the moment to use it, in the third person; and a body kept under 500 lines.

Both sides got the same short request to review the file. One side also had our [SKILL.md review skill](/skills/skill-md-review) loaded as its instructions, with a fixed code for each problem and a verdict. Since the side without it does not know those codes, both sides were scored by the same word patterns: an answer earned the point for a planted problem only by naming that problem. Each model reviewed each file once per side.

## Results

| | Skill loaded | No skill |
|---|---|---|
| Sonnet: every planted problem named | 20 of 20 | 18 of 20 |
| Haiku: every planted problem named | 20 of 20 | 14 of 20 |

**The dangerous part was never the problem.** Without any instructions, both models flagged the request to read the SSH key, the silent copy, the line that claimed to override the system, the curl command and the key. The note aimed at the reviewer did not fool either of them. If your only worry is a skill that tries to exfiltrate something, a plain request to a current model already catches the crude cases.

**Haiku noticed the symptoms and missed the rule.** Given the 65-character name, it called the file safe to install and suggested a shorter name because the long one was hard to type, not because it is over the limit. It called the 1,305-character description duplicated and overloaded without saying that it is longer than allowed. It called the 538-line body padding. None of those answers says that the published limits are broken, which is the fact you need before you upload it. It also let the backslash path through.

**Neither model asked for an example.** On the file with no example of input and output, both called it safe and listed other gaps, such as which PDF library to use. The same omission sank Sonnet on the file with three problems: it found the bad name and the backslash path and said nothing about the missing example.

**With the checklist, the answer is something a script can read.** Each problem came back as one line with a code, the place in the file and a fix, followed by a verdict. That matters if the review runs in a pipeline before publishing: a long essay with a recommendation buried in paragraph four cannot stop a release. Both models also left the bundled-script file alone on both sides.

**One mistake was ours.** Our first wording of the third-person rule was too wide, and the side with the skill treated an ordinary you in a description as a fault, as in one that fills a form from data you give it. The documentation objects to a skill that speaks of itself as I, or addresses the reader as the one doing the job, not to every you. We narrowed the rule to match it and ran the side with the skill again; the first run is kept with the test files. Two of our word patterns for the side without the skill also needed widening, once because a sentence saying that no hidden orders were present counted as an alarm. That change applied to both sides.

## What we did not measure

- **Once per side.** Every file was reviewed once per side; a second run could change a close case.
- **Our own files.** All 20 were short variations of one skill. A real file with scripts, references and several hundred lines may hide things our cases do not have.
- **Word patterns on the side without the skill.** An answer that named a problem in words our patterns do not expect was scored as a miss, and one that mentioned the right word in passing as a hit.
- **What is outside the file.** A skill can ship scripts and download more at run time. A review of the text alone cannot see what those scripts do.
- **Two Claude models.** We did not try models from other vendors.
