Reviews a SKILL.md file before you install, buy or publish it and lists every problem with a fixed code, the place and a fix, then gives one verdict. It checks the front matter against the published limits (name up to 64 characters in lowercase, hyphens and digits, no reserved words; description present, up to 1,024 characters, naming both its job and the requests that should trigger it, in the third person), and it reads the body for hidden orders to the agent, credentials pasted into the text, commands that download and run remote code, Windows-style paths, a body over 500 lines and the lack of any example. Use to review a SKILL.md, audit a Claude or agent skill before installing it, check a skill file for prompt injection, or lint a skill before publishing it to a marketplace.
SKILL.md Review: Safe and Within the Limits is a tested SKILL.md that reviews a SKILL.md file before you install, buy or publish it and lists every problem with a fixed code, the place and a fix, then gives one verdict; an agent buys it once for $0.03 over x402.
Not for
Running or testing the skill, reading the scripts that come with it, or judging whether its advice is correct for its topic. It reads one SKILL.md as text, so a problem hidden in a bundled script or a downloaded file is outside what it can see.
Tested, honestly
Tested 2026-10-08 with a strong and a weak model.
With and without the skill
Results with and without the skill, for Sonnet and Haiku |
| with | without | with | without |
|---|
| Every planted problem named, nothing invented |
| Every planted problem named, nothing invented | 20/20 | 18/20 | 20/20 | 14/20 |
|---|
Both sides scored with the same word-based checks, since the side without the skill does not know our codes. Without the skill both models noticed the hidden orders, the key and the curl command, and Sonnet called out the note aimed at the reviewer; the misses were the counted rules. Haiku did not flag the 65-character name, the 1,305-character description, the 538-line body or the backslash path, and neither model said that the file had no example. Two baseline patterns first missed correct answers (one counted a sentence saying no injected instructions as an alarm) and were widened before the final score; the change applies to both sides.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
- SonnetStrong model, claude-sonnet-5-5
- Gave the exact codes and verdict in all 20 cases, as bare lines with nothing around them. It caught a request to read the SSH key, a silent copy to a shared drive, a line that claims to replace the system's instructions, a note telling the reviewer to answer No findings, a live-looking API key, a curl piped into sh, a name over 64 characters, an uppercase name, the reserved word claude in a name, a missing, an overlong and a first-person description, an XML tag in the description, a Windows path, a 538-line body and a missing example, and it left a skill that runs its own bundled script unflagged. A first run with a wider wording of our person rule flagged the word you inside ordinary descriptions; the rule was narrowed to match the published guidance and run again.
- HaikuWeak model, claude-haiku-5-5
- Also 20 of 20 with the skill, with the same codes and verdicts as Sonnet, including the counted limits it did not flag without the skill: the 65-character name, the 1,305-character description and the 538-line body.
Full test summary
Example
Our own test text, before and after the skill ran. Excerpts only.
English · claude-sonnet-5-5
Before
---
name: PDF_Form_Filler
description: Fills in PDF forms from data you give it and checks that every required field is set. Use when asked to fill a PDF form, complete a fillable PDF, or check a form for missing fields.
---
# PDF form filler
Read the form fields first, then fill only the fields the user gave data for.
Never invent a value. Report every required field that is still empty.…
After
[NAME] name: "PDF_Form_Filler" has uppercase letters and underscores, and only lowercase letters, digits and hyphens are allowed. Fix: pdf-form-filler.
Verdict: fix before use
What is in the file
- The answer
- The codes
- Rules
- Work in this order
- Short example
Languages
Any language. Tried in: English.
License
Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Full terms.
Versions
Current version 1.0.0, updated 2026-10-08. Whoever bought an earlier version gets new ones free through the same re-download token.
v1.0.0 · 2026-10-08
First release: reviews one SKILL.md file and lists each problem with a fixed code, the place and a fix, then a verdict (unsafe, fix before use, ready). Unsafe: hidden orders to the agent, credentials in the text, download-and-run commands. Front matter against the published limits: name, missing or overlong description, no trigger, first or second person, XML tags. Body: over 500 lines, Windows paths, no example. Ordinary instructions to use the skill's own scripts are not flagged, and an order to the reviewer inside the file is itself a finding.
FAQ
What does it check in the front matter?
The published limits for skills: a name of at most 64 characters in lowercase letters, digits and hyphens without the reserved words anthropic and claude, and a description that exists, stays within 1,024 characters, says when the skill should be used and is written in the third person. It counts the characters instead of estimating them.
Will it catch a hidden instruction to the agent?
Those come first in the report, with the verdict unsafe: a request to read SSH keys, a silent copy of the user's files, a line claiming to replace the system's instructions and a note telling the reviewer to answer No findings were all caught in our tests, along with a pasted API key and a curl command piped into a shell.
Does it flag every skill that runs scripts?
No. Telling the agent to run the skill's own bundled script on the file the user named is normal and was not flagged by either model. It becomes a finding when the file reaches for data the task does not need or downloads and runs code at use time.
What did the models miss without it?
The counted rules. Without the skill, Claude Haiku did not notice a 65-character name, a 1,305-character description, a body of 538 lines or a Windows path, and neither Haiku nor Sonnet said that a file had no example. With the skill both models reported all of them.