# Write a SKILL.md your agent uses, and test it first

> What makes an agent pick up a SKILL.md, how to keep its rules checkable, and how to test it on a stronger and a weaker model, with a check that can fail.

Published 2026-09-30 · https://aiskills402.com/blog/skill-md-write-and-test

A SKILL.md gets used when its description matches what the person asked, and it is worth trusting only after you have run it on real inputs and read what came back. Everything else follows from those two facts. The description decides whether the file is loaded at all. The body decides what happens once it is. The test decides whether you know either of those things or just hope.

This guide is our method, in the order we use it. The format facts come from the [public skills documentation](https://code.claude.com/docs/en/skills); the testing advice comes from running our own skills that way.

## 1. Write the description for the moment of choosing

A skill is a folder with a `SKILL.md` inside. The file opens with a block of frontmatter between two lines of three dashes, and the documentation says the `description` field is what Claude uses to decide whether to apply the skill. Two traps: the block counts only when the dashes open line one, and a misspelt key is dropped without a word. Invalid YAML is gentler still, since the skill loads with every field empty. To the author, a typo in a key looks like a skill that never fires.

The same page says that the description and the optional `when_to_use` text are cut off at 1,536 characters in the skill listing, so the main use case goes first. [The skills page](https://code.claude.com/docs/en/skills) also gives a pattern: say what the skill does, then say when to use it, in the words a person would type.

Our rule of thumb, which is ours and not the documentation's: write the description as the sentence the user would say, not the sentence the author would put on a product page. "Use when someone asks to make a text sound less machine-written" matches a request. "A powerful text transformation toolkit" matches nothing anyone types. Add the negative too: what the skill is not for. An agent that picks the wrong skill wastes a turn, and a line about exclusions stops that cheaply.

## 2. Keep the rules short, and only write rules you can check

Every rule in the body should pass one test: could a script tell whether the output obeyed it? "Keep every number and name unchanged" can be checked with a list of strings. "Write in a warm but professional tone" cannot, so it will drift unnoticed. That does not make tone rules forbidden, but it tells you where your testing will be blind and where you will have to read by hand.

Short also means a limit on what you ask for. In our observation a weaker model applies each instruction too hard: ask for short sentences and it chops the text into fragments. Put a ceiling next to each direction ("shorter, but keep at least most of the length") so the rule has a stop.

And write the skill in your own words. We started one skill from an outside file, and our overlap check, which counts shared runs of seven words, found 117 of them. We rewrote the whole wording without reopening the outside text, the check found none afterwards, and the skill passed its five cases again. The lesson is not about licences only. A skill assembled from pieces you did not write contains rules whose reasons you do not know, and you cannot test rules you cannot explain.

## 3. Run it with the tools switched off

If the skill is someone else's, or you are not sure about every line, never test it inside an agent that can touch your machine. Hidden instructions in a file are instructions. Our harness launches the command-line client once per case, with tools off and only local settings loaded. Your personal instructions cannot leak in, and nothing the file demands can be executed. The script has no option to bring tools back, so nobody can switch them on by accident.

For your own skill the same setup has a second benefit: the model sees only the skill and the input. If the output is good, the skill did it, not your global configuration.

## 4. Write cases that can fail

A case pairs an input with mechanical checks. Ours are few: strings to preserve, strings to forbid, an allowed range for output length relative to input, the expected language, a pattern that has to match, and a standing test that some answer came back.

Three habits make the cases useful.

**Plant the errors yourself.** For a proofreading skill we wrote texts with known mistakes, such as "smoother then" or "recieved", and a pattern that fails if any planted wrong form is still in the output. Then we know exactly what a pass means. A text you found and never annotated tells you nothing when the output looks fine.

**Add a keep-list.** Names, figures and phrases that were already right must come out exactly as they went in. That catches a model which "corrects" what needed no correction.

**Include a control that should come back untouched.** One of our cases is an already correct text, and the expected result is no change. Without it a skill that rewrites everything scores perfectly.

Then test the tests. Our harness has a self-test of two cases. One should pass; the other is doomed by design, because it requires a string that cannot occur and several times the input length. Two greens would mean the tester is broken, not the skill. Build your own doomed case before you believe any green. A check nobody has watched go red is decoration.

## 5. Try a strong model and a weak one

Put each case through two models of different strength, at minimum. The strong one tells you whether the instructions are sound. The weak one tells you how much the instructions have to carry on their own, which is what matters if your users run something cheaper.

In our proofreading test the strong model fixed every planted error in all three rounds. The weak one left a planted Bulgarian misspelling of the word for schools in place in all three rounds, and a second one in two of the three. The mechanical score for the weak model on other languages looked fine, so a single average may have hidden it.

## 6. Read the outputs, not only the scores

Checks are mechanical. They do not see an invented claim, a condition rewritten into something else, or a politeness level dropped. After each run we read the texts and wrote a verdict per model in plain words. For one rewriting skill both models passed all five cases, and only the reading showed that the weak model still shortened the text heavily and appended sign-off lines nobody requested, while the strong one added a mild remark of its own in two outputs.

Reading also exposes a faulty check. Our first German proofreading failures came from our own forbid rule, which ignored letter case when it should not have. Until a person has looked at why a case is red, it says nothing about the model.

If a borderline case passes on one run and fails on the next, run it again. The weaker model is not deterministic, and one run per case is an indication, not a measurement.

## 7. Change one thing, then rerun

When a case fails, change the wording or the check, never both at once, and rerun. Keep the old reports next to the new ones so you can see whether a fix made something else worse. Then add a line to the skill's changelog saying what changed and why. Next month you will not remember.

## When this does not apply

This method is tuned for skills that take a text or a small input and return text, where a few mechanical checks say something. It is weaker for skills that run long multi-step work with tools: those need tools on, which we deliberately switch off, and a different kind of test in a sandbox you can throw away.

Our own evidence has limits. The cases are ours and short; we ran one model pair, and the humanize runs fell on two days, 29 and 30 September 2026. Between rounds we fixed our own check (the German forbid rule that ignored letter case), so the later runs are not a clean repeat of the first. The overlap check only finds copied wording, not copied ideas. And a plain reading of outputs by one person is not an independent review.
