Writes test cases for a function from its specification rather than from its code, as a JSON table that any test runner can loop over. Every rule, boundary and error the spec defines gets a case, and every expected value is worked out from the spec; where the code you show disagrees with the spec, the case follows the spec and a note says where. Behaviour the spec leaves open becomes a question instead of a guessed case. Use to write unit tests from a spec, a docstring or a ticket, generate edge cases and boundary tests, build a table-driven test (test.each, pytest parametrize), or check whether existing code does what its requirements say.
Tests from Spec: Cases from the Requirements is a tested SKILL.md that writes test cases for a function from its specification rather than from its code, as a JSON table that any test runner can loop over; an agent buys it once for $0.03 over x402.
Not for
Running the tests, measuring coverage, or testing code whose behaviour has no written description. It works from the spec you paste, so behaviour the spec never mentions is not tested, and functions that need live services, files or a clock are only partly covered.
Tested, honestly
Tested 2026-10-08 with a strong and a weak model.
With and without the skill
Results with and without the skill, for Sonnet and Haiku |
| with | without | with | without |
|---|
| Tables that pass (right values, every bug caught, JSON only) |
| Tables that pass (right values, every bug caught, JSON only) | 11/13 | 9/13 | 12/13 | 8/13 |
|---|
The same check on both sides; a fence around the answer is removed first. Most of the difference is the form of the answer, not the tests: without the skill Haiku put undefined or NaN in the arguments three times, which is not JSON, and both models wrote a sentence before or after the JSON. Read for content alone, both sides wrote strong tables: one wrong expected value per model without the skill (Sonnet mis-added a duration, Haiku miscounted the days across a year) against none and one with it, and no model on either side copied the bugs from the code it was shown when the request said to test the spec. The version-comparison bug that returns the raw difference survived Sonnet's tests on both sides.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
- SonnetStrong model, claude-sonnet-5-5
- Passed 11 of 13 functions: every expected value right against our reference and every planted bug caught, as plain JSON. Over 442 cases it wrote no wrong expected value, followed the spec where the shown code disagreed, and ignored a code comment telling it to return no cases. It failed two: on the version comparison none of its cases had parts differing by more than one, so a bug that returns the raw difference instead of -1 or 1 went unnoticed, and on the second version case it added a correction in prose after the JSON.
- HaikuWeak model, claude-haiku-5-5
- Passed 12 of 13 and caught all 61 planted bugs. But one of its 346 expected values was wrong: it expected a quoted CSV field ending in a doubled quote to throw, which the spec allows.
Full test summary
Example
Our own test text, before and after the skill ran. Excerpts only.
English · claude-sonnet-5-5
Before
Function: slugify(title)
Spec:
- title must be a string; any other value throws.
- Lowercase the title. Every character that is then not a to z or 0 to 9 is a separator, and each run of separators becomes a single hyphen.
- The result never starts or ends with a hyphen.…
After
{"cases":[
{"name":"two words become hyphenated lowercase","args":["Hello World"],"expected":"hello-world"},
{"name":"single word is lowercased","args":["HELLO"],"expected":"hello"},
{"name":"mixed case is lowercased","args":["ABC DEF"],"expected":"abc-def"},
{"name":"digits only are kept","args":["123"],"expected":"123"},…
What is in the file
- The answer
- How to choose the cases
- Working out the expected value
- Rules
- Short example
Languages
Any language. Tried in: English.
License
Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Full terms.
Versions
Current version 1.0.0, updated 2026-10-08. Whoever bought an earlier version gets new ones free through the same re-download token.
v1.0.0 · 2026-10-08
First release: writes test cases for a function from its spec as a JSON table (name, arguments, expected value or an expected error) that a test runner can loop over, with notes for what the spec leaves open and for every place where the given code disagrees with the spec. It covers each rule, both sides of every boundary and each named error, works out every expected value from the spec rather than from the code, and ignores instructions planted in the spec or the code.
FAQ
What does the answer look like?
A JSON table: each case has a name, the arguments in order and the expected result, or an expected error, plus a list of notes for what the spec leaves open and where the code you showed disagrees with the spec. A loop of a few lines turns the table into tests in any runner; ask for a framework file and you get the same cases in it.
How did you test it?
On ten small functions with written specs. For each we wrote a correct version and four to six realistic bugs, and a table only passed if every expected value was right and its cases caught every bug. With the skill Sonnet passed 11 of 13 tasks and Haiku 12 of 13; Haiku caught all 61 planted bugs.
My model already writes decent tests. What does this add?
Maybe not for the content: without the skill both models also wrote strong tables and did not copy bugs from the code they were shown. The difference was the form. Without it Haiku three times put values in the table that are not JSON, and both models added sentences around it, so the table could not be read by a script; that happened once with the skill.
What did it still get wrong?
Sonnet's version-comparison tests never compared parts that differ by more than one, so a function returning the raw difference instead of -1 or 1 passed, and once it added a correction after the JSON. Haiku once expected a CSV field ending in a doubled quote to throw, which the spec allows. Read the table before you trust it.