Code & Engineering

Tests from Spec: Cases from the Requirements

Writes test cases for a function from its specification rather than from its code, as a JSON table that any test runner can loop over. Every rule, boundary and error the spec defines gets a case, and every expected value is worked out from the spec; where the code you show disagrees with the spec, the case follows the spec and a note says where. Behaviour the spec leaves open becomes a question instead of a guessed case. Use to write unit tests from a spec, a docstring or a ticket, generate edge cases and boundary tests, build a table-driven test (test.each, pytest parametrize), or check whether existing code does what its requirements say.

Tests from Spec: Cases from the Requirements is a tested SKILL.md that writes test cases for a function from its specification rather than from its code, as a JSON table that any test runner can loop over; an agent buys it once for $0.03 over x402.

Tested 2026-10-08No code, no hidden instructionsv1.0.0 · 6.3 KB · perpetual license

Not for

Running the tests, measuring coverage, or testing code whose behaviour has no written description. It works from the spec you paste, so behaviour the spec never mentions is not tested, and functions that need live services, files or a clock are only partly covered.

Tested, honestly

Tested 2026-10-08 with a strong and a weak model.

With and without the skill

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Tables that pass (right values, every bug caught, JSON only)11/139/1312/138/13

The same check on both sides; a fence around the answer is removed first. Most of the difference is the form of the answer, not the tests: without the skill Haiku put undefined or NaN in the arguments three times, which is not JSON, and both models wrote a sentence before or after the JSON. Read for content alone, both sides wrote strong tables: one wrong expected value per model without the skill (Sonnet mis-added a duration, Haiku miscounted the days across a year) against none and one with it, and no model on either side copied the bugs from the code it was shown when the request said to test the spec. The version-comparison bug that returns the raw difference survived Sonnet's tests on both sides.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

SonnetStrong model, claude-sonnet-5-5
Passed 11 of 13 functions: every expected value right against our reference and every planted bug caught, as plain JSON. Over 442 cases it wrote no wrong expected value, followed the spec where the shown code disagreed, and ignored a code comment telling it to return no cases. It failed two: on the version comparison none of its cases had parts differing by more than one, so a bug that returns the raw difference instead of -1 or 1 went unnoticed, and on the second version case it added a correction in prose after the JSON.
HaikuWeak model, claude-haiku-5-5
Passed 12 of 13 and caught all 61 planted bugs. But one of its 346 expected values was wrong: it expected a quoted CSV field ending in a doubled quote to throw, which the spec allows.

Full test summary

Example

Our own test text, before and after the skill ran. Excerpts only.

English · claude-sonnet-5-5

Before

Function: slugify(title) Spec: - title must be a string; any other value throws. - Lowercase the title. Every character that is then not a to z or 0 to 9 is a separator, and each run of separators becomes a single hyphen. - The result never starts or ends with a hyphen.…

After

{"cases":[ {"name":"two words become hyphenated lowercase","args":["Hello World"],"expected":"hello-world"}, {"name":"single word is lowercased","args":["HELLO"],"expected":"hello"}, {"name":"mixed case is lowercased","args":["ABC DEF"],"expected":"abc-def"}, {"name":"digits only are kept","args":["123"],"expected":"123"},…

What is in the file

  • The answer
  • How to choose the cases
  • Working out the expected value
  • Rules
  • Short example

Languages

Any language. Tried in: English.

License

Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Full terms.

Versions

Current version 1.0.0, updated 2026-10-08. Whoever bought an earlier version gets new ones free through the same re-download token.

  1. v1.0.0 · 2026-10-08

    First release: writes test cases for a function from its spec as a JSON table (name, arguments, expected value or an expected error) that a test runner can loop over, with notes for what the spec leaves open and for every place where the given code disagrees with the spec. It covers each rule, both sides of every boundary and each named error, works out every expected value from the spec rather than from the code, and ignores instructions planted in the spec or the code.

FAQ

What does the answer look like?

A JSON table: each case has a name, the arguments in order and the expected result, or an expected error, plus a list of notes for what the spec leaves open and where the code you showed disagrees with the spec. A loop of a few lines turns the table into tests in any runner; ask for a framework file and you get the same cases in it.

How did you test it?

On ten small functions with written specs. For each we wrote a correct version and four to six realistic bugs, and a table only passed if every expected value was right and its cases caught every bug. With the skill Sonnet passed 11 of 13 tasks and Haiku 12 of 13; Haiku caught all 61 planted bugs.

My model already writes decent tests. What does this add?

Maybe not for the content: without the skill both models also wrote strong tables and did not copy bugs from the code they were shown. The difference was the form. Without it Haiku three times put values in the table that are not JSON, and both models added sentences around it, so the table could not be read by a script; that happened once with the skill.

What did it still get wrong?

Sonnet's version-comparison tests never compared parts that differ by more than one, so a function returning the raw difference instead of -1 or 1 passed, and once it added a correction after the JSON. Haiku once expected a CSV field ending in a doubled quote to throw, which the spec allows. Read the table before you trust it.

Read more

Share

Read this page as Markdown: /skills/tests-from-spec.md.

  • Code Review

    Code & Engineering

    SKILL.md · v1.0.4 · 6.0 KB

    Reviews a code change (a diff or a changed file pasted as text) and reports real problems ranked by severity — correctness bugs, security holes, data loss, missing error handling, edge cases, resource leaks — each with location, reason and a concrete fix in words. Use when asked to review code, a diff or a pull request, check a change for bugs, or find security problems in a snippet.

    $0.01once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 30 Sep 2026

  • Recommended

    Done Means Done: Honest Agent Status Reports

    Agents & Protocols

    SKILL.md · v1.0.2 · 9.9 KB

    Stops an AI agent from reporting false success, or hallucinated completion: work it calls done that did not happen. Every action in its status report gets one of five states (done and verified, done but not confirmed, partly done, failed, not done), backed by what the tool results of the session show. Failures and unknowns come first, counts come from the results, and no id, receipt or hash is invented. When a result does show success, the agent says so plainly. Use whenever an agent reports on actions it took with tools, such as messages, batches, tests, builds, deploys, file edits, data updates, payments or API calls.

    $0.05once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 3 Oct 2026

  • JSON Repair: Fix It, Keep Every Value

    Data & Analysis

    SKILL.md · v1.0.0 · 6.9 KB

    Repairs broken JSON so a program can parse it, without changing any value. Fixes trailing and missing commas, comments, single quotes, unquoted keys, Python and JavaScript literals, smart quotes used as delimiters, raw line breaks and stray quotes inside strings, a code fence or chat text around the JSON, and output cut off in the middle. A value that was cut off becomes null instead of a guess, numbers JSON cannot hold as written are kept as strings, and when the structure can be read two ways the answer says so instead of picking one. Use when a tool, an API or another model returned JSON that does not parse, or when asked to fix, clean up, validate or close invalid or truncated JSON.

    $0.01once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 7 Oct 2026