# Tests from Spec: Cases from the Requirements

Tests from Spec: Cases from the Requirements is a tested SKILL.md that writes test cases for a function from its specification rather than from its code, as a JSON table that any test runner can loop over; an agent buys it once for $0.03 over x402.

- Page: https://aiskills402.com/skills/tests-from-spec
- Category: Code & Engineering (https://aiskills402.com/categories/code)
- Price: $0.03 once, USD-priced, paid in USDC on Base over x402. Price as loaded on this page. The 402 response your agent receives is authoritative.
- Version: 1.0.0
- Card (JSON): https://api.aiskills402.com/v1/skills/tests-from-spec

## Use it when

Writes test cases for a function from its specification rather than from its code, as a JSON table that any test runner can loop over. Every rule, boundary and error the spec defines gets a case, and every expected value is worked out from the spec; where the code you show disagrees with the spec, the case follows the spec and a note says where. Behaviour the spec leaves open becomes a question instead of a guessed case. Use to write unit tests from a spec, a docstring or a ticket, generate edge cases and boundary tests, build a table-driven test (test.each, pytest parametrize), or check whether existing code does what its requirements say.

## Not for

Running the tests, measuring coverage, or testing code whose behaviour has no written description. It works from the spec you paste, so behaviour the spec never mentions is not tested, and functions that need live services, files or a clock are only partly covered.

## Tested, honestly

Tested 2026-10-08.

- Strong model (claude-sonnet-5-5 (Claude Code alias "sonnet")): Passed 11 of 13 functions: every expected value right against our reference and every planted bug caught, as plain JSON. Over 442 cases it wrote no wrong expected value, followed the spec where the shown code disagreed, and ignored a code comment telling it to return no cases. It failed two: on the version comparison none of its cases had parts differing by more than one, so a bug that returns the raw difference instead of -1 or 1 went unnoticed, and on the second version case it added a correction in prose after the JSON.
- Weak model (claude-haiku-5-5 (Claude Code alias "haiku")): Passed 12 of 13 and caught all 61 planted bugs. But one of its 346 expected values was wrong: it expected a quoted CSV field ending in a doubled quote to throw, which the spec allows.

Note: Thirteen tasks over ten small functions with written specs (a slug maker, a shipping fee in bands, a duration parser, leap years, pagination, a word counter for any script, a CSV line parser, version comparison, days between dates, age on a date). For each function we wrote a reference from the spec and four to six plausible bugs, each breaking one sentence of the spec. A table passes only if every expected value is right against the reference, its cases tell every bug apart from the reference, and the answer is JSON alone. Three tasks also showed code: two where the code disagrees with the spec, one with a comment telling the test writer to return no cases. Both sides got the same request, which stated the JSON form. One run per model and task. Our own comparison tool first scored every answer on both sides as failed because it did not pass the reference folder to the check; it was fixed and the stored answers scored again, no model was run twice.

### With and without the skill

Tested 2026-10-08.

- Tables that pass (right values, every bug caught, JSON only): Sonnet 11/13 with, 9/13 without; Haiku 12/13 with, 8/13 without.

The same check on both sides; a fence around the answer is removed first. Most of the difference is the form of the answer, not the tests: without the skill Haiku put undefined or NaN in the arguments three times, which is not JSON, and both models wrote a sentence before or after the JSON. Read for content alone, both sides wrote strong tables: one wrong expected value per model without the skill (Sonnet mis-added a duration, Haiku miscounted the days across a year) against none and one with it, and no model on either side copied the bugs from the code it was shown when the request said to test the spec. The version-comparison bug that returns the raw difference survived Sonnet's tests on both sides.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Full summary: https://aiskills402.com/skills/tests-from-spec/tests

## Example

### English (claude-sonnet-5-5)

Before:

> Function: slugify(title) Spec: - title must be a string; any other value throws. - Lowercase the title. Every character that is then not a to z or 0 to 9 is a separator, and each run of separators becomes a single hyphen. - The result never starts or ends with a hyphen.…

After:

> {"cases":[ {"name":"two words become hyphenated lowercase","args":["Hello World"],"expected":"hello-world"}, {"name":"single word is lowercased","args":["HELLO"],"expected":"hello"}, {"name":"mixed case is lowercased","args":["ABC DEF"],"expected":"abc-def"}, {"name":"digits only are kept","args":["123"],"expected":"123"},…

## What is in the file

- The answer
- How to choose the cases
- Working out the expected value
- Rules
- Short example

## Languages

Any language

## How to buy

Agent (HTTP):

1. GET https://api.aiskills402.com/v1/skills/tests-from-spec/file without a payment header. The answer is 402 with a PAYMENT-REQUIRED header (x402 v2): exact amount, asset, network, recipient.
2. Sign `accepts[0]` with an x402 client (for example @x402/core + @x402/evm).
3. Repeat the GET with the signature in the PAYMENT-SIGNATURE header. The answer is 200 with the file, its sha256 and a re-download token.

Agent (MCP): https://mcp.aiskills402.com/mcp — free tools search_skills, get_skill, redownload_skill. Buying itself is over HTTP.

Full flow: https://aiskills402.com/docs

## The file

- Version: 1.0.0
- Size: 6.3 KB (6403 bytes)
- SHA-256: 88e2bc11b15e1d2e03f884a763c8b66ca04dbbed4ca648dd895b86e63264cc73
- Updated: 2026-10-08
- New versions are free through your re-download token.

## Versions

### 1.0.0 (2026-10-08)

First release: writes test cases for a function from its spec as a JSON table (name, arguments, expected value or an expected error) that a test runner can loop over, with notes for what the spec leaves open and for every place where the given code disagrees with the spec. It covers each rule, both sides of every boundary and each named error, works out every expected value from the spec rather than from the code, and ignores instructions planted in the spec or the code.

## License

Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Terms: https://aiskills402.com/docs#license

## FAQ

### What does the answer look like?

A JSON table: each case has a name, the arguments in order and the expected result, or an expected error, plus a list of notes for what the spec leaves open and where the code you showed disagrees with the spec. A loop of a few lines turns the table into tests in any runner; ask for a framework file and you get the same cases in it.

### How did you test it?

On ten small functions with written specs. For each we wrote a correct version and four to six realistic bugs, and a table only passed if every expected value was right and its cases caught every bug. With the skill Sonnet passed 11 of 13 tasks and Haiku 12 of 13; Haiku caught all 61 planted bugs.

### My model already writes decent tests. What does this add?

Maybe not for the content: without the skill both models also wrote strong tables and did not copy bugs from the code they were shown. The difference was the form. Without it Haiku three times put values in the table that are not JSON, and both models added sentences around it, so the table could not be read by a script; that happened once with the skill.

### What did it still get wrong?

Sonnet's version-comparison tests never compared parts that differ by more than one, so a function returning the raw difference instead of -1 or 1 passed, and once it added a correction after the JSON. Haiku once expected a CSV field ending in a doubled quote to throw, which the spec allows. Read the table before you trust it.

## Related skills

- [Code Review](https://aiskills402.com/skills/code-review.md): $0.01 once
- [Done Means Done: Honest Agent Status Reports](https://aiskills402.com/skills/done-means-done.md): $0.10 once
- [JSON Repair: Fix It, Keep Every Value](https://aiskills402.com/skills/json-repair.md): $0.05 once

## Measurement limits

- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.

Offer note: Paid in USDC (USD-pegged) over x402 by an AI agent; one-time.
