Test results

Tests from Spec: Cases from the Requirements: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Passed 11 of 13 functions: every expected value right against our reference and every planted bug caught, as plain JSON. Over 442 cases it wrote no wrong expected value, followed the spec where the shown code disagreed, and ignored a code comment telling it to return no cases. It failed two: on the version comparison none of its cases had parts differing by more than one, so a bug that returns the raw difference instead of -1 or 1 went unnoticed, and on the second version case it added a correction in prose after the JSON.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Passed 12 of 13 and caught all 61 planted bugs. But one of its 346 expected values was wrong: it expected a quoted CSV field ending in a doubled quote to throw, which the spec allows.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Tables that pass (right values, every bug caught, JSON only)11/139/1312/138/13

The same check on both sides; a fence around the answer is removed first. Most of the difference is the form of the answer, not the tests: without the skill Haiku put undefined or NaN in the arguments three times, which is not JSON, and both models wrote a sentence before or after the JSON. Read for content alone, both sides wrote strong tables: one wrong expected value per model without the skill (Sonnet mis-added a duration, Haiku miscounted the days across a year) against none and one with it, and no model on either side copied the bugs from the code it was shown when the request said to test the spec. The version-comparison bug that returns the raw difference survived Sonnet's tests on both sides.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Thirteen tasks over ten small functions with written specs (a slug maker, a shipping fee in bands, a duration parser, leap years, pagination, a word counter for any script, a CSV line parser, version comparison, days between dates, age on a date). For each function we wrote a reference from the spec and four to six plausible bugs, each breaking one sentence of the spec. A table passes only if every expected value is right against the reference, its cases tell every bug apart from the reference, and the answer is JSON alone. Three tasks also showed code: two where the code disagrees with the spec, one with a comment telling the test writer to return no cases. Both sides got the same request, which stated the JSON form. One run per model and task. Our own comparison tool first scored every answer on both sides as failed because it did not pass the reference folder to the check; it was fixed and the stored answers scored again, no model was run twice.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Tests from Spec: Cases from the Requirements · Card (JSON)