Yes, and the expected values are mostly right. Thirteen tasks over ten small functions, each described by a written spec, went to Claude Sonnet and to Claude Haiku, and both were asked for unit tests in one form: a JSON table of test cases. Then we ran every table against our reference code and against the bugs we had planted in copies of it. Without any help, each model got exactly one expected value wrong across all thirteen tasks. What failed more often was the answer as a file: undefined and NaN written straight into the arguments, or a sentence of explanation before or after the table. A test runner that loops over the file stops at both, so those tables count as failures even when every case inside them is sound.

Method

The ten functions are the kind every codebase has: a slug generator, shipping cost by weight band, a parser for durations written like 1h30m, leap years, the range of items on a page, a word counter meant to work for any script, a parser for one CSV line, version comparison, days between two dates, and someone's age on a given date. For each we wrote a short spec, a reference implementation that follows it, and between four and six wrong versions, 61 bugs in all. Every wrong version breaks one sentence of its spec and nothing else: a boundary moved by one, an error that is never raised, a rule applied in the wrong order.

Three of the thirteen tasks also showed code next to the spec. In two of them the code disagrees with the spec, so a model that reads its expected values off the code writes tests the bug passes. In the third, a code comment instructs whoever writes the tests to return none.

A table passes only when three things hold. Every expected value matches what the reference returns. The cases tell every planted bug apart from the reference, which means at least one case gives a different result on the bug; a bug that gives the same results everywhere has survived. And the answer is JSON with nothing around it. Both sides got the same request, and it named the JSON form; one side also had our tests-from-spec skill loaded. Each model answered every task a single time on each side.

Results

With the skill Without it
Sonnet: tables that pass 11 of 13 9 of 13
Haiku: tables that pass 12 of 13 8 of 13

The logic was good on both sides. Without the skill, Sonnet made one arithmetic slip: for a duration with large values in every part it expected 38800 where the spec gives 43000. Haiku counted a span going backwards over a year that is not a leap year as 366 days instead of 365. With the skill, Sonnet wrote 442 cases and none of them expected the wrong thing; Haiku wrote 346 and got one wrong, a CSV field in quotes whose last characters are two quote marks, which it expected to raise an error although the spec allows it.

Most failures were in the wrapping. Without the skill, six answers of twenty-six could not be read as JSON. Haiku wrote undefined or NaN as an argument three times, in the slug maker, the shipping fee task with the disagreeing code, and the CSV parser. In JavaScript that looks harmless; to a JSON parser it is a broken file, and the runner loads nothing. Sonnet once added text after a finished table. With the skill, one answer of twenty-six had this problem: Sonnet appended a correction in prose.

No model copied a bug from the code. On the two tasks where the code shown disagreed with the spec, no table on either side expected what the code does instead of what the spec says. The comment asking for no tests was ignored by every run, and every run wrote a full table. Two of the runs without the skill did announce that they were ignoring it, in a first line above the JSON, and that line is what made their answers fail. Seeing the trick and still breaking the file is a very ordinary way to fail.

One bug survived Sonnet on both sides. The version spec says the result is minus one, zero or one. One planted bug returns the raw difference of the first parts that differ instead, so 2.1 against 2.0 still gives one, and the bug only shows when two parts differ by two or more. None of Sonnet's cases had such a pair, with or without the skill. Haiku's did, on both sides. A rule worded as "returns one of three values" invites one case per value, and that is not enough to see whether the value is computed or merely happens to be right.

What we did not measure

  • Each task ran once. A difference of two to four tables out of thirteen could shrink or grow on a repeat.
  • Our own small functions. Each had a clear written spec. Real specs are vaguer, and the skill's habit of turning gaps into open questions was not scored at all.
  • Most of the gain is the form of the answer. If your runner forgives text around the JSON, or you want the cases written out for your own test framework, the difference may be close to nothing.
  • Our bugs, not every bug. A table that catches all 61 planted bugs can still miss one nobody thought to plant.
  • Our scoring tool had a bug of its own. The first comparison failed every answer on both sides, because it did not hand the folder with the reference code to the check. We fixed it and scored the saved answers again; no model was run twice.
  • Two Claude models. Nothing from other vendors.

Read this post as Markdown: /blog/llm-tests-from-spec-json-first.md · Atom feed.