Test results

Diff From Before and After: A Patch That Applies Exactly: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on 21 of 22 testable file pairs, read by hand, but it missed one: on a Makefile it miscounted a hunk header. The diffs applied with git-style checks for new files, emptied files, blank context lines, a final line break added or removed, and a planted instruction ignored. Two Windows files with carriage returns failed for every model on both sides and are left out; the carriage returns do not survive the trip through the command line we test with.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on all 22 testable file pairs, read by hand, with diffs that applied.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Diffs that apply (22 testable pairs)21/2222/2222/2221/22

Same request on both sides, a fence removed first. Sonnet without the skill wrote a correct diff for all 22 testable pairs; with the skill it miscounted one Makefile hunk, so for Sonnet the skill measured no gain. Haiku without the skill broke one hunk at a blank line inside a planted-instruction case; with it all 22 applied.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-four pairs of before and after files written by us (17 with a trap, 7 plain): tabs in a Makefile, blank context lines, a final newline added or removed, trailing spaces, new and emptied files, Windows line endings and a planted instruction. Each answer is applied to the before file by our patch code and must produce the after file exactly; counts use the verdicts recorded at run time, because re-reading a stored diff loses a trailing blank line. The two Windows cases failed for both models on both sides and are left out of the counts: we suspect the carriage returns are lost on the way to the model. No check was widened. One run per model and pair.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Diff From Before and After: A Patch That Applies Exactly · Card (JSON)