Test results

Unified Diff Repair: Fix the Counts, Keep the Change: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on 22 of 23 diffs, read by hand: recounted hunk headers, restored a lost blank context line and a lost tab in a Makefile, put hunks in order, kept the no-newline marker, refused hunks that do not fit the file instead of guessing, ignored a planted instruction, and left the added lines as written. The one it did not get was a Windows file with carriage returns, which no model got right on either side.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on 22 of 23 diffs, read by hand, with the same repairs and refusals as Sonnet, but it missed the same Windows file with carriage returns.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Diffs repaired or refused right (23 diffs)22/2322/2322/2318/23

Same request on both sides, a fence removed first. Sonnet without the skill repaired and refused exactly as well as with it: 22 of 23 both ways, so the skill gives Sonnet nothing measurable here. Haiku without the skill missed four more: it broke a hunk on a blank context line, mishandled an append after a line with no final newline, patched a hunk that is not in the file instead of refusing, and answered one refusal in prose with no diff where a diff was due.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-three broken unified diffs written by us, each with the file it is meant for: wrong hunk counts, lost whitespace and tabs, a missing newline marker, hunks out of order, a context line that is not in the file, overlapping hunks, git headers, a diff inside a chat answer and a planted instruction. Each answer is applied to the file by our own patch code and the result is compared with the expected file; a hunk that cannot be placed must be refused. The counts use the pass or fail recorded at run time: re-reading the stored answers later drops a trailing blank context line and wrongly fails right patches, which we found while checking. The first bare run was cut short when the claude program was replaced by an update mid-run; that run was discarded and repeated in full. The Windows case (carriage returns in the file) failed for both models on both sides; we suspect the carriage returns do not survive the round trip through the command line, so that case says little. No check was widened. One run per model and diff.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Unified Diff Repair: Fix the Counts, Keep the Change · Card (JSON)