AISKILLS402

Proofread: test results

Tested 2026-09-30, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-09-30
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Fixed all planted errors in en, bg, de, es and ru in every one of three runs; left a correct English text untouched ("Changes: none"), kept all numbers, names and British spelling. Only failure: length limit on one Spanish text, where it correctly flagged a real ambiguity in the test text and the extra line pushed the output to 1.32 of the input.
Weak · claude-haiku-4-5-20251001 (Claude Code alias "haiku")
Good on en, de, es, ru and the correct-text case. Unreliable on Bulgarian: in three runs it missed 'училищята' (should be 'училищата') every time and 'засадание' (should be 'заседание') in two of three. One English run was too verbose (length 1.33). No meaning, number or name was changed by either model.

Note

Three full runs, one run per model and case each time. The first run's German failures were a bug in my own check (case-insensitive forbid), fixed before run two. Sonnet and Haiku differ between runs on borderline cases, so treat single results as indicative. The planted errors per language are listed in test/planted.json.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Proofread · Card (JSON)