Test results

XML Repair: Fix It, Change No Text: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on all 24 documents, read by hand: it escaped a raw ampersand in text and in a link value, mapped nbsp to a numeric reference without a second escape, quoted bare attribute values, closed the one missing end tag, moved a misplaced declaration to the first byte and kept the comment, answered with bare XML instead of the fenced chat reply, left a planted comment and a sentence that talks to the AI as plain text, and returned the six already-valid files untouched. It gave the one-line CANNOT REPAIR answer for overlapping tags, a cut-off feed, an unknown entity name, an HTML page, two roots, a repeated attribute and the items with no end tags.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on all 24 documents by the skill's own checks, read by hand. Every repair kept the text, the whitespace and the CDATA as written, the declaration was moved and not added, and the seven risky documents (overlap, cut-off, unknown entity, HTML page, two roots, repeated attribute, items without end tags) each got the fixed refusal line instead of a repair.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Documents handled right (24 documents)24/2417/2424/2418/24

Same request on both sides; the bare side is scored on content, fence removed first. Read by hand, Sonnet without the skill did every plain repair and closed the two unclosed items as siblings without touching the text, which we accept. It missed the cases where the right answer is to stop: it re-nested overlapping tags, completed a cut-off title and feed, wrote the unknown entity as literal text, turned an HTML page into XHTML, wrapped two roots in an invented element and dropped one of two repeated attributes. It also rewrote the quotes and end tag of an already-valid SVG. Haiku without the skill showed the same six misses; its HTML conversion also moved elements. With the skill both models refused those seven cases with the one line and left valid files unchanged.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-four documents written by us (16 traps, 8 controls): seven that must be refused, nine that need a repair, and eight controls (six already valid), with Bulgarian element names and an Atom feed with CDATA. Each answer is scored by code: a strict XML parser (errors and warnings fail), the text and attribute values compared node by node, kept comments and CDATA, no fence, and the exact CANNOT REPAIR line for refusals. Grammar facts re-checked on 2026-10-08 in the W3C XML 1.0 Recommendation (fifth edition); nothing here is an owner-measured platform fact. One check was widened after the run for both sides: for two items with no end tags, closing them as siblings with every character of text unchanged is accepted as well as the refusal (a nested reading or any changed text still fails). One run per model and document; the bare side is scored with a fence removed first.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to XML Repair: Fix It, Change No Text · Card (JSON)