Test results
YAML Repair: Fix It, Keep Every Value: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right on all 22 cases, answering with bare YAML every time: tabs became spaces, a cut-off summary was closed with its quote and kept as cut, a value with a colon inside and a pattern starting with a star got quotes, 1,250 stayed a string, NO and yes stayed strings, the on: key of a workflow stayed on, 1.10 stayed unquoted, and for a YAML 1.1 reader NO and off were quoted. Front matter came back without the dash lines or the Markdown, an HTML error page and a plain error message each got the single line # not yaml, and valid input kept its values.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Right on 21 of 22 cases, but not on the one with a note addressed to the AI: it left the value, which contains a colon and a space, without quotes, so the file still does not parse. One more answer, the YAML 1.1 case, came in a code fence, which a program cannot read as it is. Everything else matched Sonnet: tabs, cut-off values, 1,250, NO, the refusals and the front matter.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Values right (code fence ignored) | 22/22 | 16/22 | 21/22 | 16/22 |
| Right and parses as it is | 22/22 | 0/22 | 20/22 | 9/22 |
The same request on both sides, asking for the YAML only. Without the skill the answers were mostly in a code fence (Sonnet 21 of 22, Haiku 12 of 22), which is why the second row is so low; the first row removes the fence. Read by hand, Sonnet without the skill made real errors: 1,250 became 1250, debug: off was left unquoted for a YAML 1.1 reader, 1.10 was put in quotes (a number became a string), draft: no became false in front matter, and an HTML error page and a plain error message were turned into invented YAML maps (error codes, retry times) instead of a refusal. Refusals in any wording count as right on the plain side.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-two cases written by us, one in Bulgarian: tabs, bad indentation, a quote or bracket that never closes, values that need quotes, a cut-off value, YAML 1.1 readers, front matter in Markdown, an HTML page and an error message instead of YAML, and seven valid inputs that must keep their values. The check parses the answer with the yaml package (YAML 1.2) and compares the whole document, types included, in any layout; the refusal cases compare one fixed line. The same request is sent with and without the skill. One run per model and case.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.