Test results
Translate Markdown: Docs That Still Build: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right in 18 of 19 cases: every code block with its comments, link target, reference label, heading id, front matter key and machine value, HTML attribute, HTML comment and template tag came back byte for byte while the prose was translated. It kept a German SQL comment German in an English translation, left a Bulgarian and a Spanish quotation alone in a Bulgarian translation, translated a line that told AI translators to reply OK instead of obeying it, and returned HTML without a fence. It failed one case: in a German README it finished the document, then wrote that it had to correct the table and printed the whole document a second time.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Also 18 of 19, keeping code comments, HTML comments and front matter values where the side without the skill changed them. But it failed one case: for a German document translated into English it put a sentence about the rules it had followed before the document, so the front matter no longer starts the file and a static site would not read it.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Structure intact and translated | 18/19 | 13/19 | 18/19 | 11/19 |
One run per model and case, the same request and the same checks on both sides. Without the skill both models translated the HTML comment addressed to the AI, translated comments inside code blocks (Python, and for Haiku the bash comments of a README), and Sonnet rewrote a German SQL comment in English and changed a link from /de/docs to /en/docs. Sonnet also wrapped both HTML fragments in a code fence; with fences removed it gets 15 of 19 without the skill. Haiku wrapped one whole document, front matter included, in a code fence. With the skill, the one Sonnet loss is a German README it printed twice.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twelve fictional documents written by us, translated into Bulgarian, German, Spanish and English, 19 cases: a README with a badge, a bash block, a table and a reference link; docs with front matter, heading ids, an admonition and an in-page link; an HTML fragment with human-text and machine attributes; MDX with imports and components; a page with an order to AI translators and an HTML comment aimed at the AI; Python and SQL blocks with comments; Liquid templates; GitHub alerts and a nested list; footnotes with reference links; a Bulgarian and a German source; a page that already quotes the target language. The check compares the structure of the answer with the source: code blocks, inline code, link targets, labels, heading levels and ids, HTML tags and attribute values, comments, templates and front matter keys must be identical, the text must be translated, and nothing may be written before or after the document. One run per model and case; a bug in our own check (footnote text read as a link target) was fixed and the recorded answers rescored without new runs.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.
Back to Translate Markdown: Docs That Still Build · Card (JSON)