# Translating docs with an LLM: the code comments go first

> Claude Sonnet and Haiku translated 19 Markdown and HTML documents for us. The prose was fine; code comments, a hidden HTML comment and a link target were not.

Published 2026-10-08 · https://aiskills402.com/blog/llm-translate-docs-code-comments

Mostly, and the parts it breaks are the ones nobody reads closely. Claude Sonnet and Claude Haiku were each given 19 short documents to translate, from a README to MDX with components, and everything that is not prose was compared with the original, symbol by symbol. Anyone planning to translate Markdown documentation with a model should look first at the same places. Without written rules Sonnet kept the whole structure intact in 13 of the 19 and Haiku in 11. The prose itself was rarely the problem. What changed were comments inside code samples, an HTML comment, one link target, and in a few answers the whole document arrived wrapped in a code fence. When the rules were supplied as a skill, both models got 18 of 19 right, and each of the two misses was text written around the document.

## Method

We wrote twelve fictional documents and asked for Bulgarian, German, Spanish or English, which gave 19 translation jobs. Each document carries the things that tend to come back changed: a bash block with comments, a Python function with a docstring, an SQL query with a German comment, front matter with a slug, a date and tags, heading ids and an in-page link, reference-style links and a footnote, an HTML fragment with class names and data attributes next to alt text and a placeholder, Liquid template tags, MDX imports and components, GitHub alert blocks, and two quotations written in the very language we were translating into. One page also told "AI translators" in plain text to reply only OK, and hid a request to add a link inside an HTML comment.

Both sides got the same one-line request. One side also had our [Markdown translation skill](/skills/translate-markdown) as its instructions. A script extracted every non-prose part from both the original and the answer: code, link destinations, reference names, the heading outline with its ids, the markup with its attribute values, comments, template tags and the front matter keys. Each of those lists had to match exactly. The prose had to change, and nothing could stand before or after the document. Each model translated each job once per side. Our own extractor first read the text of a footnote as if it were a link target; we fixed that and scored the stored answers again without running the models a second time.

## Results

| | Rules loaded | No rules |
|---|---|---|
| Sonnet: nothing broken, text translated | 18 of 19 | 13 of 19 |
| Haiku: nothing broken, text translated | 18 of 19 | 11 of 19 |

**Code comments were translated as if they were prose.** Without the skill, both models turned the English comments and the docstring of the Python example into Spanish, and Haiku did the same to the bash comments in a README. In a German page translated to English, Sonnet rewrote the SQL comment in English. That reads like a kindness and is a defect: the sample no longer matches the code that ships, a diff now shows changes nobody made, and the reader who copies the snippet gets one comment in a language that matches none of the code around it. With the skill, every code block came back unchanged.

**The hidden comment was translated too.** The comment aimed at the AI never shows on the page. Without the skill, both models translated it anyway, which changes a part of the file the author never meant anyone to touch. Neither model added the link it asked for, and both translated the visible "reply only OK" line as ordinary text, on both sides. The difference the skill made was to leave the comment exactly as it was.

**One link moved to a page that may not exist.** Translating a German page into English, Sonnet without the skill changed the link target from `/de/docs/backup` to `/en/docs/backup`. It is a sensible guess about a site's layout and still a guess: if the English page does not exist, the reader lands on an error, and a build that checks links fails. With the skill, link targets stayed as written.

**Fences around the answer.** Sonnet without the skill put each HTML fragment inside a code fence, and Haiku wrapped a whole Markdown document, front matter included, the same way. A fenced answer pasted into a page renders as a grey box of source code. Remove the fence and Sonnet's score without the skill rises to 15 of 19.

**What went wrong with the skill.** Sonnet finished a German README, then announced that it needed to correct a table and printed the entire document a second time. Haiku opened a German-to-English job with a sentence describing the rules it had followed, so the front matter was no longer the first thing in the file, which is where static site generators look for it. Both are text around the document, the one thing the instructions forbid most plainly.

## Before you run a translation pass over your docs

- **Diff everything that is not prose.** Extract the code blocks, link targets and front matter keys from the source and the translation and compare the lists. It takes a few lines of script and catches every defect above.
- **Check the first line.** If the translation does not start where the source starts, something was written in front of it.
- **Decide your link policy before you translate.** If localized pages live under a language prefix, rewrite the links in a separate, deliberate step, not as a side effect of translation.
- **Leave comments to the people who own the code.** Translating code comments is a code change and belongs in a code review.

## What we did not measure

- **One pass each.** Every answer was generated once per side; another attempt could flip a close case.
- **Twelve short documents written by us.** Real documentation is longer, with nested components and custom syntax we did not try.
- **The quality of the prose.** We measured what must not change. Whether the translated sentences read well was not scored.
- **Anchors that depend on heading text.** A link to an anchor generated from a heading breaks when the heading is translated. Our cases kept explicit heading ids, and the skill does not repair such links.
- **Two Claude models, four languages.** No other vendors, no right-to-left scripts.
