Mostly because the pasted text outweighs the request. When someone writes one short line in Bulgarian and pastes a long English build log under it, the model reads far more English than Bulgarian, and when it sums up the paste, the summary often comes out in whatever language the paste was written in. A neighbouring language makes it worse: German next to Dutch, Portuguese next to Spanish, Russian next to Bulgarian. Across 24 such requests, Claude Sonnet answered in the wrong language three times when nobody told it whose language counts, and not once after a short rule said so.
Why we cared
We work in Bulgarian. For weeks, replies from our coding agent carried stray Russian words and letters into Bulgarian sentences, and now and then a whole English line. Each slip is small. Reading them all day is tiring, and a text for a client with one Russian word in it looks careless. The two languages share most of their alphabet, so the slip is easy to miss on a quick read, and hard for a model to notice on its own.
So we did two things. We measured how often a model drifts away from the person and toward the pasted material, and we wrote a check that reads every finished reply before a person sees it.
Method
We wrote 24 single-turn cases. Every case pairs a short question written in one language with a longer paste written in another, most often a close relative:
- a Bulgarian question over three kinds of paste: English build output, a post from a Russian forum, and an error message in Russian;
- a Ukrainian question on a Russian passage, a Polish one on Czech, a Czech one on Slovak;
- Spanish and Portuguese in both directions, Italian about Spanish, German about Dutch and about English;
- Russian about Bulgarian, and Turkish, French and Japanese about English.
A few cases test something else. A line hidden inside the paste asks the model to reply in English. One user explicitly asks for an email in English. A Bulgarian user types in Latin letters. Three controls use a single language throughout.
Each case ran once per model, Claude Sonnet and Claude Haiku, through Claude Code with no tools and no personal instructions: first the bare request, then the same request with the rule added. Code judged every answer on two things: the detected language, and the absence of letters and words that belong to the other language. After the first run we widened two checks, for both sides, each time after a correct answer had been refused: when we asked for an English email, a line of explanation to the user in their own language and a Subject header are now allowed, and the one-sentence controls may run a little longer.
Results
| model | without the rule | with the rule |
|---|---|---|
| Claude Sonnet | 21 of 24 | 24 of 24 |
| Claude Haiku | 18 of 24 | 21 of 24 |
Reading the misses by hand told us more than the totals. Without the rule, Sonnet replied in Dutch when the question was German, in Spanish when it was Portuguese, and in Bulgarian when it was Russian. All three were summaries, and each followed the text being summarised rather than the person asking. It also replied in Latin letters to the Bulgarian user who had typed in Latin letters; our check accepts either script for that case, so it is not counted as a miss.
With the rule, Haiku still slipped three times, always toward Bulgarian: once for a question in Ukrainian over Russian material, and twice for questions in Russian and in English about Bulgarian material. With the rule, both models ignored the line hidden in the paste that asked for English, and both wrote the English email when the user asked for one.
The rule we tested is packaged as reply-language-lock. What it changed for Sonnet was narrow and specific: the three summaries that had followed the paste now followed the person.
A second net: checking the reply after it is written
A rule in the prompt lowers the rate. It does not make it zero, and in a long session full of tool output the paste is often much longer than in our cases. So we also check every finished reply with a Stop hook. Claude Code runs it when Claude finishes a turn; if the hook answers with a block and a reason, Claude does not stop, reads the reason and continues, in our case with a short correction (Claude Code hooks reference).
What the check looks for in a Bulgarian reply, and what each part taught us:
- Letters Bulgarian does not have. Russian
ы,эandё; Ukrainianі,їandє; Serbian and Macedonianј,љandњ. None of them occurs in Bulgarian, so each is a clean signal. The soft sign is the exception: Bulgarian usesьonly beforeо, as inшофьор, so the check flags it anywhere else. - Russian verb endings Bulgarian lacks, such as
-ает,-етсяand-ются. This rule gave our only false alarms on real Bulgarian text: three ordinary nouns,силует,пируетandменует, end the same way, so they are now exceptions, each as a whole word. - A short list of Russian function words, matched as whole words. Here JavaScript bit us. The
\bboundary counts only Latin letters, digits and the underscore as word characters, so a pattern like\bуже\bnever matches Cyrillic text at all (MDN on the word boundary assertion). The boundary has to be written as a look for letters on either side, with\p{L}. A check that silently matches nothing looks exactly like a clean result. - English lines. A line with no Cyrillic and enough English function words to read as a sentence is flagged; a list of file names is not. Trailing punctuation has to be removed first, or
bundle.looks like a host name and the whole word is skipped.
Quoted text passes on purpose. Anything inside backticks or quotation marks is left alone, so we can still cite a Russian error message or an English sentence when we mean to.
Two details only showed up in daily use. The hook's own message is written in Bulgarian and shows the offending words transliterated into Latin letters; otherwise the guard itself fills the console with the very words it exists to stop. And it reads only what was written after the person's last real message, or after the last time the hook spoke. Without that boundary, a report arriving from a background agent started a fresh check that flagged lines we had already corrected. The stop_hook_active flag in the hook input covers only the continuation right after a block, not that case.
What we did not measure
- Every case ran a single time on each model. A repeat could move any single result, and three misses out of 24 are too few to rank the models finely.
- The cases are ours and single-turn. The drift we actually saw happened in long sessions with tool output, which this test does not reproduce.
- Two models, both from one vendor. Other model families may drift more or less.
- Language detection on very short answers can misjudge, and we loosened two checks after the first run, for both sides. A looser check can also hide a real slip.
- The hook was tuned on Bulgarian replies written for one team. For another language its letter lists, endings and word lists must be rebuilt, and we have not measured its false alarms anywhere else.
- We did not count how many extra turns the hook adds in a working day.
Read this post as Markdown: /blog/llm-replies-in-wrong-language.md · Atom feed.
