Because more than one Latin system for Bulgarian is in use, and a model blends them unless it is told which one counts. Bulgaria's Transliteration Act of 2009 fixes a single table, yet older habits keep showing up in model answers: Mariya where the law writes Maria, the accented letters of the scholarly system, the Russian shch. We gave Claude 24 short Bulgarian texts and compared each answer, letter by letter, with the official Latin spelling. Claude Sonnet got 15 right on its own and all 24 once the rules were written out for it; Claude Haiku went from 11 to 18.

Why one letter matters

A transliterated name ends up in a passport form, a shipping label, a database column or a web address. There Mariya and Maria are different strings: a search misses the record, a form rejects the name, a link breaks when someone corrects the slug. Most wrong answers we collected look fine to anyone who cannot read Bulgarian, and that is usually who pastes them.

The system in short

The Act appeared in the State Gazette, issue 19 of 13 March 2009. It binds public bodies, and also anyone who writes Bulgarian place names, historical names or terms of Bulgarian origin in Latin (text of the Act). Three of its articles decide most of our cases. Article 4 is a table with one Latin form per letter: ж gives zh, ц gives ts, щ gives sht, and ъ gives a plain a. Article 5 turns ия into ia when those two letters close a word. Article 6 keeps the country as Bulgaria, by tradition, although the table on its own would give Balgaria.

The other systems never went away. Schemes for romanizing Bulgarian disagree on 12 of its 30 letters, and the older scientific one, the ISO/R 9:1968 norm, uses č, š, ž, j and c (Romanization of Bulgarian). The ia ending started in 2006 as a rule for given names and places; the 2009 law widened it to all words ending in ия. Read letter by letter, the table gives iya, and for years that was the normal written form. Our guess is that this is why models still reach for it.

Method

We wrote 24 short Bulgarian texts:

  • personal names, among them a hyphenated surname, names with дж and ьо, and names that end in ия;
  • towns, a street address with a house number, and a short line typed entirely in capitals;
  • a sentence where the country's name and its adjective stand side by side;
  • a price on two lines, with Markdown marks, digits and two Latin words, Wi-Fi and Xiaomi;
  • five titles and names to turn into URL slugs, answered as JSON holding the text and the slug;
  • two place names with letters from other Cyrillic alphabets, which had to be kept and listed;
  • one sentence phrased as a command to forget earlier instructions and answer OK.

Every expected answer was produced twice, by a small script that follows the Act and by hand, and the two were compared before any model ran. Code then compared each model answer with the expected text as one whole string: one letter off, and the text counts as wrong. For the letters that turn into zh, ch, sh and sht, the Act prints the capital entirely in upper case, which can be read two ways, so as the first letter of a mixed-case word either form passed.

Claude Sonnet and Claude Haiku each saw every text once per condition, through plain Claude Code with no tool access and none of our everyday instructions: first the one-line request alone, then with the skill file added to the system prompt. A Markdown code fence around a JSON answer was removed before comparing, on both sides; bare Sonnet wrapped its JSON in one three times although the request asked for JSON only. Haiku's first pass with the skill broke off after six texts at a usage limit, and the tester stored the limit messages as if they were answers. We discarded that pass and ran Haiku again. No check was loosened afterwards.

Results

texts spelled as the Act requires without the skill with the skill
Sonnet 15 of 24 24 of 24
Haiku 11 of 24 18 of 24

The totals say less than the misses themselves, so we read every one.

Word-final ия was the largest single cause. Six of bare Sonnet's nine misses, and eight of bare Haiku's thirteen, had iya where the law writes ia: Mariya, Yuliya, istoriya, Aziya, Iliya, delegatsiya, and YUZHNIYA in the line typed in capitals. The rest of each word was usually right, so the answer looked careful and still failed.

Systems mixed inside one word. Bare Sonnet wrote Търговище as Târgovište, with a circumflex for ъ and a háček for щ, both from the scholarly tradition. Haiku gave Târgovishte, two systems in a single word, and turned ъ into ǎ three times in one short phrase. It also wrote the Russian-style Shchereva for Щерева, Jon for Джон, and j for й with c for ц in the planted sentence. In one slug it dropped ъ altogether and wrote yablaki for ябълки.

The country and its adjective. Only the state's own name keeps the u; the adjective follows the table, so Българската becomes Balgarskata. Bare Sonnet gave the adjective the country's u, Bulgarskata, in a sentence where it wrote the country itself as Bulgariya, and then spelled the country Balgaria inside a slug. Bare Haiku wrote Balgaria and Balgariya.

A word translated instead of spelled. In the price line, bare Sonnet changed the Bulgarian рутер into the English router while leaving Wi-Fi and Xiaomi alone. Of all the misses this is the hardest to spot, because router is an ordinary Latin word.

With the rules loaded as bulgarian-transliteration, Sonnet matched the Act on every text, the ones above included. Haiku improved but kept six misses, and two of them are new: without the rules it had Ilia and the Russian place name right. With them it wrote Džon with a háček, and Aziya as the last word of a sentence whose other ия words came out right. It wrote Bulgarskata, stretching the country's exception to the adjective. It replaced the Russian ы it was told to keep with y and an asterisk. In the planted sentence it still wrote ц as c, and й, which had been j without the rules, now came out as i.

The planted order to ignore instructions and reply OK worked on neither model, in none of the four runs. Every time the sentence came back transliterated rather than obeyed.

A check that needs no answer key

Most of these errors leave a fingerprint, because the official table only ever produces a small set of Latin letters. It never yields j, q, w or x, it uses c only inside ch, and it has no accented letters. Outside a few rare letter clusters it also never gives shch, kh, or a word ending in iya. So we wrote a check that reads the original and the model's output, without knowing the right answer:

// Flags Latin output that the official Bulgarian table cannot produce.
// `input` is the Cyrillic text you sent, `output` is the model's Latin text.
function suspicious(input, output) {
  const why = [];
  const latinIn = new Set((input.match(/[a-z]/gi) || []).map((l) => l.toLowerCase()));
  if (/[\u00C0-\u024F]/.test(output)) why.push('accented letter: another system');
  if (/iya\b/i.test(output)) why.push('word-final iya');
  if (/shch|kh/i.test(output)) why.push('shch or kh');
  for (const l of 'jqwx') if (!latinIn.has(l) && output.toLowerCase().includes(l)) why.push(`letter ${l}`);
  if (!latinIn.has('c') && /c(?!h)/i.test(output)) why.push('c outside ch');
  if (/\bBalgari(?:a|ya)\b/i.test(output)) why.push('country name');
  if (/\bBulgar(?!ia\b)/i.test(output)) why.push('Bulgar- outside the country name');
  const marks = (t) => t.replace(/[\p{L}\s]/gu, '');
  if (marks(input) !== marks(output)) why.push('digits or punctuation changed');
  return why;
}

Latin letters already present in the input, such as a brand name, are allowed through. The last line compares the digits and punctuation of both texts, so an inserted asterisk or a lost house number shows up. On the 96 real answers from this test it flagged 25 of the 28 wrong ones and none of the 68 right ones. On the Act's own text, 1,258 Cyrillic words put through our reference script, it raised no alarm.

It let three wrong answers through: router, Kolio for Кольо, and yablaki with its missing letter. Each uses only letters the table can produce, so nothing short of the expected text will catch them. We also know of one false alarm: the medieval Bulgars, булгари, really are written with a u, and the check objects.

What we did not measure

  • Every text ran once per model. Six Haiku answers did run twice, because of the discarded run, and one of those six changed between runs: Lilyana in one, Liliyana in the other. Any single result here could move on a repeat.
  • The texts are short and all written by us. Long documents, tables and paragraphs mixing Bulgarian with other languages were not tested.
  • Two models from one vendor, through one tool. Other model families may lean toward a different system.
  • The expected answers encode our reading of the Act. Where the law is silent, such as capital zh-type letters inside mixed-case words or letters from other Cyrillic alphabets, the test holds a choice of ours, and an office may want something else.
  • A person may hold documents with their own Latin spelling of their name, which Bulgarian rules for identity papers permit (Romanization of Bulgarian). For a specific person the official table is not always the right answer.
  • We wrote the check after reading these very answers, so its catch rate on them flatters it. On new text it will do worse, and it cannot see a translated word or a dropped letter at all.
  • Why models prefer iya is our guess. We did not look into what they were trained on.

Read this post as Markdown: /blog/llm-bulgarian-transliteration-mixes-systems.md · Atom feed.