Translate App Text (JSON, YAML, PO): test results
Tested 2026-10-03, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-03
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Passed all 9 cases. JSON, ICU, i18next, PO, Rails YAML and HTML came back with every key, placeholder, ICU keyword and tag intact; it added the Russian few and many forms inside ICU messages and as i18next keys, switched the Rails root key to es and the PO Language header to bg, translated an injected instruction as plain text, and left already translated values, numbers and the brand untouched. Read by hand: natural Bulgarian, German and Russian, Spanish in the informal tú.
- Weak · claude-haiku-4-5-20251001 (Claude Code alias "haiku")
- The structure was right in all 9 cases, including the Russian plural forms, but it wrapped the JSON in a code fence in 4 of 9. After we moved the no-fence rule to the top, 1 of those 4 reruns still had the fence, so strip one before saving the file. Read by hand: weaker wording (Bulgarian 'browse' became 'разгледайте' instead of 'изберете'), and a fractional Russian form picked wrongly in one 'other' branch.
Note
Nine fictional files translated from English into Bulgarian, German, Russian and Spanish: placeholders of six libraries, ICU plural and select, i18next plural keys, gettext PO, Rails YAML, HTML attributes, an injected instruction and a partly translated file. Machine checks parse the JSON and compare every key and placeholder; wording was read by hand. One run per model and case; the 4 Haiku failures were rerun once after the fix. The Sonnet run used the text before three edits: the no-fence rule moved to the top, the Spanish default changed to tú (matching what Sonnet already did), and tag examples reworded in plain words so the file carries no HTML.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.