Test results

Extract UI Strings Into i18n Keys: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on all 24 files, read by hand: it moved every visible text from JSX, Vue templates and HTML into the map byte for byte, kept a sentence with a link as one message with the markup inside, turned variables into named placeholders and a count into a plural form, took aria labels, alt and title texts but left class names, ids, URLs, test ids and data values alone, gave a repeated text one key, wrote keys of small letters, digits and dots only, left texts already in t() untouched, kept Bulgarian and Spanish texts exactly, and treated a comment that spoke to the AI as code.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on all 24 files, read by hand, with the same messages, keys and rewritten files as Sonnet.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Files extracted right (24 files)24/2422/2424/2421/24

Same request on both sides; the task itself names the key form and the output shape. Read by hand, Sonnet without the skill already extracted the visible texts, attributes, links and placeholders correctly. It missed two: it built a count message from pieces instead of one plural form, and it wrote a key with an underscore that the task forbids. Haiku without the skill missed three: typographic quotes changed, an aria label left in the code, and the same count message.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-four JSX, Vue and HTML files written by us, in English, Bulgarian and Spanish, with traps: attributes that must be extracted next to ones that must not, a ternary, a sentence split by a link, a variable, a count, typographic quotes, repeated text, texts already translated, dynamic data that must stay in the code and a planted comment. Each answer is parsed as JSON and checked by code: the map against the expected texts, the key form the task states (small letters, digits and dots), and the rewritten file against the original with the calls in place. The first run with the skill showed a gap in the skill itself: its key rule did not repeat that underscores are not allowed, and Sonnet wrote keys like save_failed on two files. The rule was made explicit and the side with the skill was run again in full for both models; the numbers use that second run. No check was widened. One run per model and file on each side.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Extract UI Strings Into i18n Keys · Card (JSON)