Test results

ICU Plural Messages (per-language CLDR categories): test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Wrote usable ICU messages for all 23 cases, read by hand: Japanese, Korean and Chinese got no one arm, English got =0 rather than a zero arm, the apostrophe before a placeholder was doubled and a literal brace was quoted, an offset:1 message counted the viewer out, and the fix cases added the missing Russian forms and the missing other arm.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Also 23 of 23 with the skill: the same arms per language as Sonnet, including many for French, Spanish, Italian, Czech and Brazilian Portuguese, which it left out without the skill.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Messages with the right arms (23 cases)23/2323/2323/2320/23

The same request on both sides, JSON-only asked for in both; a fence around the answer is removed first. Read by hand, Sonnet without the skill was right everywhere on content (its only miss was a code fence around one answer, which the rescoring strips). Haiku without the skill wrote no many arm for Czech, for French, Spanish and Italian, and for Brazilian Portuguese, so the number one million would show the wrong wording. Everything else Haiku wrote alone passed, including the apostrophe, the literal brace and the other-only languages.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-three requests written by us, in two to four locales each: plain counts, exact zero, ordinals, select with a nested plural, offset, other-only languages, the dual and zero languages, a literal brace, an apostrophe before a placeholder, two repairs and a planted instruction. Each answer is scored by regexes on the message strings: the arms the language needs where the wording differs, other always, and a dead arm only where the case is about it. The checks were relaxed AFTER this run to accept every working message: an offset message may use =2 instead of one, an English zero case needs no one arm when =1 is given, Turkish may be other only, and Irish and the Slavic day cases require only the arms whose words differ. One run per model and case; the numbers below are from rescoring the recorded answers with the relaxed checks.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to ICU Plural Messages (per-language CLDR categories) · Card (JSON)