Test results
Shorten to a Limit: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right on 18 of 23 texts, read by hand, but it missed five: every answer kept to the word or character limit, yet on five texts it dropped a fact it should have kept, a date such as 5 November or 30. November, a figure such as 30 000, 260.000 or 1,450, or a product code. Everything else held: the limit counted the way the request defines it, a text that already fits returned unchanged, negations and conditions kept, a quotation left whole, a web address and handles never cut, the main fact kept complete when the limit was very small, and an order inside the text cut like filler, never obeyed.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Right on 20 of 23 texts, read by hand, but it missed three: within the limit, it dropped a date (3 June) and two figures (260.000 and 1,450 with a product code).
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Texts shortened right (23 texts) | 18/23 | 10/23 | 20/23 | 6/23 |
Same request on both sides; the request states the limit and how to count. Read by hand, Sonnet without the skill wrote good short texts but went over the limit on nine of them, plainly miscounting words, and on others dropped a fact (a date range, a qualifier) or rewrote a text that already fitted. With the skill it never went over the limit, but it still dropped a fact five times. Haiku without the skill went over the limit on most texts.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-three texts written by us in English, Bulgarian, German and Spanish, each with a word or character limit stated in the request (16 with a trap, 7 plain): negations, quotations, attributions, up-to qualifiers, conditions, product codes, local number formats, hyphenated words, repeated points, a limit too small for the whole message, a text that already fits, and a planted order. Each answer is counted by our script the way the request defines a word or a character, and checked for the facts that must survive and for anything that must not appear. Word boundaries follow the request's own definition (runs between spaces), close to the Unicode text-segmentation rules. No check was widened. One run per model and text.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.