Test results
Regex From Examples for Any Script: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Wrote a regex that passes all 22 tasks: every must-match and must-not-match string, and long near-miss inputs inside 200 ms. Word boundaries in Cyrillic, Greek and accented text use lookarounds with Unicode properties and the u flag, validators are anchored, multi-line input gets the right flag, and repeated groups avoid nested quantifiers.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Also 22 of 22, with the same kinds of pattern as Sonnet.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Regexes that pass every string (22 tasks) | 22/22 | 22/22 | 22/22 | 22/22 |
No measurable gain on these tasks, for either model. The same request, asking for JSON only, went to both sides. When the task lists the strings that must match, both models already wrote lookarounds with p{L} and the u flag for Cyrillic and Greek instead of , anchored the validators and avoided nested quantifiers. What the skill adds is the measured notes on why those choices matter and a fixed JSON answer; it may help more when no examples are given or on weaker models than the two tested.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-two tasks written by us, each with strings that must match, strings that must not, and long near-miss strings that must finish within 200 ms. About a third are plain tasks where the obvious regex is right; the rest set traps: word boundaries in Bulgarian, Greek and accented text, whole-value validators, multi-line input, names with letters outside A-Z, and patterns that hang on nested quantifiers. Each returned pattern was compiled and run against all strings. The JavaScript facts the skill states were measured in Node 24 on 8 October 2026. One run per model and task.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.