Test results
Timeline From Text: test results
Tested 2026-10-09, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-09
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right on 23 of 24 texts, checked by code; it missed the format once: on a text with two date styles it sent a first list out of order, then a sentence and a corrected list, so the answer is no longer one JSON object. Everywhere else it held: relative dates resolved only from a dated anchor, month lengths and a leap year, an ambiguous 03/04 kept with a null ISO value, year-only and month dates not padded, ranges, no year borrowed, a recap not counted twice, undated events kept apart, quotes copied exactly and planted notes in English and Bulgarian ignored.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Right on 21 of 24, checked by code, but it dated a repair two days after the wrong event, put a month before the year-only date it follows, and listed a closing wish as an undated event.
With and without the skill
Tested 2026-10-09.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Timelines right (24 texts) | 23/24 | 23/24 | 21/24 | 21/24 |
Same request and same checks on both sides. The request already states the JSON shape, the order and what counts as dated, and with it Sonnet without the skill made the same single slip as with it: on the text with two date styles it corrected itself after the first list. Haiku missed three on each side, different ones: without the skill it miscounted a relative month, fenced one answer and listed an event twice; with it the slips above.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-four texts written by us in English, Bulgarian, German and Spanish: 16 with traps (relative dates with and without an anchor, a leap year, two date styles, an ambiguous numeric date, year-only and month dates, a range, a date that belongs to another event, planted notes, dates without a year, a recap of an earlier event, Spanish order) and 8 plain controls, one with no event at all. Each answer is parsed as JSON; the ISO values, the order and the undated list are compared with ours, and every quote must appear in the text character for character. No check was widened. One run per model and text.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.