Test results
Source Quotes Extract: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right on all 23 texts, read by hand: one claim per fact with the shortest word-for-word quote behind it, hedges such as probably and expected kept in the claim, numbers and units kept in the text's own format (1 250 stays 1 250), a fact that spans two sentences kept as one claim, reported speech and a minister's opinion left out, a cause the text only hints at not stated as a cause, claims written in the text's language (Bulgarian, German, Spanish), and a sentence planted in the text treated as text.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Right on 22 of 23 texts, read by hand, but it missed one: on a Spanish text with a planted sentence it wrote the claims in English instead of Spanish, although the quotes were right.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Texts extracted right (23 texts) | 23/23 | 19/23 | 22/23 | 15/23 |
Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already quoted word for word and left reported speech out. It missed four, three of them in Bulgarian: it dropped the hedge from a probable figure, rewrote 1 250 in another number format, split one fact across two claims, and on a plain German text paraphrased the claim so that the word for free was lost. Haiku without the skill missed eight, mostly the same hedges, number formats and plain cases in other languages.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-three short texts written by us in English, Bulgarian, German and Spanish (16 with a trap, 7 plain): hedged and rounded figures, numbers with spaces as thousand separators, units, a fact split over two sentences, reported speech, an implied cause, and planted sentences. Each answer is parsed as JSON and checked by code: every quote must appear word for word in the text, the claims must carry the hedge, the number and the unit as written, the number of claims must fit, and the claims must be in the language of the text. The quoting rule follows the public guidance on grounding answers in direct quotes. No check was widened. One run per model and text.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.