Test results

Summary Fidelity Check: test results

Tested 2026-10-09, skill version 1.0.1 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-09
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on all 24, checked by code: a rounded number and a changed unit, dropped hedges, added causes, two facts merged into a false one, dropped speakers, outside facts, a changed quantifier and planted notes in English and Bulgarian, each flagged with the sentence copied exactly; an empty list for eight faithful summaries, and plain words with no list when the summary was missing.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on all 24, checked by code, with the same flags and the same empty lists as Sonnet.

With and without the skill

Tested 2026-10-09.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Summaries judged right (24 pairs)24/2424/2424/2423/24

Same request and same checks on both sides. The request already spells out what unsupported means and the JSON shape, and with it Sonnet without the skill was right on all 24, faithful summaries included. Haiku without the skill missed one: a Bulgarian summary that dropped who made a claim about unemployment.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-four source and summary pairs written by us in English, Bulgarian, German and Spanish: 15 summaries with a planted unsupported sentence (rounded numbers, units, hedges, causes, merged facts, dropped speakers, outside facts, a quantifier, planted notes), one request with no summary, and 8 faithful summaries built to tempt a false flag. Each answer is parsed as JSON; the flagged sentences must be exactly the planted ones, copied character for character. No check was widened. One run per model and pair.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Summary Fidelity Check · Card (JSON)