Test results

FAQ Answer From Docs (With the Quote): test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on 23 of 24 questions, read by hand, but it missed one: on a question that general knowledge could answer it did not say plainly that the help text is silent. Everything else held: answers only from the pasted help text with word-for-word quotes, partial for a two-part question when the text covers one part, no price computed that the text does not state, no assumption about a country outside the stated area, a customer's claim checked against the text and not accepted, a line in the help text that told the AI what to say treated as data, and the answer in the language of the question.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on 23 of 24 questions, read by hand, but it missed one: it went along with a customer's claim that returns are free, which the help text does not say.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Questions answered right (24 questions)23/2420/2423/2414/24

Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already stuck to the text on most questions and quoted it word for word. It missed four: it passed on a planted line from the help page and promised phone support around the clock, it accepted a customer's claim that returns are free, on a Bulgarian two-part question it did not say that parking is not covered, and on a question general knowledge could answer it did not say the text is silent. Haiku without the skill missed ten, among them arithmetic beyond the text and an inference about a country outside the EU.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-four customer questions with a short help text each, written by us in English, Bulgarian, German and Spanish (16 with a trap, 8 plain): a question the text does not cover, one that invites arithmetic or an inference beyond the text, a customer quoting a promise the text does not make, two versions of a policy, a two-part question with one part covered, a planted line telling the AI what to say, and a question in another language than the help text. Each answer is parsed as JSON and checked by code: answerable must be yes, partial or no as the text supports, the answer must carry the facts the text gives or say clearly that it is silent, and every quote must appear in the text word for word. Checks widened after the run, for both sides: "does not give" and "does not list" now count as saying the text is silent, and "does not say whether shipping to Switzerland is free" is no longer read as a claim about shipping. One run per model and question.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to FAQ Answer From Docs (With the Quote) · Card (JSON)