Test results

Agent Report Audit: Which Claims the Log Supports: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on 23 of 24 reports, read by hand, but it missed one: after a correct JSON answer it added a note about how it quoted the log, so a program reading JSON only would reject it. Everything else held: a claim with no tool call behind it not supported, a dry run not counted as an upload, a failed test count read from the output, an edit made after the last green test run caught, an honest report of bad news supported, every quote copied exactly from the log, and reasons in the language of the report.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on all 24 reports, read by hand, with the same verdicts and quoted lines as Sonnet.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Reports audited right (24 reports)23/2422/2424/2422/24

Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already judged almost every claim right. It missed two on the quoted evidence: it gave the call line (shell npm test, edit src/totals.ts) where the output line is what decides the claim. Haiku without the skill made the same two slips.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-four agent reports with their tool logs written by us, in English and Bulgarian (16 with a trap, 8 plain): prose claims with no call, dry runs, failed tests reported as green, edits after the last test run, honest bad news and a planted instruction. Each answer is parsed as JSON and checked by code: the verdict per claim, the deciding log line copied exactly, the reason. The first run with the skill showed a gap in it: Sonnet wrote its reasons in Portuguese and French for English reports twice. The rule was made explicit (the language of the report's claims, never a third language) and the side with the skill was run again in full. No check was widened. One run per model and report on the final version.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Agent Report Audit: Which Claims the Log Supports · Card (JSON)