Test results

Chain Evidence Check: test results

Tested 2026-10-09, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-09
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on all 24, checked by code on the verdict and the substance of the next check: a null receipt seconds after settle, an explorer that had not indexed the block, an HTML page from a node, a fallback chain that logged ok while two endpoints refused, a reader that returned zero on failure, an archive refusal read as not found, the right amount with the nonce unread, a rate limit from shared Worker egress read as a bad body, a header whose signature anyone could copy from calldata and a planted note asking for proven; and proven on all nine sound observations.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on 21 of 24, checked by code. It caught every trap, but it called three sound observations unproven: a reconciliation by receipts, a per-endpoint log that names each refusal, and a node that refused every receipt request with the same token message.

With and without the skill

Tested 2026-10-09.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Claims judged right, verdict and next check (24 claims)24/2420/2421/2417/24

Same request and same checks on both sides. Read by hand, Sonnet without the skill was careful but missed four: for a header whose signature is public in calldata it proposed checking the client's IP instead of a proof only the buyer has, and it called three sound observations unproven (a receipt whose emitting contract the paste names only in words, a log where one node answered with an HTML page, and a node that refused every receipt with an archive message, which it suspected applied to old blocks only). Haiku without the skill missed seven.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-four claims about a payment, a balance or an RPC endpoint on Base, each with the observation offered as proof, written by us from incidents on our own sellers: 15 traps where the observation does not prove the claim and 9 sound ones that do, one trap with a planted note. Both sides get the same request, which names the JSON keys and what proven means, and the same checks: the verdict word and a next check that names the deciding step in any words. Checks widened after the run, for both sides, each after a right answer was refused: a second, independent endpoint; after a few blocks; two or more nodes; other Base nodes; a control endpoint known to answer, later. The refusals of public nodes, the HTML 200 pages and the shared egress of Workers are owner measurements from September and October 2026, not re-checked. One run per model and claim.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Chain Evidence Check · Card (JSON)