Test results
x402 Settlement Review: Replay and Double Charge: test results
Tested 2026-10-09, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-09
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right on all 24 snippets, checked by code on the finding codes and the verdict line: an expiry check before the nonce lookup, repeat paths that hand over the file with no proof or with a foreign body, a repeat window taken from the buyer's validBefore, settlement_pending treated as terminal, a public header that lifts a ceiling or skips a check, a bucket keyed on the claimed from, verify with no global ceiling, a health lamp fed by another facilitator or calling a verdict down, 400 reasons logged as unknown, a network kept in two places and a planted comment; No findings. on all eight sound snippets.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Right on 22 of 24 snippets, checked by code. It found every planted defect, but on two sound snippets it reported one that is not there: a receipt returned on a repeat, which the skill itself prescribes, and verify calls whose ceiling sits in code the paste only calls.
With and without the skill
Tested 2026-10-09.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Planted defects found (16 snippets) | 16/16 | 16/16 | 16/16 | 16/16 |
Same request on both sides, a fence removed first. Without the skill both models named every planted defect in their own words, so the count shows no gain. What it does not show: without the skill each answer is a long review of 4 to 7 thousand characters, and on the six sound snippets we read, both models listed high-severity items; some are real issues outside the ten codes, such as a validBefore that is not a number skipping the expiry check. With the skill Sonnet answered No findings. on all eight sound snippets, in one line each.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-four snippets of seller-side x402 code written by us: 16 with a planted defect (one with a planted comment, one with two defects) and 8 sound ones, among them a buyer-side script that is not a register. With the skill each answer is scored by code on the finding codes and the verdict line; without it the same request is scored on the concept in any words, and on the sound snippets the bare side has no check, so the comparison rests on the 16 faulty ones. Changes after the first run: one sound adapter put settlement_pending under the same kind as a refusal, which both models flagged, so it now gets a kind of its own; the skill now says that the verdict is the last line and that a part not pasted at all gets no finding and no note (Sonnet had added notes after the verdict twice). The side with the skill was run again in full; the numbers use that run. Facts on the facilitator's 400 body, the forged-from bucket, the random-signature repeat and the health lamp are owner measurements, not re-checked. One run per model and snippet.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.
Back to x402 Settlement Review: Replay and Double Charge · Card (JSON)