Test results

x402 Listing Review: Catalogs, Scanners, Buyers: test results

Tested 2026-10-09, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-09
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on 20 of 23, checked by code on the codes and the verdict line. It named every planted problem: a tariff advertised as one row or as four, a price above the default per-payment cap with no note where agents read, a well-known file the scanner no longer parses, a bearer scheme under an unknown name, a presence script reading items, checks with no control that must come back empty, one clock for several rows, a crawler mailed as a customer, discovery that queries the database, and a planted comment. Twice it added a STALE-RECORD line where the paste holds no plan that expects the scanner page to change, and once it missed the verdict, giving the first-group one for a crawler mailed as a customer.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on 20 of 23, checked by code. It named every planted problem, but it added a PARETO-ROWS line on two snippets whose tariff was not in question and reported an unstated cap on a sound listing with two rows.

With and without the skill

Tested 2026-10-09.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Planted problems named (16 cases)16/1610/1616/166/16

Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill wrote long, sensible reviews but missed six planted problems, mostly facts it could not know: that the standard buyer client stops at one dollar per payment (two cases), that the catalog answers under resources and not items (two), and that a presence check needs a control query that must come back empty; nor did it say that four rows where two carry the tariff only clutter the catalog. Haiku without the skill missed ten.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-three listings, scripts, plans and code snippets written by us: 16 with a planted problem (one with a planted comment, one with two) and 7 sound ones. With the skill each answer is scored by code on the finding codes and the verdict line; without it the same request is scored on the problem named in any words, and the sound snippets have no check on that side. Changes after the first run: one sound snippet was dropped, because both models found real flaws in it (one explorer page cannot reach a hundred recipients, and the follow-up count runs before the payment and includes the paid request); the skill's fix for a crawler buyer now says to follow the explorer's pages and to decide "no later request" in a delayed job. The side with the skill was run again in full. Checks widened for both sides, each after a right answer was refused: bigger packs that "do not exist as listings" or "each pack its own route", a raw key sent without the word Bearer, and, for the scheme name, a 402 answered before the token check accepted as another working fix. The catalog fields, the default cap of the buyer client and the scanner indexer's reading of OpenAPI were read on public pages on 8 October 2026; the name-based scheme classification, the stale scanner record and the crawler buyers are owner observations from September, not re-checked. One run per model and case.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to x402 Listing Review: Catalogs, Scanners, Buyers · Card (JSON)