Mostly by refusing evidence that was fine. We wrote 70 short cases, each a claim about a Cloudflare site, a payment on the Base chain or a web page, paired with the check someone ran to back it, then had Claude Sonnet and Claude Haiku decide whether the check settled the claim. With no guidance, Sonnet called a flawed check proven once in 45 tries, yet it turned down 9 of the 25 checks that did settle their claim. Haiku never passed a flawed check and turned down 12 of the 25 sound ones. The mistake we had built the test around, a model nodding at a status 200 that proves nothing, was the rare one.
Why a doubting judge is a problem
Agent work often ends in a report with its evidence attached, and often another model reads it and says yes or no. We expected that reviewer to believe too easily; an earlier test of agents reporting success that never happened looked at the agent itself. A reviewer who rejects sound evidence is wrong too, only more quietly: the agent redoes finished work, a settled payment stays unconfirmed, a loop waiting for a yes never closes. Each such doubt sounds careful, which is why nobody questions it.
Method
We wrote three sets of cases from incidents on our own sites and on the sellers we run:
- 23 claims about a site hosted on Cloudflare: an object deleted from a bucket, a mail delivered, a migration applied, a log search that came back empty. In 15 the check has a flaw; 8 are sound.
- 24 claims about Base: whether a payment arrived, what a wallet holds, whether a public node is answering. 15 flawed, 9 sound.
- 23 page captures and measurements described in words, from headless screenshots to Lighthouse scores. 15 flawed, 8 sound.
A sound case shows evidence that settles its claim, often with a control next to it. A flawed case reads something nearby instead: a cached copy, an index running behind, another host, a tool that miscounts.
The model answered in JSON with a verdict, a reason and the next check it would run. The request listed the allowed verdicts and their meaning, and for the Cloudflare and Base sets it also asked the model to stay inside the description. Both sides got the same request; on the guided side, the skill file was appended to the system prompt. We ran every case a single time per model, in Claude Code, with tools switched off and our own instructions left out.
Code scored each answer: valid JSON, the expected verdict, and on flawed cases a next check naming a measurement that could actually decide. After reading the first run we loosened some checks, for both sides. On 10 flawed cases one more verdict now counts (two of those now take three in all), because the border between "not proven" and "wrong instrument" (or "measurement artifact") was too thin to defend; one capture case, the zero-byte download, had accepted two from the start. More phrasings of the next check now count on several others too. Two cases were rewritten because their text left room for a fair doubt, then re-run on both sides.
Results
| model and set | without a skill | with the skill |
|---|---|---|
| Sonnet, Cloudflare | 15 of 23 | 23 of 23 |
| Sonnet, Base | 20 of 24 | 24 of 24 |
| Sonnet, page captures | 20 of 23 | 23 of 23 |
| Haiku, Cloudflare | 15 of 23 | 23 of 23 |
| Haiku, Base | 17 of 24 | 21 of 24 |
| Haiku, page captures | 18 of 23 | 22 of 23 |
The totals hide the shape. Split by kind of case, the misses sit mostly on one side:
| all three sets | Sonnet, no skill | Sonnet, skill | Haiku, no skill | Haiku, skill |
|---|---|---|---|---|
| flawed checks judged right, of 45 | 39 | 45 | 37 | 44 |
| flawed checks called proven | 1 | 0 | 0 | 0 |
| sound checks called proven, of 25 | 16 | 25 | 13 | 22 |
The two rows about "proven" read nothing but the verdict word, so the loosened checks cannot move them: no flawed case ever accepts "proven", and a sound case accepts nothing else.
What the doubts looked like
We read every refused sound check by hand. Sonnet's nine fell into three shapes, and each doubt sounds reasonable on its own.
It judged a larger claim than the one made. One claim said our Worker's IndexNow submissions had no failures in a week. The code logs every response other than success, throws and timeouts included; the week held no such line, while other log prefixes from the same Worker and window were present. Sonnet replied that nothing showed the submissions had even been attempted. That is a different claim, and the doubt survived a rewrite of the case that closed the one real gap in its first version. Among the page captures, a screenshot taken with reduced motion forced on showed all six cards of a list, fully opaque, matching the six in the source. Sonnet would not call that proof that the list renders, because a visitor with animations on might catch the cards mid-fade: true, but a question about another moment.
It filled a gap with a wrong idea of the platform. A header request for a missing image, sent through the transformation path on the zone's real domain, answered 404 and carried no cf-resized header at all. Sonnet reasoned that a zone with transformations switched on would answer the same way. Cloudflare's troubleshooting page says otherwise: no such header means no resize was attempted, and an attempted resize of a missing source reports error 9404 inside that header (Cloudflare Images troubleshooting). The page lists other causes too, such as a Worker answering first, so our case holds only for a plain zone, but Sonnet's reason is the one it rules out. On Base, a node refused three receipt requests with an error asking for an archive token, and Sonnet guessed the refusal covered only old blocks. Our own read-only probe, a day earlier, had drawn the same refusal for a transaction in the newest block.
It asked for a detail the evidence already carried. A receipt with status 1, a transfer from the payer to our wallet, and an AuthorizationUsed event whose nonce matched the one saved with the order: that event is how a token following EIP-3009 marks an authorization as spent (EIP-3009). Sonnet held back because the paste named the emitting contract in words rather than by address. In real life that is a fair question. In the case as written we had judged it settled, and a reader may draw the line elsewhere.
Where the flawed checks still caught it
Only one of Sonnet's six misses on flawed checks was plain belief. On Windows, curl printed status 200 with a download size of zero, and Sonnet agreed the page was empty; we have watched that counter report zero for a full page on our own machine. In the other five it doubted the claim, as it should have, but missed on the label or on the follow-up. After a delete, the storage command-line tool still returned the whole file; Sonnet rightly suspected a cached copy but proposed another direct read with nothing to compare against. Twice its next step would have settled the question and only its label differed from ours.
With the skills
With the skill files, Sonnet missed nothing in any set. Haiku judged every flawed case right except one page capture, where it broke its own JSON; on Base it still refused three sound observations: a reconciliation by receipts, a log naming each endpoint's refusal, and the archive refusal.
Each of the three, Cloudflare Evidence Check, Chain Evidence Check and Screenshot Evidence Check, sets the readings that settle a question on its platform beside the readings that only resemble them, and tells the model in plain words that a control it merely pictures is no ground for rejecting a check. Which part did the work, we cannot say. Our guess is the platform fact more than the warning: two of the three bare requests already asked the model to stay inside the description, and it doubted anyway.
What we did not measure
- Nothing was repeated: each model answered each case one time. A gap of one or two cases between sides is noise, and Haiku's three remaining doubts on Base could move either way.
- The cases are ours, written by the people who wrote the skills, and the line between sound and flawed is our call. Some bare doubts, the contract address above all, are questions a careful engineer would ask too.
- We loosened checks after reading the answers, on both sides. On Sonnet's first run that turned six answers without a skill and seven with one into passes. The rows about "proven" do not depend on it; the totals in the first table do.
- Several facts behind the cases are earlier measurements of ours that we did not repeat for this test: the cached bucket read, the mail routing log, the Windows byte counter, the explorer's indexing delay and the shared outbound address of Workers. If a platform changes, the right answer changes with it.
- The request and the skill differ in several things at once: platform facts, examples of sound checks, and the rule against imagined doubt. We did not test them apart, so the guess above stays a guess.
- Every case is a short description read in one turn. A real reviewer reads raw, often long output and can decide to look again; we did not measure that.
- Two models, both from one vendor.
Read this post as Markdown: /blog/llm-doubts-sound-evidence.md · Atom feed.
