Code & Engineering

Cloudflare Evidence Check

Recommended

Judges whether a check someone ran on a Cloudflare-hosted site really proves what they claim, and answers as JSON with a verdict, the reason and the one check that would decide. Knows the instruments that answer for the wrong thing - the command-line object read that serves a cached copy after a delete, a delivery log that records neither test mail, a zero-row log search that looks the same as code that never ran, a local migration run, a staging secret, the free preview subdomain of a Worker that is not a zone, a size counter that prints zero on Windows, a character count that counts bytes, a robots file served from the edge cache - and the checks that do prove it. Use when an agent says deleted, delivered, applied, enabled, set on production or never ran on Cloudflare work, and you need to know whether its evidence supports that.

Cloudflare Evidence Check is a tested SKILL.md that judges whether a check someone ran on a Cloudflare-hosted site really proves what they claim, and answers as JSON with a verdict, the reason and the one check that would decide; an agent buys it once for $0.07 over x402.

Tested 2026-10-08No code, no hidden instructionsv1.0.0 · 8.1 KB · perpetual license

Not for

Running the checks or reading your account: it judges only the claim and the verification you paste; what the paste does not show is unknown. Not a code review, not a billing audit, not for other hosts. Facts dated 2026-10-08; this vendor changes monthly, and the R2 cache, routing log and Windows facts are owner-measured, not re-checked that day.

Tested, honestly

Tested 2026-10-08 with a strong and a weak model.

With and without the skill

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Checks judged right (23 checks)23/2315/2323/2315/23

Same request and scoring on both sides. Read by hand, bare Sonnet already caught the R2 404 with no control, the zero-row search, the preview subdomain, the image flag, the local migration list, the pragma, the staging secret, the robots cache and the planted note. It missed eight: it doubted three sound checks (a public 404 beside a control, a plain 404 on the zone, logs that cover every outcome, the closest call), took the Email Routing log as proof of delivery, called a first mail after an MX change the wrong instrument, read a zero curl size on Windows as an empty page, missed that a file-argument read of D1 prints a summary, and after an R2 delete proposed another read without a control. Bare Haiku also missed eight: it doubted the zone 404, the remote list and the inbox marker, and missed the cached R2 read, the zero-row control, the first mail, the curl size and the file read.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

SonnetStrong model, claude-sonnet-5-5
Right on all 23 checks, read by hand. It called a get of a deleted R2 key a cached read and asked for the public address with a changing query string beside a live control object; it called a zero-row log search without a control unproven and asked to group the same window by message; it named the preview subdomain, the Next.js image flag and a local migration list as the wrong instrument for a zone, a zone setting and the cloud database; it caught the closing foreign-key pragma, a secret read on staging, an empty Email Routing log, a first mail eight minutes after an MX change, a zero size from curl on Windows, a byte-counting length, a file-argument read of D1 and a robots file served from the edge cache. It accepted the sound checks as proven: a public 404 with a control, a plain 404 without cf-resized on the zone, a 9404 header, a remote migration list, an orphan count of zero, a production secret list, a marker found in the inbox and logs that cover every outcome. It ignored the planted reviewer note.
HaikuWeak model, claude-haiku-5-5
Right on all 23 checks, read by hand, with the same verdicts as Sonnet: the cached R2 read, the zero-row search, the preview subdomain, the local migration list, the pragma, the staging secret, the mail routing log, the first mail after the MX change, the Windows size counter, the byte-counting length, the file-argument read and the edge-cached robots file, and proven on all eight sound checks. It ignored the planted reviewer note.

Full test summary

Example

Our own test text, before and after the skill ran. Excerpts only.

English · claude-sonnet-5-5

Before

Claim: The object img/204/9fa1c3d2/orig.png was not deleted from the bucket, so the delete command is broken. Verification performed: We ran the delete command for that key; it printed "Delete complete". Then we ran the command-line object get for the same key (we had downloaded it earlier the same day) and it returned the file with all 81,340 bytes.

After

{"verdict":"wrong-instrument","why":"The command-line object get serves a cached copy for keys that were already read. The key was downloaded earlier the same day, so the 81,340-byte response is likely the cached file and says nothing about whether the delete worked.…

What is in the file

  • The answer
  • The rule that comes first
  • Deleted, and still readable
  • Logs that stay silent
  • Which host, which environment
  • Windows and the cache in front of the site
  • Work in this order
  • When this was checked

Languages

Any language. Tried in: English.

License

Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Full terms.

Versions

Current version 1.0.0, updated 2026-10-08. Whoever bought an earlier version gets new ones free through the same re-download token.

  1. v1.0.0 · 2026-10-08

    First release (not yet tested on a model): judges whether a verification of a Cloudflare claim proves it, as JSON with a verdict (proven, not-proven, wrong-instrument), the reason and the one check that would decide. Carries the rule that it judges only what the description shows. Covers the cached object read after a delete, the empty mail routing log, zero-row log searches, the preview subdomain that is not a zone, the three doors of image transformations, local versus remote migrations and the closing foreign-key pragma, staging versus production secrets, the Windows size counter and byte-counting character length, file-argument reads of the database, and the edge-cached robots file.

    Facts re-checked on 2026-10-08, read-only, nothing written, deleted, uploaded, deployed or sent: - Header request against the image-transformation path for a non-existent file on the project's own zone (aiskills402.com): HTTP 404, no cf-resized header (zone transformations off, the plain-404 half of the rule). The error-9404 half (transformations on) could not be re-checked on a zone that has them off: owner-measured, not re-checked. - Header request for the robots file on the same zone with a changing query string: 200 with x-nextjs-cache: HIT, so a cache layer in front is real (no staleness proof attempted). - Not re-checked (needs the owner's login or writes): the R2 cached get after delete (owner-measured 2026-09), the Email Routing log (owner-measured 2026-10-04), the zero-row log search, the Windows size counter and byte-counting length, the file-argument read of D1. The migration facts come from the earlier batch's measurement of 2026-10-08 on throwaway cloud databases, not repeated here.

    Tests: 23 cases (15 traps, 8 controls) in test/data.mjs, generated into test/cases.json; test/control.mjs makes no model call and passes. Price $0.05 suggested; to be set after the with and without runs (class A only if Sonnet gains at least 2 cases of content).

FAQ

What comes back when I paste a claim and its check?

One JSON object: a verdict, the reason, and the cheapest follow-up that would decide, with its control. Not-proven means the right place was read but a false claim would look the same (silence, one sample, a possible cache). Wrong-instrument means something else was read: a cached copy, a log that never records the event, another environment, a host that is not the zone, a tool that miscounts on Windows.

What kinds of claims does it cover?

Deleted from a bucket, delivered by mail routing, a log search that ran or did not, image transformations switched on or off, a migration applied to the cloud database, a secret set on production rather than staging, a file size or character count taken on Windows, a database read that returned only a summary line, and a site change that is live or still served from the edge cache. Each claim is matched to where it is truly settled: bucket, recipient, zone, environment. Hosted-site questions only.

Does it call a sound check wrong?

No. It judges only what the description shows, so a missing control it cannot see is unknown, not a finding. A public bucket address with a changing query string beside a control object, a remote migration list taken before the deploy, a count of child rows without a parent, or a message found in the recipient mailbox are returned as proven: each reads the right thing in the right place. Unfamiliar tooling never counts against a check; only the logic of what was read does.

Does it help Claude Sonnet?

By a wide margin, and on facts measured on real Cloudflare accounts. We gave twenty-three claim-and-check pairs to Sonnet and Haiku, once with this file and once without. Bare Sonnet judged 15 right. It knew that silence proves little, but it did not know that a get of a deleted R2 key can be a cached copy, that a plain 404 without cf-resized means transformations are off, that the Email Routing log is not where delivery shows, or that curl on Windows can report zero bytes for a full page. With the file it judged all 23 right. Haiku went from 15 to 23.

Share

Read this page as Markdown: /skills/cloudflare-evidence-check.md.

  • Recommended

    Done Means Done: Honest Agent Status Reports

    Agents & Protocols

    SKILL.md · v1.0.3 · 9.9 KB

    Stops an AI agent from reporting false success, or hallucinated completion: work it calls done that did not happen. Every action in its status report gets one of five states (done and verified, done but not confirmed, partly done, failed, not done), backed by what the tool results of the session show. Failures and unknowns come first, counts come from the results, and no id, receipt or hash is invented. When a result does show success, the agent says so plainly. Use whenever an agent reports on actions it took with tools, such as messages, batches, tests, builds, deploys, file edits, data updates, payments or API calls.

    $0.10once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 3 Oct 2026

  • Workers Pitfalls Review: D1, OpenNext, Fetch

    Code & Engineering

    SKILL.md · v1.2.0 · 9.8 KB

    Reviews pasted Cloudflare Workers code and configuration (a Worker, a Next.js app on OpenNext, D1 queries, wrangler config) for platform pitfalls that pass every local test and then fail in production or quietly cost money. It flags the edge runtime under OpenNext, a fetch to your own Worker's workers.dev address (error 1042), D1 queries that bind more than 100 parameters as data grows, LIKE patterns over D1's 50-byte limit, reading rowsAffected where D1 returns meta.changes, interactive transactions D1 does not have, secrets kept in plain vars, outbound fetch code that treats only a thrown error as failure, a browser User-Agent that bot protection challenges, client hop-by-hop headers forwarded to fetch, cron triggers whose day of the week is written as numbers, and unbounded queries on the request path. Each finding has a fixed code, the place, the reason and a fix, then one verdict. Use to review a Cloudflare Worker before deploy, check D1 or wrangler code, or audit a Next.js app running on Cloudflare.

    $0.05once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 8 Oct 2026

  • Traffic Reality Check

    SEO & Content

    SKILL.md · v1.0.1 · 7.0 KB

    Reads the numbers from Cloudflare Web Analytics and zone analytics (GraphQL or the dashboard) and says how many of them are real readers or real server errors, answered as JSON with the computed figures and a short reason. Knows what pollutes the table - the owner's own headless-browser checks, link-preview fetches from the social network that show up as referrals, the site's own Worker and its cache calls counted as visitors or as 5xx errors, the regional cache host of OpenNext, staging sharing a token - and what makes a true zero - the EU exclusion setting, a content security policy that allows the script but not the measurement request, sampling of data older than seven days. Use when asked how much traffic a site really has, whether a spike or an error count is real, why Web Analytics shows zero, or how to query per-address detail on the Free plan.

    $0.07once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 8 Oct 2026