# Screenshot Evidence Check

Screenshot Evidence Check is a tested SKILL.md that reads a description of how a web page was captured or measured - a headless browser capture with a window size, a tall screenshot, a document width reading, a local Lighthouse run, a PageSpeed Insights run, an automation session that has been open for hours, a command-line download size on Windows - and says what that evidence proves, what it cannot prove and the one measurement that would decide, answered as JSON with a verdict; an agent buys it once for $0.03 over x402.

- Page: https://aiskills402.com/skills/screenshot-evidence-check
- Category: Code & Engineering (https://aiskills402.com/categories/code)
- Price: $0.03 once, USD-priced, paid in USDC on Base over x402. Price as loaded on this page. The 402 response your agent receives is authoritative.
- Version: 1.0.0
- Card (JSON): https://api.aiskills402.com/v1/skills/screenshot-evidence-check

## Use it when

Reads a description of how a web page was captured or measured - a headless browser capture with a window size, a tall screenshot, a document width reading, a local Lighthouse run, a PageSpeed Insights run, an automation session that has been open for hours, a command-line download size on Windows - and says what that evidence proves, what it cannot prove and the one measurement that would decide, answered as JSON with a verdict. Knows the ways the capture itself produces the symptom - headless windows under about 500 pixels that lay out wider and crop the picture, no touch emulation behind a window-size flag, blank images in very tall captures, scroll-reveal sections caught before they appear, a width reading that stays clean while a container clips, scores that move with machine load, a long automation session that stops running page scripts, a download size of zero at a successful status. Use when an agent or a person is about to report that a page is broken or fixed from a screenshot, a score or a command-line reading.

## Not for

Looking at your page or running a browser: it reads only the capture or measurement you write down, and whatever the description leaves out stays unknown to it. It does not diagnose or repair the page. Facts dated 2026-10-08; headless Chrome flags and window rules change between versions, and most measurements are the owner's own, not re-checked.

## Tested, honestly

Tested 2026-10-08.

- Strong model (claude-sonnet-5-5 (Claude Code alias "sonnet")): Right on all 23 captures, read by hand: it called the 420-pixel headless window a crop and not a phone defect, the blank hero image in a very tall capture, the pale logo that no window-size flag can test, the document-width reading beside a clipping container, the empty reveal sections, the local Lighthouse score that fell while observed paint got faster, the automation session that had stopped running scripts, the zero download size from curl on Windows, the production-only look at a media-host setting and the hydration date. It also called the sound captures proven (element measurement, hosted audit, reduced-motion capture, touch emulation with its in-page check, a normal-height hero, a steady idle Lighthouse run, a control page that worked, a staging lightbox whose image host differed from the production fallback) and did not invent doubt.
- Weak model (claude-haiku-5-5 (Claude Code alias "haiku")): Right on 22 of 23 captures, but it missed one: on the local-versus-hosted accessibility score its reply was not valid JSON (a comma where a colon belongs), and its verdict there was not-proven where the check expects contradicted. Everything else matched, including the cropped 420-pixel window, the tall capture, the zero curl size and the sound captures, the staging lightbox among them, which it called proven.

Note: Twenty-three descriptions of how a page was captured or measured, written by us (15 with a flaw in the evidence, 8 sound), answered as JSON with four keys. Each answer is scored by code: valid JSON with exactly the four keys and a fixed verdict; on the flawed ones the next check must also name a measurement that can decide. Facts re-checked in a local headless Chrome 154 on 2026-10-08: a window of 420 and 460 laid the page out at 500 and the 420 picture is a crop; hover: none is false in a headless window; scrollWidth equals clientWidth beside an overflow-x: clip child. Owner-measured and NOT re-checked: the blank images in a 9000 px tall capture, the Lighthouse spread, hosted versus local scores, the automation session that stops running scripts, the zero download size on Windows, the staging, date and client-bundle facts. Checks widened after the run, for both sides: the verdict on five setups accepts a second defensible verdict (tall capture, Lighthouse faster, automation session, 390-pixel breakpoint, everyday browser speed); the next check on the grid case accepts a capture of the commit before the grid change or a one-column view; the next check on the zero-size case accepts reading the size in the storage itself. One setup (staging lightbox) did not say what the fallback value is, so Haiku could defensibly doubt it; the input was fixed and re-run with both models on both sides, and it is counted. One run per model and case.

### With and without the skill

Tested 2026-10-08.

- Captures judged right (23 captures): Sonnet 23/23 with, 20/23 without; Haiku 22/23 with, 18/23 without.

Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already knew every flawed capture: the cropped narrow window, the tall capture, touch emulation, the clipping container, the noisy Lighthouse score, the zero size on Windows. What it lacked was trust in good evidence: it doubted three sound captures (a reduced-motion capture showing all six cards, an emulated touch session whose in-page check returned true, a control page that worked). Haiku without the skill missed five: it called the 420-pixel crop not-proven, took a production-only look for a tool artifact, doubted a sound hero image blamed the tool where the control worked, and doubted the staging lightbox although its host could only have come from the variable.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Full summary: https://aiskills402.com/skills/screenshot-evidence-check/tests

## Example

### English (claude-sonnet-5-5)

Before:

> Claim: "On small phones the main call-to-action button overflows the screen on the right." How it was looked at: a screenshot was taken with headless Chrome using --window-size=420,900. In the picture the full-width button ends exactly at the right edge of the image and its right border is missing. The paragraph above it has its normal left margin and a visible right margin of about 18 px.…

After:

> {"verdict":"measurement-artifact","proves":"At 420 px the headless window lays the page out at about 500 px and crops the picture, so the button's missing right border and the paragraph's 18 px margin are what a good page would show; the 500 px capture shows the button ending 18 px before the edge, matching the text margin.","cannot_prove":"It cannot show how the page lays out at a real phone…

## What is in the file

- The answer
- The four verdicts
- The rule before everything else
- What the capture itself can produce
- How to decide

## Languages

Any language

## How to buy

Agent (HTTP):

1. GET https://api.aiskills402.com/v1/skills/screenshot-evidence-check/file without a payment header. The answer is 402 with a PAYMENT-REQUIRED header (x402 v2): exact amount, asset, network, recipient.
2. Sign `accepts[0]` with an x402 client (for example @x402/core + @x402/evm).
3. Repeat the GET with the signature in the PAYMENT-SIGNATURE header. The answer is 200 with the file, its sha256 and a re-download token.

Agent (MCP): https://mcp.aiskills402.com/mcp — free tools search_skills, get_skill, redownload_skill. Buying itself is over HTTP.

Full flow: https://aiskills402.com/docs

## The file

- Version: 1.0.0
- Size: 9.3 KB (9478 bytes)
- SHA-256: fda0d602d4dcd672cf6c1b1ee1c3da6c809be720e552073da39063d8f3edd86b
- Updated: 2026-10-08
- New versions are free through your re-download token.

## Versions

### 1.0.0 (2026-10-08)

First release: reads a description of how a page was captured or measured (headless window size, tall capture, touch emulation, document width reading, Lighthouse runs, hosted PageSpeed Insights run, long automation session, download size on Windows, staging versus production, time-zone dependent dates, everyday browser) and answers as JSON with a verdict (proven, not-proven, measurement-artifact, contradicted), what the evidence shows, what it cannot show and the one deciding check. Carries the rule that only the description counts and unseen things are unknown.

Facts re-checked on 2026-10-08 in a local headless Chrome 154.0.8037.98 on a local page (free, read-only, no network): - window-size 420 and 460 both laid the page out at 500 (innerWidth 500, a full-width button with 18 px padding ended at x=482), and a screenshot taken with 420 was 420 pixels wide, so the picture is a crop; 500 gave 500; - the headless window reported hover: none false, hover: hover true and pointer: fine; - document scrollWidth equalled clientWidth (500) while a 900 px child sat inside an overflow-x: clip container; the element measurement (getBoundingClientRect().right above clientWidth) listed that child. Not re-checked (owner-measured 20-21 September 2026, `platform/frontend.md`): blank images in a 9000 px tall capture, reduced-motion capture of reveal animations, the 96/91/91 Lighthouse spread, hosted versus local Lighthouse, the automation session that stops running scripts, the zero download size on Windows, the staging/date/client-bundle facts. Cases: 23 (15 traps, 8 controls); the control script makes no model calls. The model test and the price check come next.

## License

Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Terms: https://aiskills402.com/docs#license

## FAQ

### Is the answer machine-readable?

One JSON object: a verdict (proven, not-proven, measurement-artifact or contradicted), a sentence on what the picture or number does show, a sentence on what it cannot show, and the single measurement that would settle the matter, written concretely enough to run today.

### Which captures does it distrust?

A headless window under about 500 pixels, a window-size flag used as if it were a phone, a very tall screenshot used to check that an image exists, a page caught before its fade-in animation, a document width reading beside a clipping container, a Lighthouse score from a busy laptop, a long remote-controlled browser session and a download size of zero on Windows. When the description shows a proper element measurement or a control that worked, it says proven.

### Will it invent problems that are not there?

It is told not to. It weighs only the text you give it, treats whatever it cannot see as unknown and never as a finding, and the test set contains sound evidence on purpose, so a good capture has to come back proven rather than suspect.

### Does it help Claude Sonnet?

Somewhat. Bare Sonnet got 20 of 23 captures right; with the skill 23 of 23. It already knew the cropped narrow window, the tall capture and the zero download size on Windows. It lacked trust in good evidence: it doubted a reduced-motion capture, an emulated touch session and a control page that worked. Haiku went from 18 to 22 of 23.

## Related skills

- [Traffic Reality Check](https://aiskills402.com/skills/traffic-reality-check.md): $0.07 once
- [Accessibility Review: WCAG Defects in Markup](https://aiskills402.com/skills/accessibility-review.md): $0.03 once
- [Bill Spike Finder](https://aiskills402.com/skills/bill-spike-finder.md): $0.05 once

## Measurement limits

- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.

Offer note: Paid in USDC (USD-pegged) over x402 by an AI agent; one-time.
