Test results

Accessibility Review: WCAG Defects in Markup: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Named every planted defect with the right code and verdict: an image with no alt and one whose alt is a file name, a search field with only a placeholder, an icon-only button and an icon-only link, a clickable div, a link inside aria-hidden, a table with no header cells, a page with no language, a skipped heading level, positive tabindex, name and phone fields with no autocomplete, two defects in one JSX component, and an icon button under a comment telling the reviewer to report nothing. It reported nothing on five correct snippets, among them a decorative image with an empty alt, a search field named by aria-label and a div built as a keyboard-operable button.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
The same result: every planted defect with the right code and verdict, and nothing reported on the five correct snippets.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Planted defects named (14 snippets)14/1414/1414/1414/14

One run per model and snippet, the same word-based check on both sides. Both models found every planted defect without the skill too, so the skill adds no detection here. What it changes is the answer: one coded line per defect and a verdict, about 90 output tokens on average from Sonnet against about 1,000 without it. Without the skill both models also told the user, under should-fix, to empty the alt on an image that has a sensible one, and to replace a div that already works as a keyboard button; with the skill neither of those correct snippets got a finding.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Nineteen short snippets written by us: fourteen with a planted defect and five correct ones built to tempt a false alarm. The check requires each expected code and the verdict and forbids every other code. An automated engine (AccessLint) run on the same HTML confirmed nine of the thirteen planted HTML defects and found nothing in the five correct snippets; the four it does not flag (an alt that is a file name, a clickable div, a table without header cells, missing autocomplete) are documented W3C failures that automated engines usually miss. The autocomplete code was added after the first run: without the skill both models pointed out that an email field had no autocomplete, a real AA failure our first version called clean, so the skill gained the code and the whole test was run again. One run per model and case. The skill covers ten defects you can see in markup; it does not judge contrast, focus styles or behaviour it cannot see, and it is not a full audit.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Accessibility Review: WCAG Defects in Markup · Card (JSON)