For problems that are visible in the markup itself, yes, and even a small model finds them. We planted 14 accessibility defects in short pieces of HTML and JSX, from an image with no alt text to a link hidden from screen readers while still reachable by keyboard. Claude Sonnet and Claude Haiku found every one of them with no instructions at all. The difference our checklist made was not what the models found but what they wrote around it: without it, a one-line fix arrived inside a review of about a thousand tokens, and correct markup got "should fix" advice too. The models without the checklist also caught a real failure our own first version missed.
Method
We wrote 19 short fragments. Fourteen had a defect. Five were about images and names: a picture with no alt attribute, another whose alt repeated its own file name, a search box whose only hint was placeholder text, and a button and a link that held nothing but an icon. Three were about the keyboard: a div that responds to clicks alone, a link sitting inside a block marked aria-hidden, and tab stops numbered above zero. Four were about structure and forms: a price list laid out as a table without th cells, a full page whose html tag carries no lang, headings that go from h2 straight to h4, and a checkout form asking for a name and a phone number without autocomplete. The last two mixed things up: a JSX signup component with two problems, and a close button under an HTML comment claiming the page had already passed an audit. The other five fragments were correct and easy to over-review: a newsletter form with a proper label, a divider image whose empty alt is right, a search box named with aria-label, a div given a button role, a tab stop and key handling, and a plain contact paragraph.
Each defect maps to a criterion in the WCAG 2.2 quick reference, or is marked as best practice where no criterion fails, as with skipped heading levels. We also ran the HTML through an automated checker. It agreed on nine of the thirteen HTML defects and stayed silent on the correct fragments. What it did not report: the alt text that merely repeats the file name, the div that reacts to clicks, the table with no header row and the fields without autocomplete. Each of those needs someone who understands what the element is for.
Both sides got the same request: review the markup before release. One side also had our accessibility review skill, with ten codes and a verdict. Because the side without it does not use codes, the planted defects were scored on both sides by word patterns; the correct snippets could only be scored strictly on the side with the skill. Every fragment went to each model one time per side.
Results
| Checklist loaded | No checklist | |
|---|---|---|
| Sonnet: planted defects found | 14 of 14 | 14 of 14 |
| Haiku: planted defects found | 14 of 14 | 14 of 14 |
| Sonnet: average answer, output tokens | about 90 | about 1,000 |
Nothing slipped past either model. The comment asking for a clean report changed nothing, the file name used as alt text was called out, and the hidden link was explained correctly: a keyboard user can reach it while a screen reader says nothing. If your question is whether a current model knows these rules, it does.
The fix was there, under a lot of other text. A typical answer without the checklist listed what was right, then the real defect, then several items to check in CSS, then optional polish, then suggested markup. All of it was reasonable. But a review that runs before every release has to be acted on, and a reader who meets the same contrast and focus reminders on every snippet stops reading at the third one.
Correct markup still got homework. On the newsletter form, both models without the checklist told us to empty the alt text on an envelope image because the heading already said enough, and put it first on their list of fixes. On the div that already worked as a keyboard button, both recommended replacing it before release. Neither is a failure of any criterion. With the checklist, both snippets came back with no findings.
Our first version was missing a rule. In the first run, the checklist had no code for autocomplete, and it called the original form clean. Without the checklist, each model noted that an email field collecting the person's own address had no autocomplete attribute, which fails criterion 1.3.5 at level AA. They were right. We added the code, gave the form the attribute, added a separate case for it, and ran the whole test again. The earlier run stays in the test folder.
What we did not measure
- No repeats. Each fragment was reviewed once per side, and a repeat could change a borderline result.
- Our own small fragments. Real pages mix templates, scripts and styles, and many defects only appear once they render.
- Anything you need a browser for. Contrast, focus visibility, zoom and what a screen reader actually announces were out of scope for both sides.
- False alarms without the checklist. Free-text advice is hard to score fairly, so that side was scored on the planted defects only; the homework above is from reading the answers.
- Two Claude models. No other vendors, and no testing with disabled users, which no review of markup replaces.
Read this post as Markdown: /blog/llm-accessibility-review-buried-fixes.md · Atom feed.
