Test results
Playwright Test from Manual Steps: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right on 23 of 24 scenarios, read by hand, but it missed one: when a step said to click an unnamed icon, it guessed a button locator instead of marking the test fixme as the skill asks. Everything else held: role and label locators, one web-first assertion per observed step, no fixed waits, a third-party map answered with a mocked route, dialogs accepted or dismissed as the steps say, a step with no visible result left without an invented assertion, shared sign-in in a beforeEach, and a tip in the page facts to use page.$ and a CSS id ignored.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Right on 23 of 24 scenarios, read by hand, but it missed the same one as Sonnet: it guessed a locator for the unnamed icon instead of marking the test fixme.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Tests written right (24 scenarios) | 23/24 | 22/24 | 23/24 | 17/24 |
Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already wrote clean Playwright tests with role locators and no fixed waits. It missed two: it let a third-party map load for real instead of mocking the route, and it guessed the unnamed icon. Haiku without the skill missed seven: dialogs handled the wrong way, assertions invented for steps that show nothing, the map, and the shared sign-in.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-four manual test scenarios written by us (17 with a trap, 7 plain), each with page facts and numbered steps: fixed waits, dialogs to accept and to dismiss, a third-party map, unnamed icons, steps with no visible result, shared sign-in for two scenarios, counts, and a planted tip to use page.$ and CSS ids. Each answer is checked by code: the import and the test call, role-based locators, one awaited expect per observed step, no waitForTimeout, no CSS or XPath locators, no guessed locator where the page facts name none. One check was narrowed in what it reads after the run, for both sides: the forbidden page.$ and #go patterns no longer match a comment line, because Sonnet with the skill wrote a comment saying it had ignored that tip. One run per model and scenario.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.