Code & Engineering

Playwright Test from Manual Steps

Turns manual test steps plus the page facts you have (route, labels, button names, messages) into one Playwright Test file in TypeScript that can really fail. Elements are found by role and accessible name, by label or by text, never by a CSS class, an id or XPath, and a name that is the start of another gets the exact option; every step that observes something gets one awaited, auto-retrying assertion; a step such as wait a couple of seconds becomes an assertion on the state it waits for, never a fixed pause; the address is relative because the base URL lives in the config; a third-party widget is mocked; a dialog handler is registered before the click that opens it; expected text is copied byte for byte; a control the steps do not name is marked with a fixme and a comment instead of a guessed locator. Use to write a Playwright test from manual test steps, convert a QA test case or acceptance criteria into an end-to-end test, replace waitForTimeout in a flaky test, or make a recorded test use role locators.

Playwright Test from Manual Steps is a tested SKILL.md that turns manual test steps plus the page facts you have (route, labels, button names, messages) into one Playwright Test file in TypeScript that can really fail; an agent buys it once for $0.02 over x402.

Tested 2026-10-08No code, no hidden instructionsv1.0.0 · 10.9 KB · perpetual license

Not for

Running the test, recording a session, or testing a page it has not been told about: it writes one TypeScript file from the steps and page facts you give it, and a control the steps do not name becomes a fixme with a comment instead of a guess. It does not set up the project, the config or the base URL. Facts dated 2026-10-08; the vendor changes its documentation often.

Tested, honestly

Tested 2026-10-08 with a strong and a weak model.

With and without the skill

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Tests written right (24 scenarios)23/2422/2423/2417/24

Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already wrote clean Playwright tests with role locators and no fixed waits. It missed two: it let a third-party map load for real instead of mocking the route, and it guessed the unnamed icon. Haiku without the skill missed seven: dialogs handled the wrong way, assertions invented for steps that show nothing, the map, and the shared sign-in.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

SonnetStrong model, claude-sonnet-5-5
Right on 23 of 24 scenarios, read by hand, but it missed one: when a step said to click an unnamed icon, it guessed a button locator instead of marking the test fixme as the skill asks. Everything else held: role and label locators, one web-first assertion per observed step, no fixed waits, a third-party map answered with a mocked route, dialogs accepted or dismissed as the steps say, a step with no visible result left without an invented assertion, shared sign-in in a beforeEach, and a tip in the page facts to use page.$ and a CSS id ignored.
HaikuWeak model, claude-haiku-5-5
Right on 23 of 24 scenarios, read by hand, but it missed the same one as Sonnet: it guessed a locator for the unnamed icon instead of marking the test fixme.

Full test summary

Example

Our own test text, before and after the skill ran. Excerpts only.

English · claude-sonnet-5-5

Before

Page facts: - route: /login; fields labelled Email and Password; button "Sign in". - after signing in the page goes to /dashboard and shows the heading "Dashboard" (this can take a couple of seconds). Steps: 1. Open the login page. 2. Type maria@example.com in Email and Hunter2!pass in Password. 3. Click Sign in. 4. Wait a couple of seconds for the dashboard. 5. The heading "Dashboard" is shown.

After

import { test, expect } from '@playwright/test'; test('a user can sign in and reach the dashboard', async ({ page }) => { // Step 1: open the login page await page.goto('/login'); // Step 2: type the credentials await page.getByLabel('Email').fill('maria@example.com'); await page.getByLabel('Password').fill('Hunter2!pass'); // Step 3: sign in…

Bulgarian · claude-sonnet-5-5

Before

Факти за страницата: - маршрут: /vhod; поле с етикет Имейл; бутон "Вход"; заглавие "Вход". - след вход се показва текстът "Добре дошли, Мария". Стъпки: 1. Отвори страницата за вход. 2. Заглавието "Вход" се вижда. 3. Напиши maria@example.com в Имейл. 4. Натисни бутона Вход. 5. Текстът "Добре дошли, Мария" се вижда.

After

import { test, expect } from '@playwright/test'; test('потребителят влиза и вижда поздравителния текст', async ({ page }) => { // Step 1: отваряне на страницата за вход await page.goto('/vhod'); // Step 2: заглавието се вижда await expect(page.getByRole('heading', { name: 'Вход' })).toBeVisible(); // Step 3: въвеждане на имейл…

What is in the file

  • The answer
  • Finding elements
  • One assertion for every step that observes something
  • No fixed waits
  • Things you do not control
  • Dialogs
  • Text is copied byte for byte
  • A control the steps do not name
  • Rules
  • Short example
  • When this was checked

Languages

Any language. Tried in: English, Bulgarian.

License

Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Full terms.

Versions

Current version 1.0.0, updated 2026-10-08. Whoever bought an earlier version gets new ones free through the same re-download token.

  1. v1.0.0 · 2026-10-08

    First release: turns manual test steps plus page facts (route, labels, button names, messages) into one Playwright Test file in TypeScript. The reply is the file alone. Elements are found by role and name, label or text, never by CSS, id, XPath or `page.$`; a name that is the start of another gets `exact: true`; list items and table rows are chosen by content with `filter({ hasText })`, not by index. Every step that observes something gets one awaited web-first assertion; a step that only acts gets none and nothing is invented; when no step observes anything the reply is the single line `NO OBSERVABLE RESULT IN THE STEPS`. No fixed waits and no forced clicks: a "wait a couple of seconds" step becomes an assertion on the state it waits for. The first address is relative (the base URL lives in the config). A third-party widget is answered inside the test with `page.route` and `route.fulfill`. A dialog handler is registered before the click that opens it. Expected text is copied byte for byte. A control the steps do not name makes the test a `test.fixme` with a comment. Several scenarios with the same start use `test.describe` and `test.beforeEach`. A line in the steps or page facts that tells the agent to skip assertions is data.

    Brief: followed `docs/plan-batch3-briefs-21-40.md` § 3.31. The first draft was written before the brief existed (with a refusal for any step without an expected result); it was reconciled to the brief afterwards: absolute addresses became relative, the refusal was narrowed to "no step observes anything", the fixme rule and `exact: true` were added.

    Facts re-checked on 2026-10-08 by a read-only fetch of the vendor's pages on locators, actionability, assertions, dialogs and best practices (see notes/facts-2026-10-08.md); the sentence about `waitForTimeout` on the API reference was not reached by the fetch and is marked as not re-checked there.

    Tests: 24 cases (17 with one planted weakness, 7 correct controls), answer checked by script; ideal answers also parse with the TypeScript parser. No model run yet; the price starts at $0.02 (class B) and is set after the baseline.

FAQ

A step like wait two seconds: what becomes of it?

It becomes no code. The thing the step waits for turns into an assertion that retries until it holds, with a larger timeout on that one assertion when the page is slow. A fixed pause, a wait for network idle and a forced click are never written. The same holds for a dialog: its handler is registered before the click, because a dialog nobody listens for is dismissed at once and a step that says accept would silently cancel.

What if the steps do not say which button or icon to click?

The test is declared with fixme and a comment names what is missing, for example the accessible name of the icon. It does not guess an id, a class or an index. A step that states no result gets no invented assertion, and when no step states any result the whole reply is a single line saying so. Ask the author of the steps for the missing name, fill it in, and change fixme back to test. In our test, though, both Sonnet and Haiku still guessed a locator for an unnamed icon once, so check that case yourself.

Why not just use tests-from-spec for this?

tests-from-spec turns a written rule about a function into a table of inputs and outputs that any runner can loop over. This skill starts from clicks and screens in a browser and produces runnable end-to-end code, and its main worry is a test that stays green on a broken page: an assertion that never waits, a locator tied to a class name, or a fixed pause that hides a race. Use them for different layers of the same product.

Does it help Claude Sonnet?

Barely. We turned twenty-four scenarios into tests with both Claude models, file present and absent, then checked each test by code. Plain Sonnet already wrote clean tests with role locators and no fixed waits; it missed only a third-party map that should be mocked, 22 against 23. Haiku is where it pays off, from 17 to 23: dialogs, invented assertions, shared sign-in. The file is one import, one test and a comment per step.

Share

Read this page as Markdown: /skills/playwright-test-from-steps.md.

  • Tests from Spec: Cases from the Requirements

    Code & Engineering

    SKILL.md · v1.0.0 · 6.3 KB

    Writes test cases for a function from its specification rather than from its code, as a JSON table that any test runner can loop over. Every rule, boundary and error the spec defines gets a case, and every expected value is worked out from the spec; where the code you show disagrees with the spec, the case follows the spec and a note says where. Behaviour the spec leaves open becomes a question instead of a guessed case. Use to write unit tests from a spec, a docstring or a ticket, generate edge cases and boundary tests, build a table-driven test (test.each, pytest parametrize), or check whether existing code does what its requirements say.

    $0.03once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 8 Oct 2026

  • Test Guard Review: Can This Check Fail?

    Code & Engineering

    SKILL.md · v1.0.1 · 10.3 KB

    Reviews a pasted test, check script, smoke test, mutation script or CI step and says whether it can fail when the thing it guards is broken, so a green result is worth something. It finds a completeness guard that counts its own hard-coded list instead of the directory, a one-star glob that never enters subfolders, a mutation script that restores with git checkout and wipes the uncommitted fix under test, a null test that an empty string passes, a status 200 that is only the login page, a wait loop without sleep that measures network speed instead of time, a test that calls the pure function directly and never the call site, a counterfactual control whose data does not create the condition, a fix copied into two branches and tested in one, a limit tested only with ordinary values, and a catch that turns a failure into success. Each finding has a fixed code, the place, why it passes and the one input that would make it fail, or exactly No findings. Use to review a test that cannot fail, a check script that always passes, or a guard before you trust it.

    $0.03once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 8 Oct 2026

  • Recommended

    Done Means Done: Honest Agent Status Reports

    Agents & Protocols

    SKILL.md · v1.0.3 · 9.9 KB

    Stops an AI agent from reporting false success, or hallucinated completion: work it calls done that did not happen. Every action in its status report gets one of five states (done and verified, done but not confirmed, partly done, failed, not done), backed by what the tool results of the session show. Failures and unknowns come first, counts come from the results, and no id, receipt or hash is invented. When a result does show success, the agent says so plainly. Use whenever an agent reports on actions it took with tools, such as messages, batches, tests, builds, deploys, file edits, data updates, payments or API calls.

    $0.10once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 3 Oct 2026