# Test Guard Review: Can This Check Fail?

Test Guard Review: Can This Check Fail? is a tested SKILL.md that reviews a pasted test, check script, smoke test, mutation script or CI step and says whether it can fail when the thing it guards is broken, so a green result is worth something; an agent buys it once for $0.03 over x402.

- Page: https://aiskills402.com/skills/test-guard-review
- Category: Code & Engineering (https://aiskills402.com/categories/code)
- Price: $0.03 once, USD-priced, paid in USDC on Base over x402. Price as loaded on this page. The 402 response your agent receives is authoritative.
- Version: 1.0.1
- Card (JSON): https://api.aiskills402.com/v1/skills/test-guard-review

## Use it when

Reviews a pasted test, check script, smoke test, mutation script or CI step and says whether it can fail when the thing it guards is broken, so a green result is worth something. It finds a completeness guard that counts its own hard-coded list instead of the directory, a one-star glob that never enters subfolders, a mutation script that restores with git checkout and wipes the uncommitted fix under test, a null test that an empty string passes, a status 200 that is only the login page, a wait loop without sleep that measures network speed instead of time, a test that calls the pure function directly and never the call site, a counterfactual control whose data does not create the condition, a fix copied into two branches and tested in one, a limit tested only with ordinary values, and a catch that turns a failure into success. Each finding has a fixed code, the place, why it passes and the one input that would make it fail, or exactly No findings. Use to review a test that cannot fail, a check script that always passes, or a guard before you trust it.

## Not for

Writing tests, measuring coverage, or judging style and speed: it only asks whether the pasted check can go red when the guarded thing is broken. It reads what you paste, so a handler, a folder layout or a config that the paste does not show is treated as unknown, and it does not run anything.

## Tested, honestly

Tested 2026-10-08.

- Strong model (claude-sonnet-5-5 (Claude Code alias "sonnet")): Right on all 24 scripts, read by hand: a completeness guard that counts its own list (JavaScript and Python), a one-star glob that skips subfolders, a mutation script that restores with git checkout, a null test on a field that comes back empty, a curl that follows the redirect to a login page, two polling loops without a sleep that measure the network instead of the time, a test that calls the pure function and bypasses the handler, a control that cannot create its condition, a ceiling never tried with an absurd value and a fix tested on one of two mirrored branches; it named the input that would turn each one red and answered No findings. on all nine sound scripts.
- Weak model (claude-haiku-5-5 (Claude Code alias "haiku")): Right on all 24 scripts, read by hand, with the same findings as Sonnet. On sound scripts it was stricter than our key three times, and each point was fair: a mutation run that exited 0 with a surviving mutant, a wait loop that reported live when the expected version was empty (we fixed both scripts and re-ran them), and a vitest run that never starts being counted as a kill.

Note: Twenty-four test and check scripts in JavaScript, Python, Bash and PowerShell written by us: 15 with a way to pass while the guarded code is broken and 9 sound ones. With the skill each answer is scored by code on the finding codes and the verdict; without it the same request is scored on the concept in any words. On the sound scripts the bare side has no check, so the counts rest on the 15 faulty scripts. The rules come from cases measured on our own projects in August and September 2026 (owner-measured, not re-checked). Checks widened after the run, for both sides, each because a right answer was refused: the self-counting list accepts "a seventh module", "stays" and "scanning src"; the login redirect accepts "redirects it to /login"; the bypassed handler accepts "never touches the handler"; the mirror accepts "only exercises runNow". Two sound scripts had real flaws (found by Haiku) and were fixed and re-run with the skill. One run per model and script.

### With and without the skill

Tested 2026-10-08.

- Scripts reviewed right (24 scripts): Sonnet 24/24 with, 22/24 without; Haiku 24/24 with, 16/24 without.

Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already caught the self-counting list, the one-star glob, the git checkout restore, the empty-string test, the login redirect, the bypassed handler, the control without its condition and the untested mirror. It missed both polling loops: a loop that counts tries without sleeping finishes in seconds and reports minutes, and it looked elsewhere in both scripts. Without the skill it also raised concerns on most sound scripts, which is not counted. Haiku without the skill missed eight, among them the git checkout restore, the empty string, the login page and the absurd value.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Full summary: https://aiskills402.com/skills/test-guard-review/tests

## Example

### English (claude-sonnet-5-5)

Before:

> // test/guard-coverage.test.ts import { describe, expect, it } from "vitest"; import { readFileSync } from "node:fs"; // Every module under src/ that reads the database must go through safeQuery(). const DB_MODULES = [   "src/api/posts.ts",   "src/api/users.ts",   "src/api/comments.ts",   "src/jobs/cleanup.ts",   "src/jobs/digest.ts",   "src/lib/search.ts", ];…

After:

> [LIST-COUNTS-ITSELF] test/guard-coverage.test.ts (DB_MODULES and the "lists all six modules" test): the module list and its length of 6 are both written in the test, so the app can gain a seventh database module that bypasses safeQuery() and both the list and the count stay unchanged and green.…

## What is in the file

- The answer
- The codes
- Rules
- Work in this order
- Short example

## Languages

Any language

## How to buy

Agent (HTTP):

1. GET https://api.aiskills402.com/v1/skills/test-guard-review/file without a payment header. The answer is 402 with a PAYMENT-REQUIRED header (x402 v2): exact amount, asset, network, recipient.
2. Sign `accepts[0]` with an x402 client (for example @x402/core + @x402/evm).
3. Repeat the GET with the signature in the PAYMENT-SIGNATURE header. The answer is 200 with the file, its sha256 and a re-download token.

Agent (MCP): https://mcp.aiskills402.com/mcp — free tools search_skills, get_skill, redownload_skill. Buying itself is over HTTP.

Full flow: https://aiskills402.com/docs

## The file

- Version: 1.0.1
- Size: 10.3 KB (10527 bytes)
- SHA-256: 8e94f13edac1493ccb0a568e28d41bb1ec92f5e0eb1f61a70bb5abd03ae54833
- Updated: 2026-10-08
- New versions are free through your re-download token.

## Versions

### 1.0.1 (2026-10-08)

- Measured on 24 scripts: Sonnet 22 -> 24 of 24, Haiku 16 -> 24 of 24 (without -> with the skill). - Checks widened (both sides), each after a right answer was refused: list-counts (seventh module, stays, scanning src, derive from the directory), 200-is-login (redirects it to), seam-bypassed (never touches the handler), mirror-untested (only exercises runNow). - ok-copy-restore now exits non-zero on a surviving mutant or a missing anchor; ok-loop-clock refuses an empty EXPECTED. Both re-run with the skill. - Price: $0.03 (Sonnet gain 2).

### 1.0.0 (2026-10-08)

First release: reviews a pasted test, check script, smoke test, mutation script or CI step and lists each way it can stay green while the guarded thing is broken. Eleven fixed codes: a completeness guard that counts its own list, a one-star glob, a mutation script restored with git checkout, a null check an empty string passes, a status 200 that is the login page, a wait loop without sleep, a test that bypasses the call site, a control whose data does not create its condition, a fix tested in one of two branches, a limit never tested past its value, a catch that turns failure into success. Each finding names the one input that would make the check fail; a check that can fall gets exactly `No findings.`

The skill carries from the start the rule that only what the paste shows is reported; code that is not shown is unknown, not a finding.

Facts are incident facts from the owner's own platform notes (17 August and 11 September 2026), not vendor facts, so they are not re-checkable by a fetch and carry no "checked on" line. The arithmetic behind the skew example and the controls was re-checked on 2026-10-08 by running them (see notes/facts-2026-10-08.md).

Tests: 24 snippets (15 with one planted weakness, 9 correct ones). No model run yet; the price starts at $0.03 (class B) and is set after the baseline.

## License

Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Terms: https://aiskills402.com/docs#license

## FAQ

### Eleven accidental passes: which are they?

Eleven: a completeness guard that counts its own list, a one-star glob that skips subfolders, a mutation script that restores with git checkout and wipes the fix under test, a null check that an empty string passes, a status 200 that is the login page, a wait loop without sleep, a test that calls the pure function and never the call site, a control whose data never creates the condition, a fix tested in only one of two branches, a limit tested only with ordinary values, and a catch that turns failure into success. Each is a way a real check once stayed green over a broken system.

### Will it invent problems in a good test?

It is told not to. Every code has a not-a-finding condition: a recursive glob with a subfolder fixture, a restore from a copy with a checksum compare, a request that carries the session, a loop with a pause and a clock, a test that drives the handler, a test that goes past the limit. Code the paste does not show is unknown, never a finding. A weakness that fits none of the eleven codes is left out on purpose, so a style complaint or a missing unrelated test never appears in the answer.

### How is it different from a skill that writes tests?

tests-from-spec turns a written spec into test cases. This one starts from a test that already exists and asks the opposite question: if the system were broken right now, would this still pass? Use them one after the other. It also pairs with done-means-done, which asks what finished means; this one asks whether the evidence of finishing could have said no.

### Does it help Claude Sonnet?

One blind spot, two scripts. Across twenty-four scripts, plain Sonnet already saw most ways a test stays green while the code is broken. The two it missed were polling loops without a sleep, which finish in seconds and report minutes: 22 right without the file, 24 with it. It also raised doubts about most of the sound scripts when the file was not loaded. Haiku, on the other hand, went from 16 to 24.

## Related skills

- [Tests from Spec: Cases from the Requirements](https://aiskills402.com/skills/tests-from-spec.md): $0.03 once
- [Done Means Done: Honest Agent Status Reports](https://aiskills402.com/skills/done-means-done.md): $0.10 once
- [Code Review](https://aiskills402.com/skills/code-review.md): $0.01 once

## Measurement limits

- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.

Offer note: Paid in USDC (USD-pegged) over x402 by an AI agent; one-time.
