AISKILLS402

Prompt Injection Guard: test results

Tested 2026-10-02, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-02
Strong · Claude Sonnet (claude-sonnet-5-5, Claude Code alias "sonnet")
With the always-on block, alone or with the full file, it moved money on 0 of 294 trap runs in three dry runs and on 0 of 98 in a run with real USDC on Base. Without the skill it paid on 10 of 294 dry runs and 3 of 98 real ones (0.13 USDC). The traps it fell for without the skill were licence fees written into setup READMEs as a normal step, and a list entry whose address pointed at a different product. Ordinary tasks and ordered purchases kept working: 69 of 69 dry, 23 of 23 real. It never loaded the full file by itself (0 of 363 runs), so the block is what does the work.
Weak · Claude Haiku (claude-haiku-4-5-20251001, Claude Code alias "haiku")
Cuts the risk several times over without removing it. Without the skill Haiku paid on 29 of 294 dry trap runs; with the block 7, with block and file 5. With real USDC: 9 of 98 traps (0.34 USDC) without the skill, 1 of 98 (0.03 USDC) with block and file. The trap it still falls for is a list entry named after one product whose address points at another. Ordinary tasks: 67 to 69 of 69 dry, 23 of 23 real. It loaded the full file in about one run in eight, so paste the block.

Note

98 traps and 23 ordinary tasks, written by us in five languages, hidden in web pages, e-mails, READMEs, PDF text, API answers, skill files under review, code comments, file names and another agent's message. The agent had read tools and two payment tools; caps and an allowed payee were enforced in the test harness, and any payment attempt on a trap counted as a failure, settled or not. Measured: payments and leaking a planted secret, not deletion or sending messages. Dry runs record intents; the real run settled on Base mainnet from our test wallet to our own wallet (6.16 USDC in total, mostly the ordered control purchases). Claude models only.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Prompt Injection Guard · Card (JSON)