# Our skills, September 2026: measured with and without

> The eight skills we released in September, run on Claude Sonnet and Haiku with the file and without it. Five clearly helped, two changed nothing, one hurt.

Published 2026-10-08 · https://aiskills402.com/blog/our-skills-september-2026

Five of the eight helped clearly, two made no difference we could measure, and one made the weaker model worse. These are the skills we put in our catalog in the last two days of September 2026. For each one we ran the same cases twice on two Claude models, Sonnet and Haiku: once with the skill's `SKILL.md` loaded and once without it, the same request and the same checks on both sides. This page is about our skills and is written by the people who sell them, which is why the losses are listed next to the wins.

## Method

Every skill gets its own small set of cases written by us, built around the mistakes that skill is meant to prevent: a summary that drops a number, an outreach email padded past what anyone reads, a proofreader that "fixes" correct text, a payment request that should be refused. The checks are mechanical, so they count only what a script can see, such as a figure written exactly as in the source, a word limit, a banned phrase or the first word of a decision. Each model answered each case once per side. The comparison without the skill was run on 3 October 2026.

The weaker model in these runs was Claude Haiku 4.5. The Claude Code alias for Haiku has since moved to a newer model, and we have not repeated these eight tests on it, so the Haiku column describes the older one.

## Results

| Skill | What it does | Sonnet with / without | Haiku with / without |
|---|---|---|---|
| [summarize](/skills/summarize) | shortens a text and keeps every figure and condition | 6 / 0 of 6 | 6 / 0 of 6 |
| [humanize](/skills/humanize) | rewrites stiff prose without changing what it claims | 5 / 0 of 5 | 5 / 0 of 5 |
| [cold-email](/skills/cold-email) | writes a short first email from the facts you give | 6 / 0 of 6 | 6 / 1 of 6 |
| [x402-seller](/skills/x402-seller) | finds why a paid API does not sell or is not listed | 6 / 2 of 6 | 6 / 0 of 6 |
| [proofread](/skills/proofread) | fixes spelling and grammar, leaves correct text alone | 6 / 4 of 6 | 5 / 1 of 6 |
| [code-review](/skills/code-review) | reviews a pasted change for bugs and security holes | 6 / 6 of 6 | 6 / 6 of 6 |
| [seo-meta](/skills/seo-meta) | writes title tags and meta descriptions inside the limits | 6 / 6 of 6 | 6 / 6 of 6 |
| [x402-buyer](/skills/x402-buyer) | decides whether an agent should pay a request | 7 / 7 of 7 | 5 / 7 of 7 |

## Where the file made the difference

**Summaries kept their numbers.** Without the skill, both models wrote longer summaries, some of them lost figures, and Haiku answered in English for three of the six texts that were not in English. Part of the gap is our strictness: the check wants each figure in exactly the form the source used, so a decimal comma where the source had a point is scored as a miss. Even read by hand, the side without the skill dropped facts the other side kept.

**Rewrites stopped making claims firmer.** The humanize check also looks for exact words, which makes the gap look wider than it is. Reading the answers, the pattern is real: without the skill, words like "usually" vanished and statements came out more certain than the original, and Haiku turned Bulgarian, Russian, Spanish and German texts into English because the request was in English.

**Cold emails got short.** With the skill the drafts ran from 57 to 102 words; without it, from 137 to 297. Sonnet's content was as good either way, so here the skill mostly buys length discipline. Haiku without it also reached for a stock opening line, and in one Bulgarian draft used the familiar form of address where a stranger expects the polite one.

**Sellers learned which field was wrong.** Each case pasted the reply of a paid API with one planted defect. Without the skill, Sonnet found the problem but not the exact field to change in four of six, and Haiku missed all six, including a missing payment header and text decoded as the wrong character set. With it, both named the field every time.

**Proofreading left correct text alone.** With no rules, Sonnet changed a text that was already correct, and Haiku, for four of the six texts, returned only its corrections as a list, not the finished text, and wrote that list in English even for Bulgarian, German and Spanish. With the file loaded, Sonnet got six of six right and Haiku five of six.

## Where it changed nothing

**Code review and SEO meta: both models were already right.** On our six code changes with planted bugs, both found every bug without help. On six pages that needed a title and a description, both kept to the language of each page and inside the character limits without help. We added five harder cases to each. Sonnet stayed right without the skill. Haiku let through a defect that silently drops orders, describing it as a sensible fallback; with the review skill it flagged the defect, though its proposed fix still dropped the order. On the harder SEO pages Haiku stretched the facts without the skill and stuck to the given ones on three of the four with it. So the honest summary is: no gain on Sonnet, a partial one on a weaker model, and our standard cases were too easy to show it.

## Where it made the weaker model worse

**The x402 buyer skill cost Haiku two simple cases.** Without the file, both models made the right call on all seven simple payment requests. With it, Haiku made the right decisions but failed two cases on wording: once it refused without naming the field that was wrong, and once it answered partly in Russian. On harder cases the skill earned its place: without it, both models were ready to settle a request in a token made to look like USDC, and Haiku signed off on a product other than the one requested. Still, on the cases that come up most, Haiku did better without it, and the skill's listing says so.

## What the whole catalog looks like

These eight are the first part of what we sell. Every skill in [the catalog](/skills) carries the same two-sided test on its own page, including where it did not help, and the October skills will be in next month's edition once October has ended.

## What we did not measure

- **Single attempts per side.** A case that flips between runs is noise; the large gaps (six to nothing) are not.
- **Our own cases, a handful per skill.** They are built around known mistakes, which favours a skill written against those mistakes.
- **Mechanical checks.** They count exact words, numbers and limits. They cannot see a summary that keeps every number and still misleads, or an email that is short and wrong.
- **An older Haiku.** The weaker model here is Haiku 4.5; the newer one behind the same alias has not been tested on these eight.
- **No other vendors.** Two Claude models only.
