AISKILLS402

Testing models honestly

Eight skills, two models: where the weaker one fails

We ran eight SKILL.md files on a strong and a weak model. Both mostly passed our checks; the weak one failed in ways checks cannot see, and used more tokens.

Georgi Kalchev5 min readreport
Two robot testers check the same stack of cards: one bench glows steady blue, the other amber with a cracked card

Mostly not where you would look. On our eight skills the weaker model passed the mechanical checks about as often as the stronger one. What it got wrong was quieter: claims nobody gave it, conditions rewritten into something else, a formal address dropped, a Bulgarian misspelling left in place every time, severity inflated. It also spent several times more output tokens on the same task, and most of those were reasoning tokens.

Method

The two models are the ones printed in every report: claude-sonnet-5-5, which we call the strong one, and claude-haiku-4-5-20251001, the weak one. Each case was sent through the command line with no tools at all and without any personal instructions; the skill file itself was given as the system prompt and the test text came in on standard input. The reports print this command at the top. Every run behind this post was made on 30 September 2026. The folder of the humanize skill also holds reports from the day before, for earlier versions of that skill, and we do not use them here.

The cases are ours, written for each skill: five for humanize, six each for code-review, seo-meta, proofread, cold-email, summarize and x402-seller, and seven for x402-buyer, which is 48 cases in total.

Each case carries mechanical checks: the answer exists, the language is right, names and figures we listed are kept, phrases we banned are absent, and the length stays within a ratio. We then read the outputs ourselves and wrote a verdict per model, which is where the failures below come from. Nothing here was judged by a third model.

One run per case and model in each round. Several skills needed more than one round, because between rounds we changed the skill wording or repaired our own checks. The reports show it: code-review two full rounds, proofread three, cold-email two, x402-seller two, x402-buyer three, and summarize four full rounds followed by four single-case re-runs.

Results

The table shows the last full round of each skill: how many cases each model passed, and what reading the outputs added.

Skill Strong Weak What the weak one did
humanize 5 of 5 5 of 5 cut the text hard and added closing remarks nobody asked for
code-review 6 of 6 6 of 6 rated severity too high and hedged; once described a loop bug inaccurately
seo-meta 6 of 6 6 of 6 plainer wording, awkward Bulgarian, filler on the thin page
proofread 5 of 6 5 of 6 missed Bulgarian misspellings in every round
cold-email 6 of 6 6 of 6 invented claims and flattery; ignored formal address in three languages
summarize 3 of 6 6 of 6 reshaped a condition and dropped a hedge
x402-seller 6 of 6 6 of 6 showed a code snippet against the rules; over-promised listing time
x402-buyer 7 of 7 5 of 7 right decision every time; one answer read as Russian, one refusal did not name the payee field

Read the table with care. The two summarize numbers say the weak model did better, and that is an artefact: our checks asked for particular figures to be kept, while the strong model cuts length by dropping entire facts and listing what it dropped in a last line. Its red marks were our check disagreeing with a design choice, and the later single-case re-runs passed after we changed the limits.

The first rounds of two skills were red for both models for reasons of our own making. All 12 cold-email runs failed the first round because our length check was wrong, and all 12 summarize runs failed their first round too, mostly on length limits and kept-figure lists that we were still tuning. A red mark is not evidence about a model until somebody has read why it is red.

Where the weak model failed

The clearest failure is Bulgarian proofreading. In three rounds the weak model never corrected the word for schools and, in two of the three, also left a misspelled word for a meeting untouched; the strong model fixed all planted errors every time. Cold email is the most worrying one for anyone who sends what the model writes: the weak model made up statements absent from the brief, for example a claim about the customers it supposedly works with, and on the thin brief did not admit that it was a first contact.

Summaries were subtler: one condition was turned from "unless the assembly decides otherwise" into "the assembly will decide", and a hedged cause was stated as a fact. Checks that only look for names and numbers do not see either. For the buyer skill the decision was right in all 21 weak-model runs, although one Bulgarian answer read as Russian to our language check and one refusal did not name the field that held the wrong payee.

Where the strong model slipped too

It added a mild closing remark of its own to the English and German humanize outputs, and its Bulgarian and Spanish outputs came out at roughly 55-60% of the original length. In proofread, one Spanish text came back at 1.32 times the input length against a limit of 1.3, because it flagged a real ambiguity in our test text. Three of its summaries ran slightly over the default limit of 100 words.

Tokens

In the last round of every skill the weak model produced more output tokens than the strong one, between 1.3 and 7.9 times more. The per-case figures in the reports were added up by us, so the ratios are our arithmetic. The largest gap is humanize, where the weak model's five outputs total 70832 tokens against 8939, and the usage lines show that 70093 of the weak model's tokens were reasoning. A cheaper model that thinks this long is not automatically a cheaper choice.

What we did not measure

We did not measure variance. Every case was run once per model per round, and the proofread rounds already disagreed with each other on borderline cases, so any single mark is indicative only. The cases are short texts we wrote ourselves, not real workloads, and they total 48. We did not test multi-file code changes, real wallet loops or long documents. The checks are mechanical and cannot see an invented claim or a changed meaning; those findings rest on our own reading, and nobody else has checked that reading.

The skills were tuned on these same cases. Rules were added after a failure showed up, so pass rates for the later rounds are friendlier than a fresh set of cases would give. We compared one model pair on one day. We did not measure time or price, only the token counts printed in the usage lines, and we did not check whether a different model pair, or the next version of either model, behaves the same way.

Read this post as Markdown: /blog/weak-model-eight-skills.md · Atom feed.