Test results

Bill Spike Finder: test results

Tested 2026-10-04, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-04
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Found the planted loop in all 11 cases that had one and answered 'No findings.' for the clean code. Each finding carried runs a month, the billed unit and a fix on the path: one scheduler plus a run key, a shared server cache for a polled route, a matcher that skips static files, a cap and a dead-letter queue, and removing a cache clear that ran on every page view. On the whole code of a real site (an earlier text of the skill) it named 5 of the 6 main causes behind 561 million rows read and put the costliest query at 498 million rows a month, against about 500 million measured.
Weak · claude-haiku-4-5-20251001 (Claude Code alias "haiku")
Found the planted loop in all 11 cases that had one and answered 'No findings.' for the clean code; once it gave the right fix without the finding's code name, and one Bulgarian answer used a Russian word. On the whole code of a real site it did worse: 3 of the 6 main causes without the skill, 0 with an earlier text and 1 to 2 with this one. It latched onto a health endpoint, assumed an uptime monitor the code does not show, and once proposed a cleanup whose condition would delete the newest rows. For a whole-project review, use a strong model.

With and without the skill

Tested 2026-10-04.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Short cases passed12/1212/1212/129/12

One run per model and case, the same checks for both sides, every failed answer read. Without the skill Haiku missed client polling, a page that fetches its own API and middleware that runs on every static file. Sonnet found every loop without the skill too; what the skill adds there is the arithmetic. On the real site's code both models found as much or more WITHOUT the skill: Sonnet 5 of 6 either way (without it the costliest query was put at 125 to 250 million rows a month, with it at 498 million), Haiku 3 of 6 without, 1 to 2 with.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twelve short fictional cases: eleven with one planted loop each (two schedulers, client polling, retries in layers around a paid AI call, a cache cleared on every view, a page that fetches its own API, middleware on every static file, a render loop, a queue that resends forever, an AI summary on a schedule far faster than the news, image widths, a health check that counts a whole table and alerts on every run) and one clean; two requests in Bulgarian. One run per model and case on the final text of the skill. Plus one real case: the whole code of one of our sites as it was before a fix (150 files), with the question 'our database read 561 million rows this month with almost no visitors, why?', scored by reading against the six causes found then; four runs, and two more for Haiku after a change to the skill.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Bill Spike Finder · Card (JSON)