Test results
SQLite Search Builder: test results
Tested 2026-10-04, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-04
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Fixed the planted problem in all 11 cases that had one and answered 'No findings.' for the clean search. It moved LIKE '%q%' and a load-the-whole-table route onto an FTS5 index with prefix search and a LIMIT, cut endings from the query only, added a dictionary for irregular plurals, and kept the breve of 'й' in a Bulgarian catalogue. On the product-code case it also stored the code without its dash, so 'ab1042' finds 'AB-1042', and it noticed that one complaint in our case could not happen with the code we showed. In 1 of 12 answers it left out the 3-letter prefix rule, which the skill asks for.
- Weak · claude-haiku-4-5-20251001 (Claude Code alias "haiku")
- Fixed the planted problem in all 11 cases that had one. It followed the skill's form less closely: five of its answers each left out one part the skill asks for (a debounce, the 2,000-item limit of the in-memory mode, cutting endings from the query only, the 3-letter prefix rule, or storing dictionary forms), and once it described a stem wrongly ('tomat' for 'tomatoes'), though the prefix still matched. On the clean search it answered 'No findings.' once; in a second run it reported a cursor pattern that cannot overflow in practice. With Haiku, check the generated code against the skill's test table.
With and without the skill
Tested 2026-10-04.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Cases passed | 12/12 | 12/12 | 12/12 | 7/12 |
One run per model and case, the same checks for both sides; we read every failed answer and widened nine checks that rejected right answers for their wording. Without the skill Sonnet gave other working fixes (a Porter tokenizer, remove_diacritics 2, an exception list). Haiku without it kept LIKE '%q%' on the server, called the loss of 'й' in Bulgarian harmless and gave no Bulgarian rules, offered two fixes for product codes that each left one complaint unsolved, said an over-eager stemmer was fine for 'universe' and 'news', and changed the tokenizer without rebuilding the table.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twelve fictional search codes: eleven with one planted problem each (an in-memory search, LIKE '%q%' on every keystroke on D1, plurals in FTS5, word order, accents, irregular plurals, an over-eager stemmer, product codes with dashes, a Bulgarian catalogue, cleaning on one side only, a route that loads the whole table) and one clean. One run per model and case, six cases run again after a fix. The steps in the skill were run on real FTS5 (node:sqlite): 70 checks that take the rules, the SQL and the examples from the skill text, two pages of 18 with no repeats, and a query plan with no full scan. The test found one gap in the skill itself, a code typed without its dash; it was fixed and checked again.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.