A large one can, a small one only partly. Twelve short, invented product-search problems were handed to the two models, Claude Sonnet and Haiku, and nothing else came with the task. Sonnet solved all twelve. Haiku solved seven and failed five, and two of the five were the expensive kind: a query that still reads every row on each keystroke, and an accent rule that quietly damages Bulgarian. With our SQLite Search Builder skill loaded, Haiku solved all twelve under the same checks.
Method
Each case is a small search with one planted fault, written around what a shop or a help centre complains about: a plural that finds nothing (the gap that makes a forgiving search worth building), an accent that hides a product, words typed in a different order, a product code typed without its dash. Eleven cases carry a fault; the twelfth is clean, and the right answer there is that nothing needs changing. Some ask for a new search, some for a review of existing code, and a few come with a complaint from a user.
Every case went to both models once without the skill and once with it, from a terminal with tools switched off and our personal settings left out. That is 24 runs per side, plus 6 reruns after we fixed things. Each answer was scored by a pattern check: does it mention the fix the case is about? The checks were identical for both sides.
The first scoring was too strict. We read every failed answer ourselves and found that 9 checks were rejecting right answers over wording, for example a pattern wanting "aggressive" when the model had written "too broad". We widened those, for both sides alike. Requirements that belong to one approach only, such as cutting word endings from the query but never from the stored text, moved to the extra checks that apply to the skill side alone. One more thing was ours to fix: the clean case contradicted the skill's own advice, so we rewrote it and ran it again.
As a separate step we executed what the skill tells a builder to do on a real SQLite FTS5 index. That produced 70 checks, among them two consecutive result pages of 18 rows that share nothing, and a query plan free of full scans.
Results
| Model | With the skill | Without it |
|---|---|---|
| Sonnet | 12 of 12 | 12 of 12 |
| Haiku | 12 of 12 | 7 of 12 |
Sonnet did not need us. Without the skill it still picked working fixes, just different ones: a Porter stemmer, an accent setting built into the tokenizer, an exception list for words like "news". Different is not wrong, and the checks accepted all of them.
Haiku's five misses were more instructive than the number. Each is a trap worth knowing whether or not you ever use our skill:
- A substring search on the server still reads everything. Asked to fix a route that loaded a whole table, Haiku kept
LIKE '%q%'. A leading wildcard cannot use an index, so every keystroke is a full read, and a database that charges per row it reads punishes that directly. - Stripping every combining mark breaks Cyrillic. Haiku called the usual decompose-and-strip accent cleanup harmless for a Bulgarian catalogue. It is not: pulled apart, "й" is the letter "и" with a breve on top, and dropping the mark leaves a plain "и". Words that differ only by that letter collapse into one, and it offered no Bulgarian ending rules at all.
- An over-eager stemmer eats real words. It judged an aggressive stemmer fine for "universe" and "news", the very words that case was written to protect. The safer pattern is to cut endings from the typed query only, never from stored text, and to search by word start instead.
- Product codes need both spellings. For codes like AB-1042 it offered two fixes, and each left one of the two complaints unsolved.
- Changing the tokenizer without rebuilding the index. An FTS5 table cannot change its tokenizer in place. Editing the CREATE statement does nothing to a table that already exists; it has to be created again and every row inserted again.
The test also found a gap in our own skill. A code typed without its dash, "ab1042", did not find "AB-1042", because our index held only the separated pieces. Sonnet, running without the skill, was the one that pointed at it on the clean case. The skill now stores the joined form too, and the 70 checks on real FTS5 cover it.
Two limits are worth stating plainly. The skill does not correct typos, since there is no edit distance anywhere in it, and it does not rank by relevance; results come newest first. If you need either, this is not the tool.
What we did not measure
- Each model saw each case once. A pass shows it can happen, not how often.
- The twelve cases are short, invented by us, and written around what the skill does. A real catalogue has dirtier data, and real traffic would show costs that no short case can.
- Only two models took part, both from one vendor.
- The pass counts use checks shared by both sides. Against the fuller checks that follow the skill's own form, Sonnet passed 11 of 12 and Haiku 6 of 12, so with the skill Haiku often left out a requested part, such as a debounce, or the size ceiling that applies when the list stays in the browser. A pass here means the planted problem was fixed, not that the skill was followed to the letter.
- We widened 9 checks after reading failures, and a different reader might have drawn the line elsewhere.
- Nobody ran the generated code against a production database; only the skill's own steps were executed.
Read this post as Markdown: /blog/llm-search-plurals-accents.md · Atom feed.
