SQL Query Optimizer for D1, SQLite and Turso: test results
Tested 2026-10-03, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-03
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Named the planted problem with the right code in all 11 cases that had one and answered 'No findings.' for the clean schema. Its fixes were working SQL: an index on (scope, created_at) for the gap, a partial index on status = 'pending' for the empty poll, a range on the raw timestamp, keyset pagination, chunks of 50 for D1's 100-parameter limit, a backfill in the same migration, the partial index's WHERE copied into ON CONFLICT. At first 2 of 12 were scored as failures because our pattern accepted only one of the two fixes the skill names; Sonnet chose the other (a latest-prices table kept by the writer; day bounds computed in SQL), both were confirmed with EXPLAIN QUERY PLAN, and the same answers rescored at 12 of 12.
- Weak · claude-haiku-4-5-20251001 (Claude Code alias "haiku")
- Found the planted problem with the right code and a working fix in all 11 cases that had one, including the Bulgarian request. On the clean schema it invented a finding under a code of its own and proposed DELETE ... LIMIT, which standard SQLite rejects as a syntax error. With Haiku, run any suggested statement before relying on it.
Note
Twelve fictional schemas and queries: eleven with one planted problem each (aggregate scan, index gap, NOT IN, OR on a flag, empty queue poll, function on a column, prefix LIKE, OFFSET with COUNT, D1 parameter limit, flag column without backfill, ON CONFLICT against a partial index) and one clean; one request in Bulgarian. Every planted plan and every accepted fix was confirmed with EXPLAIN QUERY PLAN on SQLite 3.53 without ANALYZE. Machine checks look for the finding code and a working fix; one run per model and case. The rescoring of stored answers after two patterns were widened is in results-2026-10-03-recheck.md.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.
Back to SQL Query Optimizer for D1, SQLite and Turso · Card (JSON)