Test results

SQLite Query From a Question for D1: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on all 24 questions in substance, read by hand: every query ran in SQLite and returned the expected rows, including pairs from a self-join without duplicates, the top rows per group, ties, NULLs, dates stored as text, money in cents and case-insensitive matches; where the tables hold no column for the question it declined to guess. On one such question it declined in its own words instead of the fixed CANNOT ANSWER line the skill asks for, which a program reading the answer would not recognise.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on all 24 questions in substance, read by hand, with the same queries and refusals as Sonnet; on one refusal it also used its own words instead of the fixed CANNOT ANSWER line.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Questions answered right (24 questions)24/2422/2424/2423/24

Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already wrote correct queries for ties, NULLs, text dates, cents and case, and declined the questions the tables cannot answer. It missed two: on the self-join it wrote the word Wait in the middle of its SQL and corrected itself in place, so the query does not run, and on top N per group it used a column named id that the table does not have. Haiku without the skill missed the self-join.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-four questions written by us, each with the CREATE TABLE statements and a small data set (17 with a trap, 7 plain): self-joins, top N per group, ties, NULL handling, text dates, cents, LIKE and case, a question in another language, a planted comment and three questions the tables cannot answer. Each answer is run in a fresh SQLite database with the data and its rows are compared with the expected ones; a question with no column to answer it must be declined. The fixed CANNOT ANSWER form is checked on the skill side; in the comparison both sides are scored on declining in any clear words. No check was widened. One run per model and question.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to SQLite Query From a Question for D1 · Card (JSON)