Mostly, yes, and the exception is the one that hurts. Thirteen D1 migration tasks went to Claude Sonnet and Claude Haiku, and every answer was applied to a database full of rows. Without help, each model got eleven right. Both missed the same one: making a column required on a table that another table points to. Both wrote the textbook rebuild, where a new table is filled and then renamed into the old one's place while foreign keys are deferred. Cloudflare D1 rejects that migration. We know because our own first version of the skill taught exactly the same rebuild, and the test caught it.

Method

The thirteen tasks are the changes a real project makes: a new status column that a job will poll, a counter that must stay correct, a required column on a table with rows, an index for a slow query, an email that may appear once among rows that are not deleted, a team that every project must belong to, removing a column that has an index on it, turning euro amounts into integer cents, two columns that must appear together, and tag names that ignore letter case. One schema held a comment telling whoever writes migrations to drop the audit log every time. Two of the tasks were written in Bulgarian.

Each answer ran in SQLite against our schema and its rows, inside one transaction with foreign keys switched on, which is how D1 applies a migration file. Then it was judged by what happened next: inserts that ought to work, inserts that ought to be refused, the state of the rows, and the plan SQLite picked for the slow query. Before trusting that setup, we created two throwaway remote D1 databases and tried the tricky statements on them directly. In every probe, local SQLite and remote D1 gave the same answer.

The request was identical on the two sides: the target is D1, and it already holds data. One side also had our D1 migration skill loaded. Any Markdown fence around the SQL was stripped from both sides first. Every task ran once per model and side.

Results

Using the skill Plain request
Sonnet: migration applies, data right 13 of 13 11 of 13
Haiku: migration applies, data right 13 of 13 11 of 13

The rebuild D1 refuses. SQLite cannot add NOT NULL to an existing column, so the table has to be built again. With PRAGMA defer_foreign_keys on, both models created users_new, copied the rows, dropped users and renamed the new table. On remote D1 that file stops with FOREIGN KEY constraint failed. What passed was keeping the name: park the rows in a spare table, drop the original, build it afresh with its old name and copy the rows back. The orders that point to those users then find them, and the check D1 postponed to the end finds nothing wrong.

The closing pragma that switches the check off. Cloudflare's documentation shows a migration that defers foreign keys first and turns the setting off again on its last line. Our first probe ended that way, and the renamed rebuild passed. So we tried a migration that was plainly broken: defer, drop a table that three comments point to, then turn the setting off. D1 applied it and left three orphaned rows behind. Turning it off does not run the postponed check; it throws it away. Without that last line, the same broken file was refused, as it should be.

How we found it. Our first draft of the skill taught the rename. With it loaded, both models wrote the rename without the closing line, and our checker failed them. We assumed the checker was wrong and took that exact file to remote D1. D1 refused it too. The skill now teaches the rebuild that keeps the name and warns against the closing pragma, and we ran all thirteen tasks again with it loaded.

Two more misses without the skill. Sonnet obeyed the comment in the schema and added a statement that deleted the audit log together with its rows. Haiku put BEGIN TRANSACTION and COMMIT around one change; remote D1 refuses those statements in a migration (error 7500), since it already treats the whole file as one unit.

What both models got right on their own. Given a new column for when an article was announced, they stamped the existing articles as done, so the job's first pass would not flood the channels with the archive. They put the equality column first in the new index and kept the old index. They made the unique email index partial with the very predicate the upsert used, and they rounded before converting euros to cents. Those were the traps we expected to catch them, and none did.

A note on triggers. On 30 September a remote D1 refused one of our own migrations, which held a trigger, with "incomplete input". On 8 October the same trigger text, with the same wrangler version, applied without complaint, alone and next to other statements. We do not know what changed. What we do now is apply every migration to a remote database before deploying the code that depends on it.

What we did not measure

  • Two versions, unequal runs. The side with the skill ran twice because its first version was wrong; the side without it ran once.
  • Our own thirteen tasks. They are small schemas with a few rows; real migrations touch tables with views, many triggers and millions of rows, where time limits matter.
  • Remote D1 on two days, with one wrangler version. D1 is a hosted service and its behaviour can change. The trigger refusal we saw only once and never again is the reminder.
  • The checker is SQLite, not D1. It matched remote D1 in every probe we ran, but we probed only the statements these tasks needed.
  • Only Claude. Sonnet and Haiku; no model from another company.

Read this post as Markdown: /blog/llm-d1-migrations-rebuild-fails.md · Atom feed.