Yes, if the model is a strong one and gets the whole code at once. Claude Sonnet named five of the six causes we found by hand in July, with or without our bill-spike-finder skill. Claude Haiku named three without the skill and fewer with it. That does not match what the owner of the site remembers, so this report also says why the two experiences may differ.
Method
In July 2026 one of our own sites, a crypto news tracker, read 561 million rows in a single month. At the time it ran on Next.js with Vercel and a Turso database; it has since moved to Cloudflare, like all our sites, and everything below is about the code as it was in July. The database was 48 MB and almost nobody visited. Turso blocked reads and the site went down. The owner recalls that the model used at the time did not find the causes alone and had to be steered, one hint after another.
We rebuilt the situation. The input was the site's complete text code from a git commit of 4 July, before any fix: 150 files, about 89 thousand tokens, with images, lock files and dependencies left out. After the code came the question a person would type: 561 million rows, a 48 MB database, almost no visitors, why? There were no tools and no live usage data, only the code. One run per model and side.
The answer key comes from the July fixes. Six causes mattered:
- Two schedulers, one in the platform config and one external and described only in a project doc, each firing every job.
- A six-hour average of trading volume, re-read from raw rows every minute. This one was worth roughly 500 million rows a month.
- A 24-hour chart rebuilt from raw rows on every page load, around 11,500 rows each time.
- A latest-value query that grouped the whole table with no
WHERE. - A table of price snapshots, about 12 thousand new rows a day, that nothing ever deleted.
- An always-on metric stored under an hourly key, so active rows kept multiplying (237 instead of 6).
We read each answer against that list and marked a cause found only if the answer named it, not merely the area.
Results
| Sonnet without the skill | Sonnet with it (first text) | Haiku without | Haiku with (first text) | |
|---|---|---|---|---|
| Main causes found, of 6 | 5 and a half | 5 and a half | 3 | 0 |
Sonnet found the same causes on both sides: two schedulers, the volume average, the chart, the ungrouped query and the missing deletion. The sixth, the hourly key, it covered only in part. So the July experience did not reproduce here. Two honest reasons are possible. The model is newer than the one the owner used. And the owner's session searched the repository with tools, a piece at a time, while this test pasted everything in front of the model at once. We cannot separate those two.
What the skill changed for Sonnet was the arithmetic. With it, the costliest query was sized at about 498 million rows a month, against roughly 500 million measured. Without it the estimate was 125 to 250 million, half or less. With the skill the answer also added up its findings to roughly 65 million rows a day with no visitors, which reaches 561 million in about nine days, and it spelled out what two schedulers do: every write twice, the digest email twice. Without the skill the two schedulers were only a "may".
Haiku went the other way. Without the skill it found three causes, though it sized the volume average at 4.1 million a month, thirty times too small. With the first text of the skill it found none. It fixed on a health endpoint and imagined an outside uptime monitor calling it every minute, a caller that appears nowhere in the code. It also judged the main table about fifty times smaller than it is: 200 rows a day instead of 11,520. Its total, 349 million, came from that invented monitor.
We then changed the skill, once. The new rule asks the model to total its findings and hold the sum against the reported figure, to size tables from the write rate it can read in the code, to treat a caller the code never shows as an assumption that never ranks first, and to look for schedulers in the docs too. The example numbers in the rule differ from this case on purpose, so they cannot leak the answer. Haiku ran twice on the new text.
| Run 1 | Run 2 | |
|---|---|---|
| Main causes, of 6 | 1.5 | about 2 |
| Table size | right | right |
| Invented caller | yes, a monitor | no |
| Missing deletion found | no | yes, but the cleanup condition would delete the newest rows |
Better than zero, still under the three it scored with no help. We stopped there and say on the skill page that a whole-project review needs a strong model.
The contrast is the lesson. On twelve short cases, one planted loop in each, Haiku passed all 12 with the skill and 9 without; Sonnet passed 12 either way. The same text that lifts a small model on a small piece of code pulled it off course on 150 files, where one tempting line in a checklist was enough to start a wrong story.
Useful even if you never touch our skill:
- Rows read are the meter, not traffic. Almost no visitors says nothing about the bill.
- With no visitors, list what runs on its own first, every scheduler, including ones named only in a doc.
- Size each table from how fast the code writes to it, not from how big it looks.
- Total your findings and hold the sum against the bill before trusting the diagnosis. A story that covers only a sliver of the number is not the cause.
What we did not measure
- Single runs on the real case, one per model and side; two for Haiku after the change. Haiku wavered between runs, so its 0 against 3 may be partly luck.
- One real project. A different codebase may favour a different model.
- The answer key comes from our own fixes. A cause we never found is not in it, and a model that found one would score nothing for it.
- Only two models, from one vendor. Sonnet was not rerun on the changed skill text.
- No live usage data was given to either model, only code. Had the models seen the platform's ranking of the queries that read the most rows, the costliest causes would likely stand out at once; we did not test that.
- Scoring was by reading, by us, with a half point for partial hits. Another reader could score a point or so differently.
Read this post as Markdown: /blog/llm-561-million-rows.md · Atom feed.
