Yes, for the everyday kind. We handed 21 spreadsheet tasks to two models, Claude Sonnet and Claude Haiku, each a short description plus a few rows of the sheet, and every answer was run in a formula engine on that sample. In the final run, with our instructions loaded, neither model got a single one of the 21 wrong. Without them Haiku also got all 21, and Sonnet got 19, losing the other two only by adding a sentence after a correct formula. The finding we did not expect was about us: on the way to that result, two of the skill's own wordings caused failures that the models did not make without it.

Method

The tasks are the ones people actually ask for. A sum and a count with a boundary, a price looked up by code in an unsorted list, the same lookup printing a fixed text when the code is not there, commission tiers, a month total that has to check the year as well, an average that counts a zero but skips an empty cell, a total that runs down the column, the working days between two dates with holidays, a lookup on two keys at once, a count of cells that contain a word, the number of different customers in an Excel 2019 workbook with a blank cell in the list, the latest order date or nothing, days overdue measured against one fixed date cell, a first name from a name that may be a single word, and a count of rows in either of two regions.

Three tasks test something other than formula knowledge. One sample has a Notes cell with a line addressed to the assistant, telling it the correct answer is =0. One is a file in Google Sheets set to Bulgarian, where numbers carry a decimal comma and arguments are separated by semicolons. One leaves out the price column the formula needs; there, the right answer is a single line beginning Missing: that names the gap.

With or without the skill, the wording of the request already said what we wanted: the formula alone, on one line, starting with =, no explanation. One side also had our spreadsheet formula skill loaded. Each answer went into HyperFormula, an open-source formula engine, at the cell the task named; where the task said the formula would be copied down, it was copied down and every row had to match its expected value. A Markdown fence around the whole answer was removed on both sides before scoring.

Our checker had two gaps of its own, and both showed up when answers we knew to be correct failed. The engine wants TRUE() where Excel accepts a bare TRUE, so an exact lookup ending in FALSE failed. And with array arithmetic switched off, a SUMPRODUCT over two comparisons counted one row instead of three. Both were fixed in the checker, not in what the models wrote.

Results

Skill on Skill off
Sonnet: correct formula, one line 21 of 21 19 of 21
Haiku: correct formula, one line 21 of 21 21 of 21

The first fourteen tasks were too easy. In the first round, before the harder ones were added, every answer on both sides was a correct formula. The only failure was Sonnet without the skill adding a remark under the formula on the task with the planted note. A test that nothing fails measures nothing, so we wrote seven harder tasks and ran everything again.

Version one broke the Bulgarian sheet. Our first wording said that a person who writes decimal commas needs semicolons between arguments. It said nothing about how to write the numbers. With that rule loaded, Haiku answered in the US form, =ROUND(B2*1.2,2), which the engine set to the same locale could not even parse. Without the skill, Haiku had written =ROUND(B2*1,2;2), which is exactly right for that sheet. In the same run Haiku also put one formula inside backticks, which turned out to be the next problem.

Version two taught the models backticks. The fix for the decimal comma went in, along with a sentence on cells that give orders. In the next run three answers came back as a formula inside backticks, one from Sonnet and two from Haiku. Every example in the skill was written in code formatting, and the models copied the format along with the content. Pasted into a cell, a formula with a backtick in front of it is just text. We took all code formatting out of the skill, wrote the examples as plain lines, and the next run with the skill had no failures.

Without the skill, Sonnet explains itself. Both of its misses in the final comparison were correct formulas followed by text. On the Bulgarian sheet it wrote the formula, then a line headed "Correction", then the same formula again. On the planted note it wrote the right sum, then explained that the cell telling it to answer =0 was data and had been ignored. Ignoring the cell was correct; saying so inside the answer is what someone pasting into a cell does not want. Haiku without the skill gave the bare formula every time.

What we did not measure

  • The skill was tuned on this test. Both fixes came from failures on these same tasks, and the side with the skill was run four times while the side without it was run once. A fresh set of tasks could show a different picture.
  • An engine, not Excel. HyperFormula evaluates the common functions the way Excel does, but not every function. In one earlier run Haiku used AVERAGEIFS, which the engine lacks, and that answer counted as failed. A correct formula the engine cannot run scores the same as a wrong one.
  • Small, clean samples. A few rows each. Real sheets have merged cells, numbers stored as text and dates that are not dates.
  • For Haiku, no gain. It solved all of them without the skill. For Sonnet the gain is two answers, and both are about extra text, not the formula.
  • Two Claude models, one run per side in the final comparison.

Read this post as Markdown: /blog/llm-spreadsheet-formulas-skill-was-bug.md · Atom feed.