Test results

Spreadsheet Formula: Excel and Google Sheets: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Wrote a correct formula for all 21 tasks, each as the single line asked for, computed by a real spreadsheet engine on the sample and copied down where the task said so: sums and counts with boundaries, exact lookups in unsorted codes, a not-found text, tiers, a month filter that checks the year, an average where zero counts and blank does not, a running total, working days with holidays, a two-key lookup, semicolons and decimal commas for a Bulgarian sheet, a Missing line when the price column was absent, distinct customers with a blank cell, a latest date or none, days overdue against one fixed cell, a first name from a one-word name, and an either-or count. It ignored a cell telling it to answer =0.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Also 21 of 21, with the same kinds of formula as Sonnet.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Correct formula, one line (21 tasks)21/2119/2121/2121/21

The same request and the same check on both sides. Without the skill Haiku got all 21, so for Haiku the skill adds nothing measurable here. Sonnet without it wrote a correct formula in every case but twice added text: a correction and a repeat after the Bulgarian-locale formula, and a remark about the cell that told it to answer =0. On content, both sides were right almost everywhere; neither used functions that only the newest Excel has. The gain is the strict one-line answer, and it rests on few cases.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-one tasks written by us, each a description plus a sample sheet; the answer is evaluated by HyperFormula 3.4 on that sheet and every target cell is compared with the expected value. The first fourteen tasks turned out easy: in a first run both models solved them with and without the skill, so seven harder ones were added and everything was run again. The skill itself was changed twice after its own failures on these tasks: it did not say that decimals are written with a comma where semicolons separate the arguments, and its own code formatting led the models to wrap the formula in backticks. The result shown is the fourth run with the skill on the same tasks, while the side without the skill was run once, so treat the difference as small. One run per model and task in each version; earlier runs are kept in the test folder. Two gaps of the engine were fixed in our checker, not in the skill: bare TRUE and FALSE, and array arithmetic for SUMPRODUCT.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Spreadsheet Formula: Excel and Google Sheets · Card (JSON)