# Summarizing with an LLM: what it adds and drops

> Claude Sonnet and Haiku summarized six short texts in five languages from a plain request. Without rules they ran long, switched language and bent conditions.

Published 2026-10-07 · https://aiskills402.com/blog/llm-summary-adds-and-drops

Partly. The names and numbers mostly come through; the length, the language and some of the conditions do not. Each of two Claude models, Sonnet and Haiku, got six short texts, two of them English and one each in Bulgarian, German, Russian and Spanish, plus the bare request "Summarize this text." Sonnet answered in the language of each text, but its summaries ran to between 68 and 89 percent of the original, and twice it appended a remark of its own about today's date. Haiku replied in English to three of the five texts that were not English, turned a condition about selling two ferries inside out, and described a price rise of about 56 percent as a near doubling. After we added our rules as a skill, both models kept the source language every time and wrote between 28 and 44 percent of the original. Haiku still reshaped some wording with the rules in place, and our checks did not notice.

## Method

We wrote six invented news items of 241 to 345 words. A ferry operator dropping a night route (English), a village school that will close unless its roof gets funded (Bulgarian), a district heating price rise (German), an olive oil cooperative's results for the year (Spanish), a clinic moving to online booking (Russian), and an English review of a bridge repair that disagrees with itself about dates, cost and injuries. Each item is dense with deadlines, hedges, conditions and claims that one party makes about another.

The side without the skill received the line "Summarize this text." and then the item. The other side had our [summarize skill](/skills/summarize) as its instructions and was asked to apply it to the same item. Both models ran through Claude Code's non-interactive mode, tools off and none of our own settings loaded, with one answer for each pairing of model, item and side, so 24 answers are compared. An earlier baseline had sent the items with no request at all; a model that is never told what to do is not a fair opponent, so those answers were discarded and the plain request was run instead.

A script judged both sides with the same checks. It looked for one or two key figures per item written exactly as in the text, for framing words such as "overall" or "the article", for a length of no more than 60 percent of the original, for the output language, and for pairs: when a figure shows up, the condition the text attaches to it has to show up too.

## Results

| Model and measure | With the skill | Without it |
|---|---|---|
| Sonnet, answers that passed every check | all six | none |
| Haiku, answers that passed every check | all six | none |
| Sonnet, length as a share of the original | 28 to 44 % | 68 to 89 % |
| Haiku, length as a share of the original | 29 to 41 % | 36 to 71 % |
| Haiku, written in the source language | 6 of 6 | 3 of 6 |

**The zero overstates the gap.** Our figure check wants the exact wording of the text. Without rules Sonnet wrote £1.9m where the text said 1.9 million, 7,5 % for 7,5 Prozent and 31,2 M€ for 31,2 millones, and each of those counted as a lost figure although nothing was lost. Its two failed condition checks were our own fault: the patterns expected the German singular Mietvertrag and the Spanish stem previst and met the plural Mietverträge and the noun previsión instead; the conditions themselves were there. Read for meaning, Sonnet without rules was mostly faithful. Its problems were size and additions.

**No length, no limit.** With only the plain request, every Sonnet summary was between 68 and 89 percent as long as its source, and the Bulgarian one came with headings and bullet lists. Haiku ranged from 36 to 71 percent. A summary that size saves a reader little time, and it leaves room to repeat every clause, which is part of why Sonnet's conditions survived. With the skill, Sonnet wrote 96 to 107 words and Haiku 86 to 102.

**Haiku slid into English.** For the Bulgarian, German and Russian items Haiku wrote its summary in English, each under an English heading; the Spanish item got a Spanish summary. The request was written in English, which is the likely pull. Sonnet answered all five non-English items in their own language. With the skill loaded, both models used the source language in all six cases.

**Conditions bent, and a sum appeared.** Without rules, Haiku changed meaning in ways no pattern of ours was built to see. In the ferry item the two ships are sold unless someone takes over the whole route; Haiku wrote that the company was "offering to sell" the ships to anyone who wanted them, which is a different deal. The cooperative pays 1,2 million euros to its members unless the assembly on 28 March chooses the new mill instead; Haiku made that a plan to pay out or to invest, with neither the assembly nor the date. The litre price went from 3,10 to 4,85 euros, a rise of about 56 percent, and Haiku called it "se duplicó casi", close to double. The German cartel office reviews such files "in der Regel" within eight weeks, and the hedge vanished from Haiku's English. Its Bulgarian summary closed with a figure of its own: the school's fate, it said, hangs on "the remaining ~700,000 leva", an amount the item never states.

**Sonnet added material of its own.** Two of its plain summaries ended with a note that the deadlines in the item had already passed, counted from the day of the run, and that it could not say what happened afterwards. That can be handy, yet the source never says it, and nothing marks the remark as separate from the summary. For the bridge review Sonnet closed with a line headed "Overall" that passed judgement on the whole report, and Haiku closed the same case with a verdict of its own. A summary that ends in an opinion invites people to quote the opinion.

**With the rules, Haiku still reshapes wording.** The skill settled length and language for both models. Sonnet's summaries with it kept who said what in every case we read, and each of the six ended with a short closing line naming the topics it had left out. Haiku passed every check as well, yet the text tells another story. In the ferry item it stated as fact that the island council was not consulted, which is the council's own claim, and presented the chief executive's view that the subsidy covered about a third of the loss as if the item said so. It kept the study's estimate of 8 to 11 million pounds a year and lost the authors' warning that the figure turns on what hauliers decide; our pattern accepted the word "estimated" and passed it. In Bulgarian it wrote that the children will move to the other school, without the "if it closes". In Spanish, a payout that happens unless the assembly decides otherwise became an assembly that will choose between paying and investing. It never wrote the closing line, in any of the six.

## If a model summarizes for you

- **Say how long.** A request with no length lets "summary" mean most of the text.
- **Say which language when the request and the text differ.** An English request on top of a Bulgarian, German or Russian item pulled the smaller model into English three times out of five.
- **Ask for who said what to stay attached, then check it yourself.** Our pattern checks passed answers that had lost a condition. For a contract, a medical letter or a financial notice, read the summary against the source before acting on it.
- **Ask for no commentary.** Remarks about today's date and closing verdicts are easy to forbid and easy to miss once they are in.

## What we did not measure

- **Each pairing of model, item and side ran once.** A repeat could come out differently, Haiku's most of all, so treat any one case that flips as noise.
- **Six short items we wrote ourselves, each in one language.** Between 241 and 345 words; no long reports, meeting transcripts, email threads or mixed-language documents.
- **Some answers with the skill come from reruns.** On the day the skill was written, three cases were run again; where a case ran more than once the comparison counts the last answer, and for Sonnet's Spanish item that was the third. We judged the earlier misses to be faults in our checks, not in the summaries, but that judgement is ours.
- **The plain request was kept plain on purpose.** One line that names a length and the output language might do much of the skill's work; that request was not part of this test.
- **Patterns miss reworded meaning.** Every bent condition above was found by reading, not by the script, so the column with the skill may hide changes we did not catch.
- **Only Sonnet and Haiku.** Both are Claude models; nothing from other vendors and no other sizes were tried.
