# Length limit or every fact: which one the AI breaks

> We had Claude cut texts to a limit and turn spec sheets into shop copy. Bare, it overran limits and added praise; with rules, it lost facts or broke a count.

Published 2026-10-09 · https://aiskills402.com/blog/llm-length-limit-or-facts

It gives up one of them, and nothing in the answer tells you which. We asked Claude Sonnet to cut 23 short texts down to a stated number of words or characters, and to turn 24 product spec sheets into shop copy with fixed counts: a title, a short text, three to five bullets. Given only the request, Sonnet kept the facts and broke the length limit: nine of its shortened texts ran past it. With a written set of rules loaded, it never ran over, but five shortened texts lost a date or a figure, and two product answers ran to six bullets where five was the cap. The invented praise you might expect from a model writing shop copy did disappear: on the bare request, 19 of its 24 product texts carried a word or claim the sheet does not support, such as lightweight, durable or quiet, and with the rules, none did.

## Method

We wrote both sets of cases ourselves, mostly in English, a few in Bulgarian, German and Spanish.

**Shortening.** There were 23 short news items and notices, each with a limit stated in the request: between 16 and 70 words, or between 160 and 280 characters for four of them. The request also told the model how to count: "e-bike" makes one word, while "1 250" makes two. Sixteen texts carried a trap, such as a negation, a quotation, a product code, a planted line telling the editor to drop all numbers, or a text that already fitted. The other seven were plain items with several figures each. A script counted every answer by the request's own definition, looked for the dates, figures and codes that had to survive, and refused hedges, ellipses and numbers absent from the source.

**Shop copy.** There were 24 spec sheets for small household goods, from a kettle to a backpack. The request asked for JSON: a title of at most 70 characters, a short text of at most 160 characters, between three and five bullets, a long text of at most 120 words, and an issues list for contradictions and for rows that give orders to whoever writes the copy. It required every figure to come from the sheet and did not forbid praise. Code compared each figure with the sheet, checked the counts and the issues, and searched for praise and for properties the sheet never states.

Two models took part, Claude Sonnet and Claude Haiku, each called once per case and per side, without tools and without our personal configuration. On the bare side a model saw nothing but the request; on the other, our written rules for the task sat in front of it as the system prompt. A person then read every miss.

## Results

| Claude Sonnet | bare request | with the rules |
|---|---|---|
| Shortened texts over the limit | 9 of 23 | 0 of 23 |
| Shortened texts that lost a date, figure or qualifier, read by hand | 1 of 23 | 5 of 23 |
| Product answers with a word or claim the sheet does not back | 19 of 24 | 0 of 24 |
| Product answers that broke a count, bullets or characters | 0 of 24 | 2 of 24 |

**Bare, the count gave way.** Sonnet's short texts read well and kept their facts, yet nine overran, by one to eight words, although the request spelled out how to count. The 16-word case came back at 19 words with every figure and date intact. The one loss we found by reading was a qualifier: "at least 20 minutes before departure" became "20 minutes early". Haiku overran on eight texts; the worst were 41 words for a limit of 16, and 200 characters for a limit of 160.

**With the rules, the facts gave way.** No answer ran over. Instead, five lost something they had to keep: a watch's launch date, the deadline and income threshold of a German heating subsidy, the cost of a Spanish study room, an e-bike's price and model code, and the five questions for installers in a guide blurb. Four of the five were plain texts, not traps. Each time the opening survived almost word for word, and a later sentence or clause was cut whole, taking its figure along. Three of the answers stopped four to seven words short of the limit, and the missing date or price needed about five. Haiku made the same trade three times.

**In shop copy the trade ran the other way.** On the bare request Sonnet kept every count: four or five bullets each time, no short text above 147 characters. What it added was sales talk: 19 of the 24 answers held a word or claim the sheet does not back, mostly adjectives such as lightweight, durable and reliable, but also quiet for a fan whose sheet gives only 38 dB, easy cleaning for a hand blender, and swimming for a tracker rated 5 ATM. Figures were rarely wrong: it used one of two capacities that disagreed (while naming the clash under issues), dropped an "up to" from one title, and summed "3 outside, 1 inside" into four pockets.

With the rules, none of Sonnet's 24 answers held a word the sheet does not back. Both misses were counts: six bullets where five was the ceiling. The kettle sheet lists six properties and each got its own bullet; the lamp got a bullet holding just its name, then one per row. Our rules want each bullet to carry one fact, so a sheet with six rows left no way to satisfy both. Haiku did the same on two sheets, and on two more it packed nearly every row into a short text of 176 and 179 characters against the limit of 160.

The shop-copy test has two published totals. Counting every unsupported word, Sonnet went from 5 of 24 to 22 and Haiku from 15 to 19. Counting only figures, conflicts, invented properties and format, the stricter measure, Sonnet went from 15 to 22 and Haiku from 18 to 19. In the shortening test Sonnet went from 10 of 23 to 18, and Haiku from 6 to 20. The rules we tested are published as the [shorten-to-limit](/skills/shorten-to-limit) and [product-description](/skills/product-description) skills.

## Where the limits come from

A count in a prompt usually stands in for a field somewhere else, and that field does its own counting. X allows 280 characters in a post, but each link counts as 23 characters whatever its real length, and each emoji as two ([X developer documentation on counting characters](https://docs.x.com/resources/fundamentals/counting-characters)). Google Merchant Center caps a product title at 150 characters and wants the description to hold only information about the product, with no promotional text ([Google Merchant Center product data specification](https://support.google.com/merchants/answer/7052112)). The model's sense of length is not the field's.

## What to check when a model writes to a length limit

- **Count in code, by the receiving system's rules.** Sonnet overran a limit whose unit the request had defined. The same gap showed up when we tested [meta descriptions](/blog/llm-meta-description-stretches).
- **Keep a must-keep list and test each item as a string.** Dates, prices, codes, and qualifiers such as "up to" or "at least", written in the source's own local format: 30. November, 260.000. A check that only counts would have passed all five answers that lost a figure.
- **Treat spare room plus a missing fact as an error.** An answer several words under the limit that dropped a date cut more than it needed to.
- **Decide which rule wins, and put that in the prompt.** If six rows must fit in five bullets, allow related rows to share one, or allow a sixth. If the facts cannot fit, ask for a short note saying so rather than a silent cut.

## What we did not measure

- **Each case was answered once per side.** Asking again might change any one verdict. Nine overruns against none is a clear direction; two sixth bullets across 24 sheets is thin.
- **Our cases, mostly in English.** Of the 23 texts, 19 are English, and so are 18 of the 24 sheets. We wrote them knowing where the traps were.
- **The must-keep lists are ours.** In the German heating notice, Sonnet kept the manager's reason for the price rise and dropped the subsidy deadline and the income limit. Some editors would choose the same; our check calls it a miss.
- **Where praise ends and an invented property begins is our call.** The stricter total of 15 of 24 for bare Sonnet comes from reading by hand; split by code on figures and conflicting values alone, only two of its bare answers fail. While writing this post we found that the published test notes described some misses in words the recorded answers did not bear out, such as a speaker called waterproof where the answer said splash resistance. The notes were corrected on 9 October 2026; no total changed.
- **The bare request allowed praise.** It asked only for figures from the sheet, so part of the bare column measures a rule that side never had.
- **Haiku's counting did not improve in shop copy.** It broke a count three times on the bare request and four times with the rules.
- **We counted words between spaces and characters one per code point.** We did not try any platform's own counter, such as the link and emoji weights on X.
- **Two Claude models**, and nothing from other vendors.
