# Humanizing AI text with an LLM: what goes missing

> We asked Claude Sonnet and Haiku to humanize one text in five languages. Without rules, all ten answers lost the word usually, and some invented facts.

Published 2026-10-07 · https://aiskills402.com/blog/llm-humanize-drops-usually

Not reliably, if all you send is the request. Claude Sonnet and Claude Haiku each received one short piece of AI-sounding text in five languages, with nothing more than the line "Humanize this text." In all ten of them the word "usually" was gone, so a meeting the company usually holds once a year turned into one it always holds. Sonnet's answers stated facts the input never gave, and Haiku's added motives nobody had claimed and came back in English for three of the four non-English texts. Once the same models had our rules for this task loaded as a skill, all ten answers kept that qualifier, stayed in the original language and returned nothing but the rewrite, though both models still cut more than the rules allow.

## Method

We wrote one paragraph of about a hundred words about an invented software company and produced it in English plus four translations: Bulgarian, Russian, Spanish and German. It reads like a generated press blurb on purpose: an opener about the fast-paced digital world, "moreover" and "it is important to note", clauses that comment on the fact just before them, and a final line calling the company a testament to remote work. Under the padding sit the facts a rewrite must not touch:

- the year the company was founded, a headcount of 85 and a presence in 12 countries;
- a yearly in-person meeting that the company *usually* holds;
- work from any time zone, *unless* the job involves customer support hours;
- a home office allowance of 150 euros a month.

Each model got the same one-line English request, "Humanize this text.", followed by the text. On one side that was all. On the other, our [humanize skill](/skills/humanize) was loaded as the system prompt. Both models ran through the Claude Code CLI with no tools and none of our usual configuration, once per model, language and side: 20 runs in all. The skill side was recorded on 30 September and the bare side on 3 October.

A script checked every answer for five things. The numbers and the company name must still be there. The words for "usually" and "unless" in that language must still be there. Three stock phrases, the opener, "in conclusion" and "testament", must be gone. The length must fall between 0.4 and 1.1 times the input. And the answer must be in the original language. The word checks match whole words exactly, and a synonym fails them; for that reason we also read every answer, and most of what follows comes from that reading.

## Results

| Model and check | With the skill | Without it |
|---|---|---|
| Sonnet, every check passed | 5 of 5 | 0 of 5 |
| Haiku, every check passed | 5 of 5 | 0 of 5 |
| "Usually" or its translation kept, both models | 10 of 10 | 0 of 10 |
| Reply in the original language, Haiku | 5 of 5 | 2 of 5 |
| Answer held only the rewrite, both models | 10 of 10 | 3 of 10 |

**"Usually" was the first word to go.** The input says the company usually meets in person once a year. Without rules, all ten answers turned that into a fixed event. Sonnet's English version simply has the whole company meeting every year, and the other languages and the other model say the same. The difference is small on the page and large for a reader: the rewrite now promises a meeting every year, and anyone who knows a year was skipped will call the text wrong. The exception for support roles fared better. Sonnet kept it in all five languages, in its own words in three of them, and Haiku kept it every time, three times in English. With the skill, both models kept the word for "usually" and the condition in every language.

**The rewrite filled gaps with facts of its own.** The input never says how many of the staff work remotely. Sonnet's answers settled that anyway, and differently in each language: "almost everyone" in English, "most" in German, "all" in Russian, and in Bulgarian a company that had worked fully remotely since the day it was founded. Four answers gave four versions of a fact nobody supplied. Haiku leaned towards motives instead. Its English answer says the allowance exists "because they actually care", its answer to the Spanish text says the company cares about people and "not just the bottom line", and its German answer credits clear communication for the arrangement working, which the input never mentions. In its answer to the Bulgarian text, Haiku also turned the exact headcount into "about 85 people".

**Three answers came back in English.** The request was in English and four of the texts were not. Haiku rewrote the Bulgarian, Russian and Spanish versions in English, with the Cyrillic company name respelled in Latin letters. Sonnet always answered in the original language. A fixed English prompt run over content in several languages is exactly the setup where this happens, and the output still reads like a good rewrite, only of the wrong page. With the skill, both models kept the original language in all ten cases.

**The answer was more than the text.** Haiku opened all five answers by announcing "a more natural" or "more humanized" version and closed each with a bulleted note on its edits. Sonnet appended a note on its edits in Bulgarian and Russian, and in Bulgarian it also opened with a line of its own and offered to write a warmer variant. Anything that takes the answer as the new text publishes those notes with it. All four forbidden-phrase failures came from such notes, where the model listed its language's "in conclusion" among the phrases it had removed; the rewritten text no longer had it. Counting the notes, Haiku's answers were 1.5 to 2.2 times as long as the input, so a request meant to tidy the text made it longer.

**With the rules, smaller slips remain.** The skill asks for a rewrite at least nine tenths as long as the input, and on this padded text neither model held to that every time: Sonnet's answers ran from 55 to 100 percent of the input length, Haiku's from 45 to 80. A few closing lines of the models' own also survived. Sonnet's English rewrite ends by calling remote work simply the way the company runs, and Haiku's says the company "shows what remote work can accomplish". Both are milder than the testament they replaced, yet the input says neither. Haiku's Bulgarian rewrite used „обична", which means beloved, where the sentence needed „обичайна", usual. A native reader sees it at once; none of our checks did.

## Checks worth adding to your own pipeline

- **List the qualifiers before the rewrite and search for them after it.** Words such as usually, can, unless and only if, in every language you publish, are cheap to look for and are the first to disappear.
- **Name the output language in the request, or check it afterwards.** One English line in front of a Bulgarian text was enough to get English back.
- **Expect commentary and cut it off.** Ask for the rewrite alone, and reject any reply whose first line talks about itself.
- **Read for new reasons and softened numbers.** A claim about why the company does something, or an "about" in front of an exact figure, is the kind of edit no style check will flag.

## What we did not measure

- **One text of about a hundred words, written by us.** It was built to be full of machine phrasing. Real drafts are longer and only partly generated, and the models may handle them better.
- **A single run for each model, language and side**, with the two sides recorded on different days. A single case flipping between pass and fail tells you little.
- **A bare request on the side without the skill.** "Humanize this text." says nothing about keeping facts or language. A prompt that asks for both in a sentence or two might close much of the gap; we did not test one. The "without it" column therefore means "with a one-line request".
- **Exact-word checks.** A synonym for "usually" counts as a miss, which can make a zero look worse than it is. Reading the answers confirmed that the meaning of "usually" was lost as well, but the invented facts and motives above were found by reading, not counted by a script.
- **Whether anyone would now take the text for human writing.** We ran no detector and asked no readers; the checks only cover what must not change.
- **Two Claude models from one family.** Models from other vendors were not tested.
