Not on its own. Two Claude models, Sonnet and Haiku, each got six short texts and a bare proofreading request: "Proofread this text." Five of the texts carried errors we had planted, in Bulgarian, English, German, Russian and Spanish; the sixth was a correct English text that needed nothing. Without further instructions Sonnet passed four of the six cases and Haiku one. Once our rules were loaded as a skill, Sonnet passed all six and Haiku five. The interesting failures were not missed typos. Sonnet returned the correct text with two edits after calling it clean, and Haiku mostly answered with a numbered list of fixes and no corrected text at all, which leaves a program with nothing to pass on.
Method
We wrote six inputs with invented people and places, each one a short news-style paragraph:
- an English note about an office move, with six planted errors such as
thenforthan,recievedandkept is word; - a Bulgarian item about a council budget, a German one about a library opening, a Spanish one about a new park and a Russian one about an exhibition, each with five planted errors of the kind a native editor would mark: a wrong article ending,
seidforseit, a missing accent oncostó,потому-чтоwith a hyphen; - an English sketch of a market square that was already correct, written in British spelling, with a semicolon, a two-word exclamation standing alone and a time given as
between 8 and 2.
The request was identical on both sides. On one side our proofread skill was loaded as the model's instructions; on the other there was only the request. The models ran through Claude Code in print mode, with no access to tools and nothing loaded from our own configuration; every model, case and side got exactly one call, 24 answers in all. For the skill side we used the third of three runs from 30 September. The plain side was run on 3 October.
The comparison uses the same two checks on both sides. Each case has a list of strings that must appear in the answer exactly: every corrected phrase, plus every number and name from the input. For the clean text the list holds the names, the year and the phrases a careful editor would leave alone. The second check asks that the answer be written in the language the text was written in. The skill side has further checks of its own, such as length and the shape of the change list; they are not counted here.
Results
| Model | Cases passed with the skill | Cases passed without it |
|---|---|---|
| Sonnet | 6 of 6 | 4 of 6 |
| Haiku | 5 of 6 | 1 of 6 |
The correct text came back changed. Without instructions, Sonnet began its answer on the market sketch by saying the text had no outright errors. It then offered a version with two edits: Mr. lost its full stop to suit the British spelling, and the hours in between 8 and 2 gained a.m. and p.m., a reading the author never wrote. Haiku praised the same text and sent back no copy of it. With the skill, both models returned the sketch word for word and said that nothing had changed. A test set without a clean text would never have shown this, and in a pipeline that proofreads everything it sees, most texts are clean.
Right fixes, wrapped in markup. Without rules, Sonnet found all six English errors, yet it set each fix in bold inside the text and put the whole paragraph in a quote block. A program that takes that answer gets asterisks in the middle of words, and our check failed on three phrases for that reason alone. The same answer added an optional comma and a remark that the April date might already be in the past. With the skill the reply was the bare corrected paragraph followed by its list of changes.
Edits that are not corrections pass a string check. Two of Sonnet's passing answers without the skill changed more than the planted errors. In the Bulgarian text it dropped one of two uses of много, which turns roads in very bad condition into roads in bad condition; it also swapped a preposition and added remarks about the currency and about a proposal the text never describes. In the Spanish text it moved the sentence about the mayor's promise above the one about the organisers, gave the town hall a capital letter, added angle quotes and shortened the opening hours. Our check looks only for strings that must be present, so it scored both as passes. With the skill, Sonnet made the same five accent fixes in Spanish, kept the order of the sentences, and added a single flagged line saying it was unclear who made the promise. That extra line took the answer past the length limit of the skill's own stricter check, the one failure Sonnet had there.
A list instead of the text. Without rules, Haiku answered four of the six cases with a numbered list of fixes and returned no corrected text; for the Bulgarian, German and Spanish inputs the list itself was in English. Only for Russian did it send back the corrected paragraph. The lists were often right: the English one named all six planted errors. But a list is something a person reads, and an agent that has to forward the corrected text cannot use it. The Bulgarian list also contained a fix for a misspelled word that does not occur in the input, while the two real misspellings, училищята and засадание, were not on it.
Bulgarian stayed the weak spot for the small model. With the skill, Haiku returned the Bulgarian text in Bulgarian and fixed three of its five errors, but left the same two misspellings in place. Over three runs with the skill it missed the first one every time and the second one twice. Sonnet fixed all five Bulgarian errors with the skill and without it. The Russian text was the one case both models got right on both sides.
What this means for an agent pipeline
- Put a clean text in your tests. The case where the right answer is to change nothing is where both models went wrong without instructions, and it is the case your pipeline meets most often.
- Ask for the text first and a separate list of changes after it. Then compare the returned text with the input yourself. A difference the list does not mention is an edit nobody asked for, and code can reject it.
- Check that the reply comes back in the input's own language. Haiku without rules answered three non-English texts in English; a one-line language check catches that before it reaches a reader.
- Do not trust a check that only looks for the right fixes. Our own check passed answers that dropped a word and reordered sentences. Comparing the whole text with the input is what finds those.
- For Bulgarian, use the larger model or a human pass. The two misspellings Haiku kept are ordinary words, not rare ones.
What we did not measure
- Single answers, recorded days apart. Every model and case got one answer per side; the skill side dates from 30 September and the plain side from 3 October, with the same model versions. In our three runs with the skill, both models answered borderline cases differently from run to run, so a single case moving from pass to fail is within noise.
- Six short texts we wrote ourselves, with planted errors of a known kind. Real texts are longer, and their errors are fewer, subtler and mixed with deliberate choices that a proofreader has to recognise.
- Our check is strict in one direction and loose in the other. An answer that lists every fix correctly but does not return the text fails, which is why Haiku's column without the skill is so low; read it as "did not return usable text", not as "missed the errors". At the same time, extra edits such as a removed word pass, as long as the required strings are present.
- Only one plain request. A longer prompt that asks for the corrected text only, in the original language, might close part of the gap without any skill. We did not try that.
- Two Claude models from one family. No other vendors, and no dedicated grammar checker for comparison.
Read this post as Markdown: /blog/llm-proofreading-changes-correct-text.md · Atom feed.
