Not reliably, unless you tell it what to do at the places where the input leaves room for a guess. We gave two Claude models, Sonnet and Haiku, fifteen pieces of broken JSON with a plain request: fix this so a program can parse it, return only the JSON. Counting the content alone, with any Markdown code fence removed, Sonnet gave back exactly the original data in nine cases and Haiku in six. The misses were not syntax errors. They were answers that parse and say something the input never said: a value that was cut off kept as if it were whole, a number rewritten, and in one case a full object built out of an error message. With our written rules loaded as a skill, Sonnet got all fifteen right and Haiku twelve.

Method

We wrote fifteen small inputs with invented data, each built around one kind of damage that tools and models really produce:

  • a model's reply with a friendly sentence before it, a code fence and trailing commas;
  • a config with comments sitting next to URLs that contain a double slash;
  • a Python dictionary with True and None, plus a line pushing for the record to be approved;
  • three replies that break off partway: inside a text value, right after a digit, and halfway through a key name;
  • unquoted keys with missing commas, and a string holding a raw line break and unescaped quotes;
  • numbers written 01234, 007, 0x1F, .5 and +3, and a price written 1,250;
  • NaN, Infinity and undefined;
  • keys and strings wrapped in curly quotes instead of straight ones, with more curly quotes inside one value;
  • a Bulgarian Python dictionary whose address contains „Витоша“ in Bulgarian quotes;
  • one input that was already valid and had to come back unchanged;
  • a rate-limit notice, HTTP 429, that contains no JSON whatsoever.

Both sides received the same request. One side also had our JSON repair skill as its instructions; the other had nothing extra. We called the models from the Claude Code command line with every tool disabled and without our own settings, once per model, case and side, which makes 60 runs. A script parsed each answer and compared the entire value with the expected one: every key, every value, its type and its order. One changed character in one string fails the case. For the rate-limit message, the side without the skill only had to avoid inventing JSON; it could not know the fixed refusal line the skill asks for.

We scored each answer twice. The first score looks at the values only, after stripping a code fence. The second is what a program sees when it hands the raw answer to a JSON parser.

Results

Model and score With the skill Without it
Sonnet, values right 15 of 15 9 of 15
Sonnet, right and parses as returned 15 of 15 5 of 15
Haiku, values right 12 of 15 6 of 15
Haiku, right and parses as returned 7 of 15 0 of 15

Truncated JSON is where the data goes wrong. A model that runs out of tokens leaves its JSON hanging, and the obvious next step is to ask a model to close it. Without rules, both kept unfinished fragments as if they were real data. Sonnet did drop a half-written key on its own, but it kept a summary that ended mid-sentence, "Revenue grew in the north and the", as a finished string, and kept a quantity that stopped at 1 as a plain 1, although the writer may have been typing 12. Haiku went further. It finished the sentence on its own with "and the south", gave the cut-off item a price of 0 that appears nowhere in the input, and turned a key that broke off at "renew into "renew": true. Every one of those answers parses. None of them matches the data. With the skill, the rule is simple: anything left unfinished turns into null, and an unfinished key name goes, together with its value. Sonnet followed it in all three cases; Haiku in two, and once kept the quantity of one.

Numbers get tidied up. Both models turned the price 1,250 into 1250, though in much of Europe that comma is a decimal point and the price may be one and a quarter. Both turned the hexadecimal 0x1F into 31. Haiku also dropped the leading zeros from the zip code 01234 and the id 007; Sonnet kept those two as strings by itself. The opposite slip happened as well: given input that was already valid, Haiku put a long numeric id into quotes and so changed its type from number to string. With the skill, both models kept 01234, 007, 0x1F and 1,250 as strings exactly as written, and left valid numbers alone.

An error page became data. The last input was a plain rate-limit message. Without rules, both models produced a tidy object from it, with an error code of 429 and a retry time of 30 seconds, Haiku even nesting it under an error key. A caller that trusts the repair step would now treat a refused request as a real response. With the skill, both replied with one line that begins CANNOT REPAIR:, a fixed prefix a caller can look for before parsing.

Quotes inside a value are part of the value. Both models, without rules, replaced the curly single quotes in "She said ‘yes’ twice" with straight ones. That looks harmless, but the string has changed, and any comparison or checksum built on it will no longer match. In the Bulgarian address, Haiku swapped the closing mark of „Витоша“ for a straight double quote, which ends a JSON string early, so the whole answer stopped parsing. Haiku kept making both mistakes with the skill loaded; Sonnet with the skill kept every quote as it was. We tried adding a worked example for exactly this to the skill and ran the three affected cases again. Haiku still straightened the quotes, so the example came out.

The planted instruction fooled neither model. The Python dictionary carried a note addressed to the model, asking it to flip approved before answering. On both sides, both models left it false and kept the comment as an ordinary string. Without the skill, Sonnet then added a sentence after the JSON to explain that it had done so, which is courteous and enough to make the reply fail a parser.

The code fence, as usual. Haiku put a Markdown code fence around every one of its fifteen answers without the skill, and around seven with it. Sonnet used a fence in five answers without the skill and in none with it. With a small model behind your agent, strip one fence from the start and end of the reply before parsing; it costs a line of code and touches no data.

What to take into your own pipeline

  • Write down the rule for cut-off values before you ask for a repair. Without one, a model reaches for the most plausible completion, and plausible is the hardest kind of wrong to spot later.
  • Check the HTTP status before anything reaches a repair step. A body that came with an error status is an error, and no repair should turn it into a record.
  • Keep any number whose meaning depends on how it is written as a string. Zip codes, ids with leading zeros and prices with a comma are the usual suspects.
  • If your damage is only ever syntax, a code library may be the better tool. A parser-based fixer such as jsonrepair gives the same answer every time and costs no model call. A model earns its place when the input needs judgement, such as text around the JSON or a decision about what to refuse.

What we did not measure

  • A single run for each model, case and side. Haiku in particular can answer differently on a second run, so a single case moving between pass and fail is within noise.
  • Fifteen short inputs we wrote ourselves, each built around one kind of damage. Real broken output is often longer and damaged in many places at once.
  • Our expected answer is the only right one. Where two repairs are defensible, for example keeping a cut-off key as null instead of dropping it, the check accepts only ours. That favours the side that knew our rules, so read the "without it" column as "without these rules", not as "wrong by every standard". The error message, the rewritten numbers and the invented values are wrong by any standard.
  • No input with two separate JSON values, or with two possible readings. The skill is written to refuse those, but this test has no case that checks it.
  • No comparison with a code library or with models from other vendors. Two Claude models from one family were the only ones tested.

Read this post as Markdown: /blog/llm-json-repair-changes-data.md · Atom feed.