Test results
JSON Lines Repair: One Record Per Line: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right on all 22 inputs, read by hand: it re-joined pretty-printed records, split glued ones, unwrapped an array, fixed single quotes and Python literals, and copied 1.10, 1e2, -0 and a twenty-digit integer as written. It refused with a CANNOT REPAIR line for the four cut-off last records, the cut inside a glued line, the repeated key, the leading-zero 007 whose neighbours are numbers and the input with no records, and it kept the planted orders inside strings and between records as plain data.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Right on all 22 inputs, read by hand, with the same refusals as Sonnet: the cut-off records, the repeated key, the ambiguous 007 and the error message with no records. It copied the number tokens, left the planted instructions alone and returned bare records with no fence.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Inputs repaired right (22 inputs) | 22/22 | 15/22 | 22/22 | 15/22 |
Same request on both sides. Read by hand, Sonnet without the skill already did every plain repair, copied number tokens, quoted a leading-zero ZIP code whose neighbours were text, and closed the record cut off only before its final brace without adding a field. It missed the cases where the data decides: it dropped a cut-off last record without saying so (twice), invented the end of two others (a closed array, a null), kept the last value of a repeated key, turned 007 into 7, and turned 'Error 503' into a record. Haiku missed the same cases and two more: it finished the cut name Ol as a record and re-wrote 1e2 as 100 and -0 as 0. It quoted 007 as text, one of two readings, picked silently.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-two inputs written by us, all data invented (14 traps, 8 controls): pretty-printed, glued and array inputs, fences and chat text, single quotes, comments, Python literals, big integers, planted instructions, five cut-off last records, a repeated key, a leading-zero number and an input with no records. Each answer is scored by code: record count, every leaf value, number tokens that must survive, and the exact CANNOT REPAIR line for the refusals. Rules re-checked on 2026-10-08 in the JSON Lines convention, the NDJSON specification and RFC 8259 (no leading zeros, repeated names unpredictable). The no-skill side was scored with the code fence removed first. One check was widened after the run for BOTH sides: for the record cut off only before its last closing brace, either the refusal or the exact closed record with no added field is accepted, because closing that one brace invents nothing. One run per model and input.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.
Back to JSON Lines Repair: One Record Per Line · Card (JSON)