Test results
JSON Schema From Examples: No Over-Constraining: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Wrote a schema for all 23 sets of examples that a draft 2020-12 validator accepted and that took every document it should and refused every document it should: a key missing from one example stayed optional, two observed values did not become a closed list, whole-number examples did not become an integer type, limits came from the rules and not from the smallest and largest example, a date or the text TBD was a choice, null was a type in a list, a pair of coordinates used positional items, and a pattern for two letters and four digits was anchored. The sentence inside an example telling the reader to require everything stayed a plain text value.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Also 23 of 23, with schemas of the same kinds as Sonnet: optional keys left optional, sets closed only where a rule closed them, anchored patterns, and nothing taken from a planted instruction.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Schema accepts and refuses the right documents (23 sets) | 23/23 | 22/23 | 23/23 | 23/23 |
The same request and checks on both sides; a fence around the answer is removed first. Without the skill both models wrote correct schemas in nearly every case, including the traps for optional keys, open sets and number types. Sonnet's one miss was a format miss: for the code pattern it wrote a schema, then 'Wait', a second schema, so the answer was not one JSON object; the final schema was right. On this test the skill adds almost nothing measurable for either model. What it gives is a written, checked rule set and an answer that is one object every time.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-three sets of example documents written by us, one with its rules in Bulgarian, each with a few plain-word rules. Every answer's schema was compiled with Ajv for JSON Schema 2020-12 with format checking and run against 162 documents in all: documents it had to accept (a new status, a decimal, a missing optional key, an extra field) and documents it had to refuse (a wrong type, a missing required key, a code with extra characters, a third coordinate). The judge is what the schema accepts and refuses, so any correct way of writing it passes. The request on both sides asked for the bare schema and for acceptance of everything the examples and rules give no reason to reject. One run per model and case.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.
Back to JSON Schema From Examples: No Over-Constraining · Card (JSON)