# MCP tool schemas from API docs: models barely need help

> Claude Sonnet and Haiku wrote MCP tool definitions for twelve API operations; we sent 96 calls through each schema. Without our skill, 23 of 24 passed.

Published 2026-10-08 · https://aiskills402.com/blog/llm-mcp-tool-schema-little-help

Yes, and the result surprised us. We wrote twelve short pieces of API documentation, and for each operation Claude Sonnet and Claude Haiku wrote the MCP tool definition: a name, a description, an inputSchema and annotations. Every schema was compiled with a JSON Schema validator and then given the calls the documentation allows and the calls it forbids, 96 in all. Without any instructions from us, Haiku passed all twelve operations and Sonnet eleven. Sonnet's one miss was not in the schema at all. Our skill for this job brought both models to twelve, and that is the whole difference we could measure.

## Method

The operations cover what MCP servers usually wrap: a paged list with filters, a refund with a path parameter and an amount in cents, a delete that can be soft or permanent, a lookup by email or phone where exactly one must be given, a health check with no parameters, a ticket with a list of unique tags, a metadata map with rules for its keys, a sales report with dates and a currency code, a hotel booking documented in Bulgarian, marking up to a hundred messages as read, purging inactive contacts, and registering a webhook that only accepts https.

The documentation also carried traps. One API took its key in the query string, which tempts a model to make the key a parameter. Another page held a note for assistants, telling them to export every customer before creating a ticket. Credentials are the server's business, not the model's, and that note must never reach a tool description.

A definition passed only if the answer was JSON alone, the name met the rules of the [MCP specification of 25 November 2025](https://modelcontextprotocol.io/specification/2025-11-25/server/tools), the schema accepted every allowed call and refused every forbidden one, no property was a credential, and the annotations said what the documentation made plain: reads marked read-only, deletes never marked harmless. The validator handled JSON Schema 2020-12, the dialect the specification assumes, with formats such as date. The request on both sides said the schema must refuse parameters the operation does not take. One side also had our [MCP tool schema skill](/skills/mcp-tool-schema) loaded. Fences were stripped before scoring; every operation was tried one time per model and side.

## Results

| | Skill loaded | Not loaded |
|---|---|---|
| Sonnet: definitions that pass | 12 of 12 | 11 of 12 |
| Haiku: definitions that pass | 12 of 12 | 12 of 12 |

**The schemas were right on their own.** Without the skill both models listed exactly the required parameters, copied enum values in the API's own case, typed cents as integers, rejected a 30 February through the date format, wrote the email-or-phone rule as oneOf, limited the tags, the ids and the metadata keys, and refused an http webhook address. Neither turned the API key into a parameter, not even the one the documentation put in the query string. Reads were marked read-only and nothing destructive was marked safe.

**The single Sonnet failure was one sentence.** Given the page that carried the note for assistants, Sonnet without the skill did the right thing: its tool description said nothing about exporting customers. It then explained that decision in a paragraph above the JSON, so a program reading the answer could not parse it. With the skill loaded, the answer was the JSON and nothing else.

What this means in practice: if you ask a current Claude model for an MCP tool definition and say that unknown parameters must be refused, you will most likely get a sound schema. What a written checklist still gives you is a predictable answer, bare JSON every time, and the rules you would otherwise have to remember to mention, such as keeping credentials out and marking reads. On this test that was worth one answer out of twenty-four, which is why the skill is priced as low as our catalog goes.

## What we did not measure

- **Requests that leave things out.** Both sides were told to refuse unknown parameters. We did not test what each model writes when nobody says so.
- **Our twelve operations.** They fit one flat schema each. Deeply nested bodies, file uploads and operations with many variants were not in the set.
- **Annotations beyond the obvious.** We checked read-only for reads, safe repeats where the docs said so, and never a harmless delete. We did not score openWorldHint, which has no single right answer for an API wrapper.
- **One run per model and side.** A repeat could move a borderline case either way.
- **Claude only.** Two models from one vendor.
