Test results
A2A 1.0 Agent Card Writer: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Wrote the right JSON for all 22 cases, read by hand: a card with supportedInterfaces and protocolBinding JSONRPC and no top-level url; SendMessage and SendStreamingMessage with ROLE_USER, a messageId and parts without a kind; the send configuration with returnImmediately and taskPushNotificationConfig; a flat push notification config with the task id; the follow-up with the task id inside the message; TASK_STATE_WORKING in the list call; structured data inside a data part of a wrapped task result; the error codes -32002, -32003, -32009, -32001 and -32602; the A2A-Version header with the value 1.0; and the well-known card path.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Also 22 of 22, with the same 1.0 names as Sonnet: PascalCase methods, ROLE_USER, parts with only text or data, the webhook details in the flat or the taskPushNotificationConfig form, and the exact error codes.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Correct A2A 1.0 JSON (22 cases) | 22/22 | 22/22 | 22/22 | 19/22 |
The same request on both sides, JSON-only asked for in both. For Sonnet the skill changed nothing on these 22 cases: it already wrote the 1.0 names without it, so it gains nothing from this skill here. The gain is for Haiku, which without the skill used the 0.3 habits on three cases: the method message/send with the role user and a kind on message and parts, and a webhook config wrapped as pushNotificationConfig or config instead of the 1.0 shape. Every name, code and shape in the skill was checked against the 1.0 specification and its protocol file (release 1.0.1), which makes it a checked reference rather than a fix for Sonnet.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-two cases written by us from the 1.0 specification and its protocol file: four agent cards and two security parts (interfaces, capabilities, bearer, OAuth client credentials, skills), nine requests (send, stream, send with a webhook, webhook create, follow-up, list, get), two replies (a task with a data part, a plain message), five error responses and two short questions (the version header and the card path). Each answer is scored by exact values and absent fields (kind, url, preferredTransport, blocking, the 0.3 wrappers) plus a JSON Schema (Ajv). A fence or text around the JSON fails. The checks accept every valid form; the 0.3 names and shapes are the strict traps. One run per model and case.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.