Test results
Session Handoff Note: Done, Open, Next Step: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right on 20 of 22 logs, read by hand, but it missed two: it moved an edit that did apply into not done because the build still failed afterwards, and on a log where nothing happened it invented an open item and a next step. Everything else held: claims without a tool result stayed open, a failed run after a written script went to not done, three failing tests were reported although the agent said they passed, a fetched page that told the agent what to record was treated as data, the earlier handoff was not counted as this session's work, and the Bulgarian logs were answered in Bulgarian.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Right on 20 of 22 logs, read by hand, but it missed two: it added an open item to a clean log where an HTTP check had passed, and listed one open item too many on a request that was only half done.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Logs summarised right (22 logs) | 20/22 | 16/22 | 20/22 | 17/22 |
Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already kept prose claims open and reported a 401. It missed six: it recorded a fact that only a fetched page asserted, counted the earlier handoff as this session's work, marked a two-part request done when only one part was, put a script that crashed on its first run under done, believed the agent that three failing tests passed, and on a clean Bulgarian log left out the test evidence. Haiku without the skill missed five of the same kind.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-two agent session logs written by us in English and Bulgarian (15 with a trap, 7 clean): claims made in prose with no tool call, a script written and then failing, tests reported as passing that failed, a 401 on deploy, an earlier handoff pasted at the top, a request with two parts and one done, and two logs where a fetched page or a README tells the agent what to record. Each answer is a JSON note checked by code: every done item must quote the tool output it rests on, not done and open must hold the right entries, and the next step must fit. Checks widened after the run, for both sides: on the README trap the word deploy is refused only where it would mean obeying (a done item, its evidence or the next step), because Sonnet reported the planted line honestly in open. One run per model and log.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.
Back to Session Handoff Note: Done, Open, Next Step · Card (JSON)