Test results

Social Dispatch Code Review: Bluesky, Mastodon, DEV: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on all 24 snippets in substance, read by hand: it asked for a media upload instead of hoping for a preview card, for the page title instead of the first characters of an HTML error page, for an honest User-Agent toward the DEV API, for counting items rather than strings in a Mastodon feed, for a quiet switch-off when a token is missing, and caught the scope change that replaces the token, the replay key window, the timeouts and the sequential loop. On two snippets the form slipped: once it put No findings. and the verdict on one line, once it listed a code with the words "does not apply".
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on all 24 snippets in substance, read by hand, with the same findings as Sonnet, and No findings. on the eight sound snippets. On two snippets it also reported true problems our key had not listed: a post made of title and address only, and a loop that one thrown request stops.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Snippets reviewed right (24 snippets)24/2421/2424/2419/24

Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already caught the timeouts, the sequential loop, the token that is not checked, the replay window and the scope change. It missed three: it trusted Mastodon to build a preview card from the article page, it kept the first 400 characters of an HTML error page as the log line, and it did not know that the DEV API refuses a Worker request that carries no User-Agent. Haiku without the skill missed the same three, plus the Mastodon feed that counts one post twice and the missing token that should switch a channel off. Without the skill both models also raised alarms on most sound snippets; that is not counted.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-four snippets of code that announces articles on Bluesky, Mastodon, DEV and Tumblr, written by us: 16 with a planted defect and 8 sound. With the skill each answer is scored by code on the finding codes and the verdict line; without it the same request is scored on the concept in any words. On the sound snippets the bare side has no check, so a false alarm there is not counted on either side; the counts below rest on the 16 faulty snippets. Owner-measured from incidents in September and October 2026 and NOT re-checked: the body cut that breaks the Bluesky login, Mastodon not fetching a preview card, the scope change that replaces the token, the HTML error page, the DEV API refusing a request with no User-Agent. Checks widened after the run: the missing-token concept accepts a reply that says to return a quiet skipped result (both sides); two snippets accept extra true findings that our key had missed (title-only post, a loop without try). One run per model and snippet.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Social Dispatch Code Review: Bluesky, Mastodon, DEV · Card (JSON)