Test results

Social API Reality 2026: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Got 24 of 24 with the skill, read by hand. It answered the Hashnode read-only question as paid, sent Pinterest pins to the first board by alphabet in both cases (Garden, Decor), put the first pins under an hour, and said the 53 Facebook posts become visible after Live with no republish. X pricing, Mastodon 202, the one-hour Idempotency-Key, the scope and card answers, the Bluesky emoji limit, the card thumbnail, the daily write budget and both Bluesky byte offsets were right. In the very first run, with the earlier wording, it made two arithmetic slips on the offsets (an end one byte too far); the wording was fixed and three further runs were all right.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Got 23 of 24 with the skill: all the Hashnode, Pinterest timing, Meta and Bluesky answers were right, including both byte offsets, but on one Pinterest question it still chose Kitchen, the first board it was told of, instead of Decor, the first by alphabet.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Answers that match the checked fact (24 questions)24/2419/2423/2419/24

The same questions on both sides. Without the skill both models failed the same five (Sonnet) or four (Haiku) of them: they believed Hashnode read-only queries are still free, sent Pinterest pins to the newest or first-listed board instead of the first by alphabet, expected the first pins after about 24 hours instead of under an hour, and thought the Facebook posts made in Development mode stay hidden and must be republished after Live. They knew X pricing, Mastodon 202, the Idempotency-Key window and Bluesky units and offsets. With the skill, Sonnet passed all 24 across four runs; Haiku still chose the first-listed board on one Pinterest question.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-four short questions written by us on six services, each answered as JSON with exact fields: 16 traps on facts that changed in 2026 or that the documentation does not state (X link price, Hashnode plan, Pinterest board and timing and image tag, Mastodon token and card and retry key, Bluesky units and offsets and write budget, Meta Live) and 8 controls where an ordinary answer is right. Each field is compared with the one defensible value; a fence or extra field fails. The task text is the same on both sides. One run per model and question. The facts were read on the services' own pages on 8 October 2026; some come from our own posting tests and are marked as such in the skill. The skill was run four times with it (the Bluesky wording changed after the first) and once without; the with-skill numbers merge the runs.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Social API Reality 2026 · Card (JSON)