Test results

Traffic Reality Check: test results

Tested 2026-10-08, skill version 1.0.1 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right in all 21 situations, read by hand: headless checks and one-view link previews at publishing hours were removed (70 readers of 203), real Bulgarian readers from the social network were kept (282 and 54), the cache host answering 504 was not counted as errors (4 real ones; 18 where they were real), the site's own Worker and its Cache API rows were taken out of the outside requests (1995), staging was left out (211), the EU exclusion setting gave zero readers and an incomplete Bulgarian dashboard, both policy cases named the missing directive, data older than 7 days was called sampled (one beacon behind a figure of 10) and newer data exact, six days needed six queries, and no separate token was asked for.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
20 of 21, with the same figures as Sonnet in every other situation, but it missed one: the small control: for 12 clean page views it wrote that the figure is not exact because the beacon misses visitors who block scripts, and so answered exact false where the sampling rule makes last-week data exact.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Right figure or setting reading (21 situations)21/2114/2120/2111/21

The same request on both sides, JSON-only asked for in both. Read by hand, Sonnet without the skill made seven real errors: it counted the social-network link previews as readers (139 instead of 70), gave no sampling answer for data older than 7 days (null), counted 84 readers under the EU exclusion setting instead of 0, thought a separate API token is needed, threw away 33 real Bulgarian readers who came from the social network (249 instead of 282), and twice called small fresh counts not exact. The other 14 situations it solved, including the cache host, the Worker rows and both policy cases.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-one situations written by us the way an owner pastes a dashboard: 14 where the obvious reading is wrong (previews counted as readers, cache and Worker rows counted as visitors or errors, a zero that hides readers) and 7 controls where real readers or real errors must stay in, so any harm from the skill would show. The answer is JSON with a computed number or a true/false, scored on the values only; the reason is not scored. The facts come from our own measurements on live sites between 16 August and 28 September 2026 and from Cloudflare's documentation, named in the skill's notes. One run per model and situation.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Traffic Reality Check · Card (JSON)