Test results
Search Console Next Action: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Chose the right action code in all 22 situations, with a reason that names the cause: Validate fix instead of Request indexing for URLs with a redirect error, the large-image robots setting in the root layout for Discover, the full sitemap address for a Domain property, a live test for the fetch label, no action for a robots.txt row that answers 404, internal links for URLs without a referring page, allowing Googlebot on www, noindex instead of Disallow, a prefix removal over a 410 folder, deleting the Host line, listing the whole URL family, purging the cache behind a 308 without Location, and the plain cases (request indexing for a new page, repair a 500, wait while validation runs).
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Also 22 of 22, with the same codes as Sonnet and reasons that name the same causes.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Right action for the Search Console situation (22 situations) | 22/22 | 17/22 | 22/22 | 16/22 |
Same task on both sides: the same 14 action codes with neutral one-line meanings. Read by hand, Sonnet's five wrong answers without the skill are two real gaps, one half right and two defensible. Gaps: for a redirect error created by the Request indexing button it chose a live test, and for a redirect over a cached 404 (a 308 with no Location) it chose a site repair, not a cache purge. Half right: for redirecting URLs it chose no action, knowing that requesting indexing would not help. Defensible: removing the Disallow line, the first of two needed steps, and no action for the informational Host warning. The honest gain for Sonnet is two to three of 22.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-two Search Console situations written by us the way a site owner describes them: 15 where the obvious action is wrong and 7 where it is right, so any harm from the skill would show. The answer is JSON with an action code from a list of 14, a reason and steps; it is scored on the code, and on a word in the reason for most cases. The facts come from our own measurements on live sites (19 and 23 September 2026) and from Google's documentation, named in the skill's notes. One run per model and situation.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.