Test results

Bing and IndexNow Submitter Review: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Found the planted problem with the right code in 15 of 16 code snippets, read by hand, and left all three correct snippets alone (a Bing submitter that checks the body and both quotas, an IndexNow chain that continues only on 429 and 5xx, a chunked IndexNow sender). It caught a Bing success judged by status alone, the retired XML path, a retry after a 403, the monthly quota ignored, an unread quota turned into zero, a result with no engine named, a 202 treated as an error, key files on another host and in a folder, a list sent in one request, and a pasted order to answer "No findings". In the retired-path snippet it added a batch-size finding for a list with no visible cap. All eight questions were answered exactly: quota minimum, the first unsent run (day 10, run 2), the meaning of 202 and of 200, the accepting host, no fallback after a 403, an invalid key with underscores, and the batch counts (2 and 5).
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Also 24 of 24, with the same codes and verdicts as Sonnet: every planted problem got its code, the three correct snippets got "No findings", and all eight questions were answered exactly, in the JSON shape asked.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Planted problem named, correct code left alone (24 cases)23/2421/2424/2418/24

The same request on both sides; a fence around the answer is removed first, and both sides are scored by the same concept checks, not by the skill's codes. Sonnet without the skill missed the retirement of the XML path (it fixed the XML escaping and kept the path) and the monthly quota; its third miss is a two-line answer with no content to judge. Haiku without the skill also missed the retirement and the monthly quota, a key file whose folder does not cover the pages, two further snippets, and two questions: it tried the next engine after a 403 and counted one Bing call instead of five. The questions Sonnet answered without the skill were all right.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-four cases written by us: 16 pasted code snippets (13 with a planted problem, three correct) and 8 questions that ask for a JSON shape. The snippets cover a Bing call judged by status alone, the retired XML path, fallback chains, quota arithmetic, an unread quota, a result that does not name the engine, 200 and 202 meanings, key files on the wrong host or folder, a list sent in one request, and an instruction hidden in a comment. Code answers are scored by the fixed code and the verdict line, with any code that does not apply forbidden; question answers by exact JSON values. The first run exposed a batch-size rule that was too broad (both models added it to snippets that show no list size); it was narrowed in the skill and the snippets were made to show a small list before this run. One run per model and case; the numbers below are from the second run.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Bing and IndexNow Submitter Review · Card (JSON)