Reads the results of probing a website with AI-crawler user agents (status codes per bot, a control request, sometimes robots.txt or the site's settings) and says what they really show as JSON - whether training bots, search bots and user-triggered agents are open, blocked, mixed or unknown, what is behind a refusal, and whether the robots.txt that answered is the site's own file. Knows the traps - a bare bot name proves nothing and lies differently for each bot, a blocked training bot is a choice while a blocked search bot is a defect, Python-urllib and libwww-perl are refused by Browser Integrity Check and not by the site's policy, since 15 September 2026 Cloudflare refuses training bots on zones nobody touched, and a robots.txt of only comments is a stand-in from the edge. Use when asked whether a site blocks ChatGPT, Claude, Perplexity or Google AI, how to read curl results with bot user agents, or why an AI crawler gets a 403.
AI Crawler Access Reader is a tested SKILL.md that reads the results of probing a website with AI-crawler user agents (status codes per bot, a control request, sometimes robots.txt or the site's settings) and says what they really show as JSON - whether training bots, search bots and user-triggered agents are open, blocked, mixed or unknown, what is behind a refusal, and whether the robots.txt that answered is the site's own file; an agent buys it once for $0.05 over x402.
Not for
Running the probe or fetching any site: it reads only the results you paste. Not for writing a robots.txt file, which is another skill, and not for judging content quality, rankings or whether a bot ought to be allowed. It does not read your Cloudflare settings, because an API token cannot; it judges by how the site actually answered.
Tested, honestly
Tested 2026-10-08 with a strong and a weak model.
With and without the skill
Results with and without the skill, for Sonnet and Haiku |
| with | without | with | without |
|---|
| Right reading of the crawler probe results (23 situations) |
| Right reading of the crawler probe results (23 situations) | 23/23 | 19/23 | 21/23 | 17/23 |
|---|
The same request on both sides: the same fields with the same neutral one-line meanings. Read by hand, Sonnet's four misses without the skill are real. Two read bare bot names as evidence: it called training "mixed" because bare ClaudeBot got 403 and bare GPTBot got 200, and called training "open" because three bare names got 200. For a zone nobody had touched since August it named an owner rule, against the facts it was given. For a Google-Extended probe it wrote "open" while its own reason said the probe proves nothing; that one is a mislabelled field, not a gap in knowledge. Haiku missed the same four plus two more.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
- SonnetStrong model, claude-sonnet-5-5
- Read all 23 probe situations correctly: bare bot names were set aside as no evidence (training unknown, cause probe error) and a control request that was refused too made every result unknown; the full GPTBot refusal beat the bare 200; the refusal of a training bot on an untouched zone was put down to the Cloudflare default, and an agent refused only on pages with ads was called mixed; a refused search bot was named as a defect; Python-urllib and the Perl and Java clients were put down to Browser Integrity Check, and a probe of Google-Extended was left unknown; a comments-only robots.txt with the 404's headers was called a stand-in, and rules in the live file that the repository lacks were traced to a Cloudflare setting.
- HaikuWeak model, claude-haiku-5-5
- 21 of 23 with the skill, up from 17, and the same readings as Sonnet on bare names, the failed control, the ads pages and Google-Extended. But one answer was cut off and is not valid JSON (the mixed training case), and for the rules injected above a robots.txt it said the cause was probably a Cloudflare setting in its text yet wrote "unknown" in the field.
Full test summary
Example
Our own test text, before and after the skill ran. Excerpts only.
English · claude-sonnet-5-5
Before
We tested https://shop.example/ with curl, one request per user agent, all within one minute. -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/141.0.0.0 Safari/537.36" gave 200. -A "GPTBot" gave 200. -A "ClaudeBot" gave 403. Does the site treat OpenAI's and Anthropic's training crawlers differently?
After
{"training": "unknown", "search": "unknown", "agent": "unknown", "cause": "probe_error", "robots_txt": "unknown", "why": "The control returned 200, but GPTBot and ClaudeBot were sent as bare names, and Cloudflare matches each bot by its full user agent string, so a 200 for one and a 403 for the other shows nothing about how the site treats OpenAI versus Anthropic.…
What is in the file
- The answer
- The three kinds of bot
- What decides the reading
- Three readings, worked through
- Where the obvious reading is right
- Work in this order
Languages
Any language. Tried in: English.
License
Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Full terms.
Versions
Current version 1.0.0, updated 2026-10-08. Whoever bought an earlier version gets new ones free through the same re-download token.
v1.0.0 · 2026-10-08
First release: reads the results of probing a site with AI-crawler user agents and returns JSON with the state of training, search and user-triggered bots (open, blocked, mixed or unknown), the cause of a refusal (a Cloudflare AI setting, Browser Integrity Check, a rule the owner wrote, or an error in the probe itself) and whether the robots.txt that answered is the site's own or a stand-in from the platform. Discards results made with bare bot names, requires a working control request, pairs each company's training and search bot, and knows that Google-Extended cannot be probed. Price: 0.05 USD.
FAQ
What comes back for a set of probe results?
A single JSON object with fixed fields that code can switch on: training, search and agent, each open, blocked, mixed or unknown; a cause such as cloudflare_ai_setting or browser_integrity_check; whether the robots.txt that answered is the site's own file or a stand-in; and a short reason in plain words. Mixed results, an unprobed bot, or a failed control are reported as mixed or unknown instead of guessed, which is what keeps an audit from publishing a claim the numbers never supported.
Why not simply read the status codes myself?
Because a probe with a bare bot name lies, and lies differently for each bot. On one site bare ClaudeBot got 403 and bare GPTBot got 200, which looked like a policy against Anthropic; with full user agents both were refused alike. The skill discards such results, insists on a control request, and pairs each company's training bot with its search bot, since a refused search bot means lost visibility while a refused training bot is usually a choice. Statuses alone cannot show this.
Does it know about the Cloudflare change of September 2026?
Yes. On 15 September 2026 Cloudflare moved zones to separate Search, Training and Agent settings, so training bots receive 403 on sites nobody touched. The skill tells that default apart from a rule the owner wrote, and from Browser Integrity Check, which turns away Python-urllib and libwww-perl with an HTML page and leaves no line in the Worker log. It also knows Google-Extended cannot be probed at all. Whoever audits many sites for visibility to assistants gets the same reading every time.
Does it help Claude Sonnet?
Somewhat, on specific points. Over 23 probe situations Sonnet got 19 right alone and all 23 with this file; Haiku rose from 17 to 21. The real gains: Sonnet treated bare bot names as proof, blamed an owner rule on an untouched site, and marked a Google-Extended probe open. Pairs, Browser Integrity Check and robots.txt stand-ins it already handled. The facts were read on 8 October 2026; Anthropic publishes only tokens, so copy Claude strings from your logs.