# AI Crawler Access Reader

AI Crawler Access Reader is a tested SKILL.md that reads the results of probing a website with AI-crawler user agents (status codes per bot, a control request, sometimes robots.txt or the site's settings) and says what they really show as JSON - whether training bots, search bots and user-triggered agents are open, blocked, mixed or unknown, what is behind a refusal, and whether the robots.txt that answered is the site's own file; an agent buys it once for $0.05 over x402.

- Page: https://aiskills402.com/skills/crawler-access-probe
- Category: SEO & Content (https://aiskills402.com/categories/seo)
- Price: $0.05 once, USD-priced, paid in USDC on Base over x402. Price as loaded on this page. The 402 response your agent receives is authoritative.
- Version: 1.0.0
- Card (JSON): https://api.aiskills402.com/v1/skills/crawler-access-probe

## Use it when

Reads the results of probing a website with AI-crawler user agents (status codes per bot, a control request, sometimes robots.txt or the site's settings) and says what they really show as JSON - whether training bots, search bots and user-triggered agents are open, blocked, mixed or unknown, what is behind a refusal, and whether the robots.txt that answered is the site's own file. Knows the traps - a bare bot name proves nothing and lies differently for each bot, a blocked training bot is a choice while a blocked search bot is a defect, Python-urllib and libwww-perl are refused by Browser Integrity Check and not by the site's policy, since 15 September 2026 Cloudflare refuses training bots on zones nobody touched, and a robots.txt of only comments is a stand-in from the edge. Use when asked whether a site blocks ChatGPT, Claude, Perplexity or Google AI, how to read curl results with bot user agents, or why an AI crawler gets a 403.

## Not for

Running the probe or fetching any site: it reads only the results you paste. Not for writing a robots.txt file, which is another skill, and not for judging content quality, rankings or whether a bot ought to be allowed. It does not read your Cloudflare settings, because an API token cannot; it judges by how the site actually answered.

## Tested, honestly

Tested 2026-10-08.

- Strong model (claude-sonnet-5-5 (Claude Code alias "sonnet")): Read all 23 probe situations correctly: bare bot names were set aside as no evidence (training unknown, cause probe error) and a control request that was refused too made every result unknown; the full GPTBot refusal beat the bare 200; the refusal of a training bot on an untouched zone was put down to the Cloudflare default, and an agent refused only on pages with ads was called mixed; a refused search bot was named as a defect; Python-urllib and the Perl and Java clients were put down to Browser Integrity Check, and a probe of Google-Extended was left unknown; a comments-only robots.txt with the 404's headers was called a stand-in, and rules in the live file that the repository lacks were traced to a Cloudflare setting.
- Weak model (claude-haiku-5-5 (Claude Code alias "haiku")): 21 of 23 with the skill, up from 17, and the same readings as Sonnet on bare names, the failed control, the ads pages and Google-Extended. But one answer was cut off and is not valid JSON (the mixed training case), and for the rules injected above a robots.txt it said the cause was probably a Cloudflare setting in its text yet wrote "unknown" in the field.

Note: Twenty-three probe situations written by us the way an auditor pastes curl results: 15 where the obvious reading is wrong and 8 where it is right, so any harm from the skill would show. The answer is JSON with three bot kinds, a cause and a robots.txt state, scored on the fields that have one defensible answer. The facts come from the vendors' crawler pages read on 8 October 2026 and from our own measurements on live sites in September 2026, named in the skill's notes. One run per model and situation.

### With and without the skill

Tested 2026-10-08.

- Right reading of the crawler probe results (23 situations): Sonnet 23/23 with, 19/23 without; Haiku 21/23 with, 17/23 without.

The same request on both sides: the same fields with the same neutral one-line meanings. Read by hand, Sonnet's four misses without the skill are real. Two read bare bot names as evidence: it called training "mixed" because bare ClaudeBot got 403 and bare GPTBot got 200, and called training "open" because three bare names got 200. For a zone nobody had touched since August it named an owner rule, against the facts it was given. For a Google-Extended probe it wrote "open" while its own reason said the probe proves nothing; that one is a mislabelled field, not a gap in knowledge. Haiku missed the same four plus two more.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Full summary: https://aiskills402.com/skills/crawler-access-probe/tests

## Example

### English (claude-sonnet-5-5)

Before:

> We tested https://shop.example/ with curl, one request per user agent, all within one minute. -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/141.0.0.0 Safari/537.36" gave 200. -A "GPTBot" gave 200. -A "ClaudeBot" gave 403. Does the site treat OpenAI's and Anthropic's training crawlers differently?

After:

> {"training": "unknown", "search": "unknown", "agent": "unknown", "cause": "probe_error", "robots_txt": "unknown", "why": "The control returned 200, but GPTBot and ClaudeBot were sent as bare names, and Cloudflare matches each bot by its full user agent string, so a 200 for one and a 403 for the other shows nothing about how the site treats OpenAI versus Anthropic.…

## What is in the file

- The answer
- The three kinds of bot
- What decides the reading
- Three readings, worked through
- Where the obvious reading is right
- Work in this order

## Languages

Any language

## How to buy

Agent (HTTP):

1. GET https://api.aiskills402.com/v1/skills/crawler-access-probe/file without a payment header. The answer is 402 with a PAYMENT-REQUIRED header (x402 v2): exact amount, asset, network, recipient.
2. Sign `accepts[0]` with an x402 client (for example @x402/core + @x402/evm).
3. Repeat the GET with the signature in the PAYMENT-SIGNATURE header. The answer is 200 with the file, its sha256 and a re-download token.

Agent (MCP): https://mcp.aiskills402.com/mcp — free tools search_skills, get_skill, redownload_skill. Buying itself is over HTTP.

Full flow: https://aiskills402.com/docs

## The file

- Version: 1.0.0
- Size: 9.5 KB (9721 bytes)
- SHA-256: 5dd2aaa4e1022d05705d3f59f2f775e0284bdb99b41198d070e6b2c1ac63250a
- Updated: 2026-10-08
- New versions are free through your re-download token.

## Versions

### 1.0.0 (2026-10-08)

First release: reads the results of probing a site with AI-crawler user agents and returns JSON with the state of training, search and user-triggered bots (open, blocked, mixed or unknown), the cause of a refusal (a Cloudflare AI setting, Browser Integrity Check, a rule the owner wrote, or an error in the probe itself) and whether the robots.txt that answered is the site's own or a stand-in from the platform. Discards results made with bare bot names, requires a working control request, pairs each company's training and search bot, and knows that Google-Extended cannot be probed. Price: 0.05 USD.

## License

Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Terms: https://aiskills402.com/docs#license

## FAQ

### What comes back for a set of probe results?

A single JSON object with fixed fields that code can switch on: training, search and agent, each open, blocked, mixed or unknown; a cause such as cloudflare_ai_setting or browser_integrity_check; whether the robots.txt that answered is the site's own file or a stand-in; and a short reason in plain words. Mixed results, an unprobed bot, or a failed control are reported as mixed or unknown instead of guessed, which is what keeps an audit from publishing a claim the numbers never supported.

### Why not simply read the status codes myself?

Because a probe with a bare bot name lies, and lies differently for each bot. On one site bare ClaudeBot got 403 and bare GPTBot got 200, which looked like a policy against Anthropic; with full user agents both were refused alike. The skill discards such results, insists on a control request, and pairs each company's training bot with its search bot, since a refused search bot means lost visibility while a refused training bot is usually a choice. Statuses alone cannot show this.

### Does it know about the Cloudflare change of September 2026?

Yes. On 15 September 2026 Cloudflare moved zones to separate Search, Training and Agent settings, so training bots receive 403 on sites nobody touched. The skill tells that default apart from a rule the owner wrote, and from Browser Integrity Check, which turns away Python-urllib and libwww-perl with an HTML page and leaves no line in the Worker log. It also knows Google-Extended cannot be probed at all. Whoever audits many sites for visibility to assistants gets the same reading every time.

### Does it help Claude Sonnet?

Somewhat, on specific points. Over 23 probe situations Sonnet got 19 right alone and all 23 with this file; Haiku rose from 17 to 21. The real gains: Sonnet treated bare bot names as proof, blamed an owner rule on an untouched site, and marked a Google-Extended probe open. Pairs, Browser Integrity Check and robots.txt stand-ins it already handled. The facts were read on 8 October 2026; Anthropic publishes only tokens, so copy Claude strings from your logs.

## Related skills

- [Robots.txt Policy Writer](https://aiskills402.com/skills/robots-txt-policy.md): $0.03 once
- [Search Console Next Action](https://aiskills402.com/skills/search-console-actions.md): $0.05 once
- [SEO Meta Writer](https://aiskills402.com/skills/seo-meta.md): $0.03 once

## Measurement limits

- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.

Offer note: Paid in USDC (USD-pegged) over x402 by an AI agent; one-time.
