The standard advice now is "block the training bots, allow the search bots" – disallow GPTBot, ClaudeBot, CCBot, Google-Extended, allow OAI-SearchBot, Claude-SearchBot, PerplexityBot. Keep your work out of training, stay citable. Good advice. I went looking at how often it is actually implemented that way, and the answer surprised me.
Method: 624 sites checked two ways, robots.txt parsed to RFC 9309 plus a paired browser/crawler probe. Then I took 14 that refuse an AI crawler at the origin and probed each with all eight named crawlers plus three controls: an ordinary browser, plain curl, and an invented crawler user-agent that is on no blocklist anywhere.
That third control is the one that decides whether a finding is real. If the made-up crawler is refused too, it is general bot protection and there is nothing to tell anyone. On all 14 the invented crawler was served normally, so a rule names these crawlers specifically.
What the 14 look like:
– All 14 return 403 to GPTBot and/or ClaudeBot while serving a browser 200.
– Not one names a user-agent introduced in 2025 or later. Claude-SearchBot, Claude-User and Perplexity-User are served normally on every single one.
– 12 of 14 say nothing about AI crawlers in robots.txt at all. The refusal is at the server, below the level robots.txt controls – usually a security plugin or a host default.
– 2 of 14 explicitly ALLOW, in a hand-written robots.txt section, the exact crawler their own server returns 403 to.
The blocked-token sets cluster oddly. Five of the 14, in five countries, on four different server stacks, refuse ClaudeBot and nothing else. Five independent teams do not separately decide to refuse Anthropic while allowing OpenAI and Perplexity, so I think one copied rule is circulating. I have not traced where it starts and would genuinely like to know.
The upshot is that these sites are simultaneously stricter and weaker than their owners believe. They refuse crawlers nobody meant to refuse, and serve the newer answer-time ones they think they blocked. Checking robots.txt will not show you any of it, because that is not where it is happening.
Two method caveats, because both cost me time:
– A 429 is not a block. A rate limit your own probing triggered looks identical from outside. I had one site that looked like it refused everything and was actually just rate-limiting me.
– Anything behind a bot-verifying CDN is unverifiable, not blocked. A user-agent string proves nothing when the CDN checks IP ranges and reverse DNS.
Not naming the sites – all 14 owners were written to privately first with the evidence and a way to check it themselves. Happy to answer method questions.
Source: r/GEO_optimization · by /u/Its_SeenSure