Skip to content
DnsLister Forum

Where domain hunters compare notes

A site’s homepage and its inner pages can give AI crawlers different answers, so most “is this site blocking AI” checks are really just homepage checks

I've been running fetch checks across a few hundred agency and small-business sites to see which AI crawlers can actually reach them. Something turned up this week that I hadn't accounted for, and it makes me distrust a chunk of what I'd already collected.

I was testing one URL per domain – the homepage – and calling that result "the site". That's wrong often enough to matter.

The block usually lives at the edge, and the edge doesn't guard every path equally.

When it's a managed WAF or bot-protection rule rather than robots.txt, it tends to get switched on for the front door first. So the homepage challenges an unrecognised datacentre IP while /blog/ and /services/ go straight through. I've also seen the reverse, where someone put a rule on a form-heavy section and left the homepage wide open. Same domain, two different answers, and which one you report depends entirely on which URL you happened to pick.

A challenge page is served as HTTP 200.

This is the one that actually cost me. When an edge decides a request looks automated it returns an interstitial – "Checking your browser", "One moment, please…" – with status 200 and a few hundred lines of HTML. Nothing in the status code says blocked. A checker looking at the status records a success. And a crawler that hits it receives that interstitial as the content of the page.

So it's silent in both directions. Uptime monitor green, access log full of 200s, and the model that fetched you has your bot-challenge text stored as your homepage.

I tested one of these directly, with the owner's permission, because I couldn't tell whether it was an AI-crawler rule or something broader. Five identities: two real crawler user-agents, a plain desktop browser UA, and one I made up on the spot that has never existed. All five got the same 200, the same byte count, the same "One moment, please…". It wasn't a rule about AI at all.

What I changed:

  • Sample paths, not domains. Homepage plus two or three inner pages of different types.
  • Compare the body, not just the status. A challenge and a real page both return 200; only one of them is 12KB of JavaScript.
  • Separate "refuses AI crawlers" from "refuses anything automated". Identical from a single probe, completely different in meaning – one is a policy someone chose, the other is collateral the owner almost certainly doesn't know about.
  • When the result is ambiguous, report it as ambiguous. An unclear result filed as a pass is worse than no result.

Rough number from what I've run so far: around 15% of the agency sites I sampled sit behind an edge that challenges an automated fetch, and in most of those cases the rule wasn't aimed at AI crawlers at all.

Does anyone here sample multiple paths as standard? And has anyone seen the reverse case – homepage open, inner pages walled – often enough to have a theory about it?

 

EDIT, half an hour later. The owner of the site I mentioned saw this, sent me the exact URLs, and asked me to publish the comparison. I re-ran the whole grid and the result corrects me on one point, so it belongs up here rather than in a comment.

It wasn't the path. It was the hostname.

grizeljboats.com/ 200 12,155 "One moment, please..." grizeljboats.com/en/ 301 795 -> www.grizeljboats.com/en/ grizeljboats.com/en/our-products/throne-15 301 795 -> www... grizeljboats.com/robots.txt 301 795 -> www... www.grizeljboats.com/hr/ 200 158,476 real page www.grizeljboats.com/en/ 200 158,872 real page www.grizeljboats.com/de/ 200 153,413 real page www.grizeljboats.com/en/our-products/throne-15 200 292,822 real page 

Every www URL is clean. Every apex path redirects to www correctly. The one exception is the apex root, which returns 200 and the challenge instead of the 301 that all of its sibling paths get. Same bytes for ClaudeBot, GPTBot, PerplexityBot, a plain desktop Chrome UA, and a bot name I invented.

So the axis I should have led with isn't path, it's hostname – and specifically the apex, which is the URL people type, print and link to. A crawler that starts there gets 12KB of challenge HTML as the homepage and no redirect out, because that one URL doesn't issue one. Every other apex path would have carried it across.

Add it to the checklist: test the apex and the www hostname separately, and don't assume one redirects to the other just because most of its paths do.

Source: r/AI_SearchOptimization · by /u/Its_SeenSure

Leave a Reply

Your email address will not be published. Required fields are marked *