I’ve been troubleshooting a strange AI web-retrieval issue on a WordPress/WooCommerce site I own (`yourseoebook.com`) and I’m looking for other site owners who may be seeing the same thing.
I’m **not claiming there is an OpenAI or Anthropic outage**. I’m trying to determine whether there is a broader issue involving AI retrieval caches, stale robots decisions, crawler infrastructure, or hosting/network layers.
Here’s what I’ve tested so far.
The affected URLs include:
[`https://yourseoebook.com/actionseo-pricing/\`\](https://yourseoebook.com/actionseo-pricing/)
# What works
The pages:
* return **HTTP 200**
* work over both IPv4 and IPv6
* have correct self-referencing canonicals
* do **not** contain `noindex` or `nofollow`
* are successfully crawled by Googlebot
* are indexed in Google
* pass Google Search Console live inspection
* are included in working XML sitemaps
* are accessible through normal browser/curl requests
* now return normal LiteSpeed public cache responses (`MISS → HIT`)
* can currently be read by Gemini
I also disabled the Hostinger CDN while troubleshooting.
# robots.txt
The current robots.txt explicitly allows the relevant bots:
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: GPTBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: ClaudeBot
Allow: /
I tested the file directly using different User-Agents, and they all receive the current robots.txt correctly.
# ChatGPT behavior
When I ask ChatGPT to open the exact product URL, I get:
Failed to fetch / Cache miss
Search can sometimes discover the page, but direct retrieval fails.
OpenAI support explained that an exact-URL prompt does not necessarily guarantee a live origin fetch and that Search may use indexed/cached retrieval.
So we performed a controlled test.
**23 Sep 2026 – 18:10 Europe/Berlin / 16:10 UTC**
URL:
ChatGPT result:
Failed to fetch / Cache miss
Hostinger then checked the available bot metrics for that exact time window.
They found:
ChatGPT-User: 0 allowed / 0 blocked
OAI-SearchBot: 0 allowed / 0 blocked
GPTBot: 0 allowed / 0 blocked
There was also **no matching request in the origin access log**.
Earlier browser/curl requests to the exact same URL appeared normally and returned HTTP 200.
Interestingly, Hostinger’s aggregated bot metrics from the previous 7 days show that OpenAI bots *can* reach the site:
GPTBot: 11 allowed
OAI-SearchBot: 8 allowed
ChatGPT-User: 7 allowed
Blocked: 0
So this doesn’t look like a general OpenAI block at the hosting level.
OpenAI support has now escalated the case to a specialist.
# Claude behavior
Claude is different.
When asked to open the same domain, Claude returns:
ROBOTS_DISALLOWED
Sometimes even when asked to retrieve the robots.txt itself.
But Hostinger’s bot metrics for the last 7 days show:
ClaudeBot: 215 allowed
Blocked: 0
So Anthropic crawling in general clearly reaches the domain.
However, Hostinger does not expose `Claude-User` and `Claude-SearchBot` as separate categories, and during interactive Claude tests I haven’t seen a corresponding request reach the origin.
I’ve sent the full diagnostics to Anthropic support as well.
# Other things checked
I also checked:
* DNS
* IPv4/IPv6
* LiteSpeed exclusions
* page-cache behavior
* Hostinger IP blocking
* crawler blocking settings
* WordPress REST API
* sitemap headers
* meta robots
* canonicals
* redirects
* external scripts
* JSON-LD
* suspicious `eval()`, `atob()`, Base64-style code, etc.
Nothing so far explains why these AI retrieval layers behave differently.
There was one unrelated structured-data issue in the site footer, but it cannot explain Claude refusing even `robots.txt`, and it does not explain ChatGPT producing a cache miss before a request reaches the origin.
# What I’m trying to understand
The pattern currently looks roughly like this:
Browser ✅
curl ✅ HTTP 200
Googlebot ✅
Google Search Console ✅
Google Search ✅
Gemini ✅
ChatGPT exact URL ❌ Cache miss
Claude exact URL ❌ ROBOTS_DISALLOWED
What I find particularly interesting is that during some failed AI retrieval attempts **there is no corresponding request at the origin at all**.
That raises the possibility of another layer between:
User prompt
↓
AI retrieval/index/cache/policy layer
↓
crawler/user-fetch infrastructure
↓
hosting/network edge
↓
origin server
A site can apparently be perfectly crawlable at the origin while an AI assistant still decides that it cannot retrieve it.
# Could other site owners test this?
If you manage a website, I’d be very interested if you could try this:
-
Pick a recently updated public URL.
-
Confirm it returns HTTP 200.
-
Ask ChatGPT to **open that exact URL and report something currently on the page**.
-
Do the same with Claude.
-
Immediately check your server/CDN logs.
I’m especially interested in cases where:
curl = 200
Googlebot = works
ChatGPT/Claude = fails
AND
no corresponding request reaches the origin
If you see this, please mention:
* hosting/CDN provider
* CMS
* ChatGPT result
* Claude result
* whether the request appeared in your logs
* whether the AI returned stale content instead
I’d like to collect enough examples to determine whether this is just a handful of unrelated configurations or a broader AI web-retrieval behavior.
Source: r/ChatGPTcomplaints · by /u/franco_yamakawa