Skip to content
DnsLister Forum

Where domain hunters compare notes

ChatGPT returns “Cache miss” and Claude says “ROBOTS_DISALLOWED” — but both are allowed, site returns 200, and no request reaches the server. Anyone else seeing this?

I’ve been troubleshooting a strange AI web-retrieval issue on a WordPress/WooCommerce site I own (`yourseoebook.com`) and I’m looking for other site owners who may be seeing the same thing.

I’m **not claiming there is an OpenAI or Anthropic outage**. I’m trying to determine whether there is a broader issue involving AI retrieval caches, stale robots decisions, crawler infrastructure, or hosting/network layers.

Here’s what I’ve tested so far.

The affected URLs include:

[`https://yourseoebook.com/actionseo-pricing/\`\](https://yourseoebook.com/actionseo-pricing/)

[`https://yourseoebook.com/product/actionseo-pro-trial-30-days/\`\](https://yourseoebook.com/product/actionseo-pro-trial-30-days/)

# What works

The pages:

* return **HTTP 200**

* work over both IPv4 and IPv6

* have correct self-referencing canonicals

* do **not** contain `noindex` or `nofollow`

* are successfully crawled by Googlebot

* are indexed in Google

* pass Google Search Console live inspection

* are included in working XML sitemaps

* are accessible through normal browser/curl requests

* now return normal LiteSpeed public cache responses (`MISS → HIT`)

* can currently be read by Gemini

I also disabled the Hostinger CDN while troubleshooting.

# robots.txt

The current robots.txt explicitly allows the relevant bots:

User-agent: OAI-SearchBot

Allow: /

User-agent: ChatGPT-User

Allow: /

User-agent: GPTBot

Allow: /

User-agent: Claude-SearchBot

Allow: /

User-agent: Claude-User

Allow: /

User-agent: ClaudeBot

Allow: /

I tested the file directly using different User-Agents, and they all receive the current robots.txt correctly.

# ChatGPT behavior

When I ask ChatGPT to open the exact product URL, I get:

Failed to fetch / Cache miss

Search can sometimes discover the page, but direct retrieval fails.

OpenAI support explained that an exact-URL prompt does not necessarily guarantee a live origin fetch and that Search may use indexed/cached retrieval.

So we performed a controlled test.

**23 Sep 2026 – 18:10 Europe/Berlin / 16:10 UTC**

URL:

[`https://yourseoebook.com/product/actionseo-pro-trial-30-days/\`\](https://yourseoebook.com/product/actionseo-pro-trial-30-days/)

ChatGPT result:

Failed to fetch / Cache miss

Hostinger then checked the available bot metrics for that exact time window.

They found:

ChatGPT-User: 0 allowed / 0 blocked

OAI-SearchBot: 0 allowed / 0 blocked

GPTBot: 0 allowed / 0 blocked

There was also **no matching request in the origin access log**.

Earlier browser/curl requests to the exact same URL appeared normally and returned HTTP 200.

Interestingly, Hostinger’s aggregated bot metrics from the previous 7 days show that OpenAI bots *can* reach the site:

GPTBot: 11 allowed

OAI-SearchBot: 8 allowed

ChatGPT-User: 7 allowed

Blocked: 0

So this doesn’t look like a general OpenAI block at the hosting level.

OpenAI support has now escalated the case to a specialist.

# Claude behavior

Claude is different.

When asked to open the same domain, Claude returns:

ROBOTS_DISALLOWED

Sometimes even when asked to retrieve the robots.txt itself.

But Hostinger’s bot metrics for the last 7 days show:

ClaudeBot: 215 allowed

Blocked: 0

So Anthropic crawling in general clearly reaches the domain.

However, Hostinger does not expose `Claude-User` and `Claude-SearchBot` as separate categories, and during interactive Claude tests I haven’t seen a corresponding request reach the origin.

I’ve sent the full diagnostics to Anthropic support as well.

# Other things checked

I also checked:

* DNS

* IPv4/IPv6

* LiteSpeed exclusions

* page-cache behavior

* Hostinger IP blocking

* crawler blocking settings

* WordPress REST API

* sitemap headers

* meta robots

* canonicals

* redirects

* external scripts

* JSON-LD

* suspicious `eval()`, `atob()`, Base64-style code, etc.

Nothing so far explains why these AI retrieval layers behave differently.

There was one unrelated structured-data issue in the site footer, but it cannot explain Claude refusing even `robots.txt`, and it does not explain ChatGPT producing a cache miss before a request reaches the origin.

# What I’m trying to understand

The pattern currently looks roughly like this:

Browser ✅

curl ✅ HTTP 200

Googlebot ✅

Google Search Console ✅

Google Search ✅

Gemini ✅

ChatGPT exact URL ❌ Cache miss

Claude exact URL ❌ ROBOTS_DISALLOWED

What I find particularly interesting is that during some failed AI retrieval attempts **there is no corresponding request at the origin at all**.

That raises the possibility of another layer between:

User prompt

↓

AI retrieval/index/cache/policy layer

↓

crawler/user-fetch infrastructure

↓

hosting/network edge

↓

origin server

A site can apparently be perfectly crawlable at the origin while an AI assistant still decides that it cannot retrieve it.

# Could other site owners test this?

If you manage a website, I’d be very interested if you could try this:

  1. Pick a recently updated public URL.

  2. Confirm it returns HTTP 200.

  3. Ask ChatGPT to **open that exact URL and report something currently on the page**.

  4. Do the same with Claude.

  5. Immediately check your server/CDN logs.

I’m especially interested in cases where:

curl = 200

Googlebot = works

ChatGPT/Claude = fails

AND

no corresponding request reaches the origin

If you see this, please mention:

* hosting/CDN provider

* CMS

* ChatGPT result

* Claude result

* whether the request appeared in your logs

* whether the AI returned stale content instead

I’d like to collect enough examples to determine whether this is just a handful of unrelated configurations or a broader AI web-retrieval behavior.

Source: r/ChatGPTcomplaints · by /u/franco_yamakawa

Leave a Reply

Your email address will not be published. Required fields are marked *