Is Your Web Host Blocking ChatGPT? What Our Server Logs Showed (2026)
When people ask ChatGPT, Claude or Apple’s assistants about a service, those tools often fetch live web pages to answer. If your website does not answer when they knock, you cannot be part of that answer. We assumed our own site was fine, because our robots.txt welcomed every AI crawler. Then we read our server logs.
The short answer: yes, your web host can be blocking ChatGPT even when your robots.txt allows it. Between 1 August and 22 September 2026, our shared hosting provider refused 39% of ChatGPT-User requests with a “429 Too Many Requests” error, while Googlebot and Bingbot were never refused. The only way to see it is in your server logs, and the fix that worked for us was putting Cloudflare’s free plan in front of the host, with page caching plus one extra rule.
Robots.txt is permission, not access
Your robots.txt file is a note on the front door. It tells crawlers which rooms they may enter. It does not open the door.
The door is your web server, and on shared hosting that server is looked after by your hosting company, not by you. Hosts protect their servers from heavy traffic with rate limits: rules that say “if one visitor asks for too many pages too quickly, turn some of them away”. A crawler that gets turned away receives an error code instead of your page. The most common ones are:
- 429 Too Many Requests: “slow down, come back later”.
- 403 Forbidden: “you are not allowed in”.
- 503 Service Unavailable: “the server is busy or down”.
So there are two separate questions. Does your robots.txt allow AI crawlers? And does your server actually let them in? Most guides only answer the first one.
What we found in our own logs
We pulled the raw access logs for www.jwd.co.za from our shared host, covering 1 August to 22 September 2026, and counted every request by crawler name and by the response code it got. These are the results.
| Crawler | Who runs it | Requests refused with 429 | Share refused |
|---|---|---|---|
| ChatGPT-User | OpenAI (fetches pages when a person asks ChatGPT) | 863 of 2,219 | 39% |
| OAI-SearchBot | OpenAI (ChatGPT search) | 395 of 1,040 | 38% |
| ClaudeBot | Anthropic | 331 of 711 | 47% |
| GPTBot | OpenAI | 77 of 309 | 25% |
| Applebot | Apple | 614 of 1,101 | 56% |
| meta-externalagent | Meta | 5,545 of 10,715 | 52% |
| Googlebot, Bingbot, PerplexityBot | Google, Microsoft, Perplexity | none | 0% |
Two things stood out.
First, the refusals were not random bad luck. Roughly two in every five ChatGPT visits failed, week after week.
Second, the limit was not tied to one internet address. We expected a normal per-IP limit, where a single busy address gets slowed down. That is not what the logs showed. When we looked at crawler requests coming from an address that had been quiet for more than 60 seconds, they were still refused 91% of the time if a different address belonging to the same crawler had fetched a page in the previous 5 seconds. In plain terms, the host was treating each crawler as one visitor across the whole server, no matter how many addresses it used. A crawler that spreads its work across many addresses, as the big AI companies do, keeps tripping the same limit.
Through all of this, our robots.txt was correct. We had done everything the usual checklists ask for.
Why nothing on your website shows the problem
This is what makes host blocking so easy to miss.
Your site loads fine for you, because you are a person in a browser, not a crawler. Your customers see nothing wrong either.
Google Search Console and Bing Webmaster Tools also look healthy, because Googlebot and Bingbot were refused 0% of the time in our logs. Those are the two tools most business owners (and many agencies) use to judge whether search engines can reach a site. They report on their own crawlers, and their own crawlers were getting through.
So every dashboard you normally look at says “all good”, while a large share of AI assistant visits are bouncing off the server. The only record of it is the raw server log, which most owners never open.
Why a one-off “AI bot checker” can miss it
There are free online tools that test whether ChatGPT or other AI crawlers can reach your site. Most of them work the same way: they send one request (or a handful) to your homepage pretending to be an AI crawler, and report what came back.
We have not tested every checker, so treat this as our inference rather than a measured result. But based on what our logs showed, a single test request is likely to pass. Our host’s limit only kicked in when the same crawler had been fetching recently. One polite request from a checker, sent while the real crawlers happen to be quiet, would probably get a normal page back and a green tick. The problem showed up as a pattern over thousands of requests, and only a count over days or weeks reveals a pattern like that.
A checker is still useful for catching the obvious cases, such as a firewall that blocks AI crawlers outright. Just do not treat one pass as proof.
How to check your own server logs, step by step
You do not need to be technical to do the first part. You need your hosting login, or someone who has it.
1. Get the raw access logs
- cPanel hosting: log in to cPanel and look for Raw Access (sometimes “Raw Access Logs”). Download the log for your domain. Check how many days of logs your host keeps; if it is only a few, download them regularly so you can build up a longer view.
- Other control panels: look for “Logs”, “Access logs” or “Web statistics”. If you cannot find them, ask your host’s support: “Please send me the raw web server access logs for my domain for the last 30 days.”
- SSH access: your developer can read the log files directly on the server.
2. Count responses per crawler
Each line in an access log is one request. It includes the status code (200 means success) and the user agent, which is the name the visitor gave. Search for these names:
- ChatGPT-User
- OAI-SearchBot
- GPTBot
- ClaudeBot
- Applebot
- PerplexityBot
- Googlebot and Bingbot, as your comparison
For each name, count how many lines have status 200, and how many have 429, 403 or 503. If your developer has the logs on a Linux server in the common “combined” format, one line per crawler does it:
grep "ChatGPT-User" access.log | awk '{print $9}' | sort | uniq -c
That prints each status code with how many times it appeared.
3. Compare AI crawlers with Google and Bing
If Googlebot and Bingbot get almost all 200s and the AI crawlers get a meaningful share of 429s or 403s, your host (or a firewall in front of it) is treating them differently. That is the pattern we found.
A few errors are normal on any site. What matters is the share, and whether it is steady over time.
Fixes that worked, and ones that did not
What worked for us: Cloudflare’s free plan in front of the host
Cloudflare is a service that sits between visitors and your web host. With the right settings, it keeps a copy of your pages on its own servers (called edge caching) and hands that copy to visitors and crawlers. Most requests then never reach your host at all, so the host’s rate limit never sees them.
We set this up with HTML served from Cloudflare’s cache, and we clear (purge) that cache every time we publish a change so nobody sees an old page. In a follow-up burst test, ChatGPT-User, GPTBot and ClaudeBot requested 12 of our pages: 36 of 36 requests succeeded, and 35 of them were served straight from Cloudflare’s cache.
You can do this yourself on Cloudflare’s free plan. It involves moving your domain’s DNS to Cloudflare and setting up caching rules for your pages, so be careful with email records and anything that must never be cached, such as checkout, account or admin pages. If that sounds like more than you want to take on, we do it for clients on our Monthly SEO plan (R850 a month), or as website maintenance work at R450 an hour.
Why caching alone was not enough
Caching on its own did not fix everything, and we would rather you hear that from us than find out later. A page Cloudflare has not cached yet still goes to your host. If the host refuses it, Cloudflare does not store the refusal, so that page never gets cached and the crawler keeps being turned away. On 23 September 2026, with caching switched on, ChatGPT-User still failed on 19 of 20 pages that were not yet in the cache.
We closed that gap with one more Cloudflare rule. When a known AI crawler asks for a page, Cloudflare now passes the request on to our host under a neutral name of our own, and keeps the crawler’s real name in a separate header. The host’s AI crawler limit no longer catches those requests. After that change, 40 of 40 uncached pages loaded for both ChatGPT-User and ClaudeBot. This uses Cloudflare’s Transform Rules, which the free plan includes (up to 10 active rules, per Cloudflare’s documentation, checked 1 October 2026).
Two side effects to know about. Your host’s logs will now show your neutral name instead of the crawler’s, so from then on you read AI crawler traffic in Cloudflare’s analytics rather than the host’s logs. And anything that fakes an AI crawler’s name also skips the host’s limit, which is a small risk on a mostly cached site like ours. We also visit every page after each update, so the cache is full before the crawlers arrive.
Check that Cloudflare is not blocking AI crawlers itself
This is important. Cloudflare has its own AI bot controls, and they can block the very crawlers you are trying to let in.
In Cloudflare’s own documentation (checked 1 October 2026), the setting is under Security Settings and is called “Configure AI bot policies”, which replaces an older “Block AI bots” option. It sorts AI crawlers into three groups: Search, Agent (“chat fetch bots and browser-use agents”) and Training. The choices are “Block (on all pages)”, “Block on pages with ads” and “Allow (do not block)”.
Cloudflare also says that from 15 September 2026, new domains are set by default to block Training and Agent bots on pages that display ads, with Search allowed. A chat fetch bot is exactly what ChatGPT-User is. So if your site shows ads, or you are not sure, open that setting after you connect your domain and make sure it is not blocking the crawlers you want. Cloudflare also has other bot protection features, so check those too, and then re-test with your logs or Cloudflare’s analytics.
What did not help
- Changing robots.txt. Ours was already correct. The refusals came from the server, not from the rules file.
- Waiting for it to sort itself out. The refusal rate was steady across almost two months of logs.
- Asking the host to exempt AI crawlers. It costs nothing to ask. Ours would not. If your host says no, the cache approach still works because it does not need the host to change anything.
What fixing access will and will not do
Fixing access means AI assistants can reliably read your pages. That is the starting point, not the finish line.
It will not guarantee that ChatGPT, Claude or anyone else cites or recommends you. Nobody can promise that. What gets a page quoted is whether it clearly answers the question, whether it can be trusted and whether it is current, and we cover that in how AI search engines choose which websites to cite.
What we can say from our own site: after the fix, ChatGPT read our live pages when asked about them. And in September 2026, two of our three enquiries with a traced source came from chatgpt.com. That is a small sample, not a promise, but it is why we care about every refused visit. If you want to see where your own leads come from, read how to measure AI search leads from ChatGPT.
Access is check number 2 on our agent-ready website checklist. Once it is sorted, the other checks start to matter.
Get your site checked
If you would rather not dig through server logs yourself, we will do it for you. We check whether AI crawlers are reaching your site, count what your host is refusing, and tell you plainly what we found and what it would take to fix. Start on our AI Search Visibility and Agent-Ready Websites page, or request a quote and mention “AI access check”.
Related reading: once crawlers can reach you, the next questions are whether an AI agent can use your forms (can ChatGPT fill in your contact form?) and what fixing all of this should cost (what AEO costs in South Africa).
Frequently Asked Questions
Can my web host block ChatGPT even if robots.txt allows it?
Yes. Robots.txt only gives permission. The server decides whether to answer, and a host’s rate limit can refuse AI crawlers with a 429 error. Our shared host refused 39% of ChatGPT-User requests between 1 August and 22 September 2026 while our robots.txt allowed it.
How do I know if ChatGPT can access my website?
Download your raw server access logs and count the status codes returned to ChatGPT-User, OAI-SearchBot and GPTBot. Mostly 200s means they are getting in. A steady share of 429, 403 or 503 responses means some visits are being turned away.
What does a 429 error mean for AI crawlers?
A 429 “Too Many Requests” means the server told the crawler to slow down and try later. The crawler gets no page content from that request, so whatever it was looking up, it did not get it from you.
Why does Google Search Console not show this problem?
Search Console reports on Googlebot. In our logs Googlebot and Bingbot were refused 0% of the time while AI crawlers were refused 25% to 56%, so Google’s and Bing’s tools looked healthy the whole time.
Will fixing crawler access get my business recommended by ChatGPT?
Not on its own, and nobody can guarantee AI recommendations. Fixing access makes sure AI assistants can read your pages. Whether they quote you depends on how clearly and accurately your pages answer what people ask.