ICODA crawled 4,689 casino, affiliate and crypto websites on 6 and 7 October 2026 to see which ones turn AI crawlers away. We analysed 4,322 of them and could measure access on 3,590. Of those, 20.9% block at least one of GPTBot, ClaudeBot, PerplexityBot or Google-Extended. Affiliates block most. Most blocks come from servers and CDNs, and few are written into robots.txt files.
Key Findings
- One site in five blocks a major AI crawler. 20.9% of 3,590 measurable sites (751 sites, 95% CI 19.6-22.3%) refuse at least one of the four core crawlers.
- Affiliates block 1.5x more often than operators. 27.4% of affiliates against 18.4% of casino and betting operators, with intervals that do not overlap.
- Servers do most of the blocking. 19.2% of sites show an AI-specific server block, while 3.5% have an AI-specific robots.txt block.
- Cloudflare is the strongest single factor. 25.9% of sites behind Cloudflare block, against 10.9% of sites elsewhere.
- Most server blocks hit training bots and miss ChatGPT’s search bot. 58.1% of server blocks let OpenAI’s search and user agents through.
- No clear link between blocking and AI citations. On 935 sites, the blocking coefficient is indistinguishable from zero once traffic and Domain Rating are controlled. This is correlation only.
Do Casino Affiliates Block AI Crawlers More Than the Casinos They Promote?
27.4% of affiliate sites (394 of 1,438) block at least one core AI crawler, against 18.4% of operators (210 of 1,142). The 95% intervals, 25.2-29.8% and 16.2-20.7%, do not overlap, and the ratio is 1.49x.
We expected the opposite. Affiliates depend on search traffic and publish structured content, so we guessed they would be the more AI-ready segment. Half of that held up: affiliates publish an llms.txt file four times as often as operators (12.3% against 3.0%), and 59.4% carry Organization schema on the homepage against 20.4% of operators. The other half did not, because they also block more often.
Traffic widens the gap. Affiliate sites that block at least one core crawler hold 44.5% of all affiliate organic traffic, though they are only 27.4% of affiliate sites. Operators run the other way: 9.9% of operator traffic sits on blocking sites, which are 18.4% of operator sites. The blockers are the large affiliates and the small operators. Traffic shares use Ahrefs organic traffic estimates.
How Many Casino and Crypto Sites Block AI Crawlers?
Only affiliates sit above the 20.9% average across segments. Crypto exchanges come second at 19.1%.
| Segment | Measurable sites | Blocks any core crawler (95% CI) | Via robots.txt | Via server | Valid llms.txt |
|---|---|---|---|---|---|
| Casino affiliates | 1,438 | 27.4% (25.2-29.8) | 4.8% | 24.7% | 12.3% |
| Crypto exchanges | 209 | 19.1% (14.4-25.0) | 1.4% | 18.8% | 21.0% |
| Casino and betting operators | 1,142 | 18.4% (16.2-20.7) | 4.8% | 16.9% | 3.0% |
| iGaming B2B vendors | 182 | 13.7% (9.5-19.5) | 2.8% | 12.8% | 18.5% |
| Crypto projects (CoinGecko top 1,000) | 619 | 13.2% (10.8-16.1) | 1.7% | 12.6% | 19.4% |
| All segments | 3,590 | 20.9% (19.6-22.3) | 3.9% | 19.2% | 11.1% |
Share of measurable sites. The robots.txt column counts any core crawler blocked, including files that also shut out Googlebot. The server column counts AI-specific blocks only. The llms.txt column uses all analysed sites in each segment.

By crawler, GPTBot is effectively blocked on 17.9% of measurable sites, ClaudeBot on 18.7%, PerplexityBot on 8.8% and OAI-SearchBot on 8.7%. Google-Extended is blocked in robots.txt on 2.6%. Sites that block hold 37.5% of exchange organic traffic but only 0.9% of crypto-project traffic, so blocking among crypto projects looks like a small-site habit.
Operators were the hardest segment to measure. Of the 1,513 we analysed, 516 (34.1%) put a wall in front of our crawler that stopped both a browser and a neutral bot. We read their robots.txt files but could not test the server. Casino sites often geo-block, and our crawler sat in France.
Is robots.txt Doing the Blocking?
An AI-specific robots.txt block, one that disallows AI crawlers and leaves Googlebot open, appears on 3.5% of measurable sites (120 of 3,469). An AI-specific server block appears on 19.2% (640 of 3,330).
A server block means the site served a normal page to a browser and to our neutral bot, then returned a 403 or a challenge page to an AI user agent. Depending on segment, these server blocks are 4 to 13 times more common than AI-specific robots.txt blocks. Of the server blocks we measured twice, 628 repeated 30 hours later and 16 vanished, so the pattern holds up.
An owner who checks only robots.txt can see a clean file and still be blocked. The file, parsed to RFC 9309 rules (Allow and Disallow groups, longest match wins), states the policy. The WAF enforces or overrides it. Few sites use the newer opt-out signals: 2.6% (113 of 4,322) publish Content-Signal lines, the ai-train, search and ai-input vocabulary Cloudflare promotes, and 3 carry a noai meta tag.
How Much Does Cloudflare Matter?
Sites behind Cloudflare block at least one core crawler 25.9% of the time (621 of 2,397), against 10.9% of sites elsewhere (130 of 1,193), a ratio of 2.4x. Cloudflare sits in front of 66.8% of the measurable sites.

The gap appears in each segment. Exchanges block 26.2% behind Cloudflare and 4.4% elsewhere, crypto projects 18.7% and 5.7%, operators 22.5% and 8.4%, and affiliates 31.9% and 16.8%. B2B vendors show the same direction at a smaller size (16.2% against 11.8%), and their intervals overlap. Of the 751 sites that block, 621 (83%) sit behind Cloudflare.
On 1 July 2025 Cloudflare changed its default to block AI crawlers unless they pay, and the same day launched pay per crawl in private beta. Our pattern is consistent with that setting being the main route. The neutral-bot control passed, so these are not blanket anti-scraping rules: the server singled out AI user agents. We cannot see anyone’s dashboard, so we cannot say how many owners switched the setting on knowingly.
Do Sites Block AI Training Bots or AI Search Bots?
Of the 640 AI-specific server blocks, 38.9% (249) hit only GPTBot and ClaudeBot, and 27.3% (175) hit every AI agent we tested. The other 58.1% let OpenAI’s search and user agents, OAI-SearchBot and ChatGPT-User, through.

So GPTBot is effectively blocked on 17.9% of measurable sites and OAI-SearchBot on 8.7%. The two do different jobs. A training crawler collects pages for training data. A search crawler builds the index an answer engine draws on, and a user-triggered agent fetches a page when someone asks about it (live retrieval). A policy on one says nothing about the other, yet a CDN rule can apply one without the owner choosing it.
Some owners do choose. In robots.txt alone, GPTBot is blocked while OAI-SearchBot stays open on 3.2% of affiliates (46 of 1,426) and 1.9% of operators (20 of 1,055). The reverse is nearly absent: one operator blocks OAI-SearchBot and leaves GPTBot open, and no other segment has any.
| Bot | Company | Purpose | Share of measurable sites blocking |
|---|---|---|---|
| GPTBot | OpenAI | Training | 17.9% |
| OAI-SearchBot | OpenAI | Search index | 8.7% |
| ChatGPT-User | OpenAI | User-triggered fetch | Not in headline measurement |
| ClaudeBot | Anthropic | Training | 18.7% |
| Claude-SearchBot | Anthropic | Search index | Not in headline measurement |
| Claude-User | Anthropic | User-triggered fetch | Not in headline measurement |
| PerplexityBot | Perplexity | Search index | 8.8% |
| Perplexity-User | Perplexity | User-triggered fetch | Not in headline measurement |
| Google-Extended | Training and grounding opt-out for Gemini, a robots.txt token | 2.6% (robots.txt only, n=3,469) | |
| CCBot | Common Crawl | Open web archive widely used for training | Not in headline measurement |
| Bytespider | ByteDance | Training | Not in headline measurement |
| Amazonbot | Amazon | Crawling for Amazon services | Not in headline measurement |
| Applebot-Extended | Apple | Opt-out token for Apple AI training | Not in headline measurement |
| anthropic-ai | Anthropic | Legacy training token | Not in headline measurement |
Block shares are robots.txt or server, n=3,590, unless noted. Bots marked “not in headline measurement” were not reported as a share.
Why Do 22 Sites Share One robots.txt?
One identical robots.txt appears on 22 domains, 18 operators and 4 affiliates, mostly white-label sites for the Indian market. It disallows dozens of AI bots, allows Googlebot by name and disallows everything else. A white-label vendor shipping one template would explain it: the operator inherits a policy it may never have read. Five smaller clusters of three to six domains share a robots.txt file in the same way. We do not name any of these sites.
Who Publishes llms.txt?
Crypto and B2B sites more often than casinos. Exchanges lead at 21.0%, then crypto projects at 19.4%, vendors at 18.5%, affiliates at 12.3% and operators at 3.0% (46 of 1,513). Across all 4,322 analysed sites, 11.1% serve a valid file and 1.4% serve llms-full.txt.

We count a file as valid when it returns a 200, is not HTML, and holds a markdown heading or a list of links. Empty and templated files count. Whether llms.txt changes anything for AI search is untested here. It costs little, so treat it as optional.
Did Blocking Cost These Sites Their AI Citations?
Ahrefs AI responses for 935 sites (ChatGPT, Perplexity, AI Overviews, AI Mode, Gemini, Copilot and Grok) show no clear link between blocking and citations per 1,000 organic visits.

The median ratios of blocking to open sites run from 0.60x to 0.99x in four segments, with every 95% interval crossing 1x. Operators show 0.08x, but half of blocking operators (50%) and 40% of open ones have zero citations, so the median rests on near-zero values.
Raw medians mislead because blocking sites differ from open ones. We controlled for traffic, Domain Rating and segment. The blocking coefficient is +0.099 for all platforms (SE 0.105), +0.152 for ChatGPT (SE 0.090) and +0.023 for Perplexity (SE 0.107). None is distinguishable from zero. Matched checks agree. OAI-SearchBot blocking against ChatGPT citations gives +0.025 (SE 0.113), and PerplexityBot blocking against Perplexity citations gives +0.013 (SE 0.141).
Google-Extended does not control AI Overviews, yet blocking it correlates with more AI Overviews citations (+0.388, SE 0.197). That is what a placebo test is meant to catch, and it is why we make no causal claims. Blocking sites appear to be bigger and more advanced than open ones, so any raw comparison is confounded.
This does not show that blocking costs citations, and it does not show that blocking is free. Three explanations are possible, and we tested none of them. Real bots arrive from verified IP addresses and may pass rules that stop a spoofed user agent. Citations in the Ahrefs panel may predate the blocks, since Cloudflare’s default began in July 2025. And answer engines also draw on other indexes, such as Bing’s.
What Should an Operator or Affiliate Do?
Check what your CDN does, then set separate policies for training and search. Five steps:
- Test the server as well as robots.txt. Request your homepage with a browser user agent and with an AI user agent, and compare the status codes.
- Decide training and search separately. Opting out of GPTBot is a statement about model training. OAI-SearchBot and ChatGPT-User concern live retrieval. We take no side on training.
- Write the policy down. The snippet below is an example of a split policy, not a recommendation.
- Expect robots.txt to lose to the WAF. If your CDN returns a 403 or a challenge, a permissive robots.txt changes nothing.
- Treat llms.txt as optional. It is cheap and unproven.
# Same page, two user agents. Compare the status codes.
curl -s -o /dev/null -w "%{http_code}\n" \
-A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.3; +https://openai.com/gptbot" \
https://example.com/
curl -s -o /dev/null -w "%{http_code}\n" \
-A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0 Safari/537.36" \
https://example.com/
# Example split policy: opt out of training, allow search and user fetch
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: PerplexityBot
User-agent: Claude-SearchBot
Allow: /
A spoofed test from your own network shows what a fake bot gets. Real bots from verified IP ranges can get a different answer, so confirm with log file analysis of what actually arrives. Full bot lists, user-agent strings and a copy-ready robots.txt are in our AI crawlers list. The GPTBot guide covers GPTBot against OAI-SearchBot in detail.
What Does This Study Not Show?
It shows how servers answer a spoofed user agent from one location. Real AI crawlers may get different answers elsewhere.
- One vantage point. Every request came from a French datacenter IP. France bans online casinos, which is why 34% of operators sat behind a wall. A US vantage point could change the wall share, though not necessarily the block shares among measurable sites.
- Spoofed user agents. Our IPs do not belong to OpenAI or Anthropic. We measure whether a server refuses the GPTBot user-agent string, not whether GPTBot can get in. The neutral-bot control rules out blanket bot blocking and does not rule out blocking of fake bots.
- Segment labels come from search signals. A manual check of 50 domains found about 6% junk, which we removed. Some brand clones are labeled as operators, and 533 segment conflicts are flagged in our data.
- The citation panel is not visibility. Ahrefs AI responses come from its own prompt panel, on a stratified subsample of 935 sites. Ratios of medians are unstable when many sites have zero citations.
- Homepage and robots.txt only. We did not test deeper pages, crawl budget, JavaScript rendering or SSR.
- The sample is 4,689 domains, not 5,000.
How Did We Run the Crawl?
We sampled from Ahrefs SERPs and brand queries in 15 countries (GB, US, CA, AU, NZ, DE, NL, SE, FI, IT, ES, BR, MX, IN and JP), CoinGecko’s top 1,000 coins, CoinGecko’s exchange list and iGaming vendor lists from our own audits. We dropped unreachable domains, 190 off-topic ones and junk found in manual checks. That left 4,322 analysed domains and 3,590 with measurable access.
For each domain we parsed robots.txt to RFC 9309, requested /llms.txt, and fetched the homepage with nine user agents: a browser and a neutral bot as controls, plus GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot and Perplexity-User. All requests used a Chrome TLS fingerprint. We ran two passes 30 hours apart and counted a server block only if it repeated. Ahrefs AI responses covered a stratified 935-site subsample. Intervals are Wilson 95% confidence intervals.
Check Which AI Crawlers Reach Your Site
The free AI Visibility checker shows whether ChatGPT, Perplexity, Claude and Gemini name your brand today. If a robots.txt or WAF review turns up a block you did not intend, our AI SEO for iGaming team can help you set the policy. We do generative engine optimization for casinos, affiliates and crypto projects.
Frequently Asked Questions
18.4% of measurable casino and betting operators (210 of 1,142, 95% CI 16.2-20.7%) block at least one of GPTBot, ClaudeBot, PerplexityBot or Google-Extended. Casino affiliates block more, at 27.4%.
Our data does not settle it. GPTBot is a training crawler, so blocking it is a decision about model training. Decide separately about OAI-SearchBot and ChatGPT-User, then check that your CDN enforces what your robots.txt says.
We found no clear link either way. After controlling for traffic and Domain Rating, the ChatGPT coefficient was +0.152 (SE 0.090), not distinguishable from zero. This is correlation on 935 sites.
GPTBot collects pages for training OpenAI’s models. OAI-SearchBot builds the search index behind ChatGPT’s answers. GPTBot is effectively blocked on 17.9% of the sites we measured and OAI-SearchBot on 8.7%.
Rate the article