Check if your brand is visible to AI Search

Who Lets AI In? AI-Crawler Access Across 4,689 Casino, Affiliate and Crypto Websites

Casino affiliates block AI crawlers 1.5x more often than the casinos they promote.

Published: October 7, 2026 – Updated: October 8, 2026

12 minutes to read

ICODA crawled 4,689 casino, affiliate and crypto websites on 6 and 7 October 2026 to see which ones turn AI crawlers away. We analysed 4,322 of them and could measure access on 3,590. Of those, 20.9% block at least one of GPTBot, ClaudeBot, PerplexityBot or Google-Extended. Affiliates block most. Most blocks come from servers and CDNs, and few are written into robots.txt files.


Key Findings

  • One site in five blocks a major AI crawler. 20.9% of 3,590 measurable sites (751 sites, 95% CI 19.6-22.3%) refuse at least one of the four core crawlers.
  • Affiliates block 1.5x more often than operators. 27.4% of affiliates against 18.4% of casino and betting operators, with intervals that do not overlap.
  • Servers do most of the blocking. 19.2% of sites show an AI-specific server block, while 3.5% have an AI-specific robots.txt block.
  • Cloudflare is the strongest single factor. 25.9% of sites behind Cloudflare block, against 10.9% of sites elsewhere.
  • Most server blocks hit training bots and miss ChatGPT’s search bot. 58.1% of server blocks let OpenAI’s search and user agents through.
  • No clear link between blocking and AI citations. On 935 sites, the blocking coefficient is indistinguishable from zero once traffic and Domain Rating are controlled. This is correlation only.

Do Casino Affiliates Block AI Crawlers More Than the Casinos They Promote?

27.4% of affiliate sites (394 of 1,438) block at least one core AI crawler, against 18.4% of operators (210 of 1,142). The 95% intervals, 25.2-29.8% and 16.2-20.7%, do not overlap, and the ratio is 1.49x.

We expected the opposite. Affiliates depend on search traffic and publish structured content, so we guessed they would be the more AI-ready segment. Half of that held up: affiliates publish an llms.txt file four times as often as operators (12.3% against 3.0%), and 59.4% carry Organization schema on the homepage against 20.4% of operators. The other half did not, because they also block more often.

Traffic widens the gap. Affiliate sites that block at least one core crawler hold 44.5% of all affiliate organic traffic, though they are only 27.4% of affiliate sites. Operators run the other way: 9.9% of operator traffic sits on blocking sites, which are 18.4% of operator sites. The blockers are the large affiliates and the small operators. Traffic shares use Ahrefs organic traffic estimates.


How Many Casino and Crypto Sites Block AI Crawlers?

Only affiliates sit above the 20.9% average across segments. Crypto exchanges come second at 19.1%.

SegmentMeasurable sitesBlocks any core crawler (95% CI)Via robots.txtVia serverValid llms.txt
Casino affiliates1,43827.4% (25.2-29.8)4.8%24.7%12.3%
Crypto exchanges20919.1% (14.4-25.0)1.4%18.8%21.0%
Casino and betting operators1,14218.4% (16.2-20.7)4.8%16.9%3.0%
iGaming B2B vendors18213.7% (9.5-19.5)2.8%12.8%18.5%
Crypto projects (CoinGecko top 1,000)61913.2% (10.8-16.1)1.7%12.6%19.4%
All segments3,59020.9% (19.6-22.3)3.9%19.2%11.1%

Share of measurable sites. The robots.txt column counts any core crawler blocked, including files that also shut out Googlebot. The server column counts AI-specific blocks only. The llms.txt column uses all analysed sites in each segment.

Bar chart of the share of sites that block AI crawlers (GPTBot, ClaudeBot, PerplexityBot or Google-Extended): casino affiliates 27%, crypto exchanges 19%, casino operators 18%, iGaming vendors 14%, crypto projects 13%.
Affiliates block AI crawlers most often, at 27% of measurable sites, and crypto projects least, at 13%.

By crawler, GPTBot is effectively blocked on 17.9% of measurable sites, ClaudeBot on 18.7%, PerplexityBot on 8.8% and OAI-SearchBot on 8.7%. Google-Extended is blocked in robots.txt on 2.6%. Sites that block hold 37.5% of exchange organic traffic but only 0.9% of crypto-project traffic, so blocking among crypto projects looks like a small-site habit.

Operators were the hardest segment to measure. Of the 1,513 we analysed, 516 (34.1%) put a wall in front of our crawler that stopped both a browser and a neutral bot. We read their robots.txt files but could not test the server. Casino sites often geo-block, and our crawler sat in France.


Is robots.txt Doing the Blocking?

An AI-specific robots.txt block, one that disallows AI crawlers and leaves Googlebot open, appears on 3.5% of measurable sites (120 of 3,469). An AI-specific server block appears on 19.2% (640 of 3,330).

A server block means the site served a normal page to a browser and to our neutral bot, then returned a 403 or a challenge page to an AI user agent. Depending on segment, these server blocks are 4 to 13 times more common than AI-specific robots.txt blocks. Of the server blocks we measured twice, 628 repeated 30 hours later and 16 vanished, so the pattern holds up.

An owner who checks only robots.txt can see a clean file and still be blocked. The file, parsed to RFC 9309 rules (Allow and Disallow groups, longest match wins), states the policy. The WAF enforces or overrides it. Few sites use the newer opt-out signals: 2.6% (113 of 4,322) publish Content-Signal lines, the ai-train, search and ai-input vocabulary Cloudflare promotes, and 3 carry a noai meta tag.


How Much Does Cloudflare Matter?

Sites behind Cloudflare block at least one core crawler 25.9% of the time (621 of 2,397), against 10.9% of sites elsewhere (130 of 1,193), a ratio of 2.4x. Cloudflare sits in front of 66.8% of the measurable sites.

Dot chart of sites that block AI crawlers behind Cloudflare versus other hosting: casino affiliates 32% against 17%, crypto exchanges 26% against 4%, operators 23% against 8%, vendors 16% against 12%, crypto projects 19% against 6%.
Sites behind Cloudflare block AI crawlers more often than sites elsewhere in every segment, 26% against 11% overall.

The gap appears in each segment. Exchanges block 26.2% behind Cloudflare and 4.4% elsewhere, crypto projects 18.7% and 5.7%, operators 22.5% and 8.4%, and affiliates 31.9% and 16.8%. B2B vendors show the same direction at a smaller size (16.2% against 11.8%), and their intervals overlap. Of the 751 sites that block, 621 (83%) sit behind Cloudflare.

On 1 July 2025 Cloudflare changed its default to block AI crawlers unless they pay, and the same day launched pay per crawl in private beta. Our pattern is consistent with that setting being the main route. The neutral-bot control passed, so these are not blanket anti-scraping rules: the server singled out AI user agents. We cannot see anyone’s dashboard, so we cannot say how many owners switched the setting on knowingly.


Do Sites Block AI Training Bots or AI Search Bots?

Of the 640 AI-specific server blocks, 38.9% (249) hit only GPTBot and ClaudeBot, and 27.3% (175) hit every AI agent we tested. The other 58.1% let OpenAI’s search and user agents, OAI-SearchBot and ChatGPT-User, through.

Stacked bar chart of server-side AI blocks by segment. Training bots only (GPTBot and ClaudeBot) account for 36% of affiliate blocks, 43% for crypto exchanges and 44% for operators. Every AI agent, including OAI-SearchBot, is 21% to 49%.
Six in ten server-side AI blocks still let ChatGPT’s search bot in.

So GPTBot is effectively blocked on 17.9% of measurable sites and OAI-SearchBot on 8.7%. The two do different jobs. A training crawler collects pages for training data. A search crawler builds the index an answer engine draws on, and a user-triggered agent fetches a page when someone asks about it (live retrieval). A policy on one says nothing about the other, yet a CDN rule can apply one without the owner choosing it.

Some owners do choose. In robots.txt alone, GPTBot is blocked while OAI-SearchBot stays open on 3.2% of affiliates (46 of 1,426) and 1.9% of operators (20 of 1,055). The reverse is nearly absent: one operator blocks OAI-SearchBot and leaves GPTBot open, and no other segment has any.

BotCompanyPurposeShare of measurable sites blocking
GPTBotOpenAITraining17.9%
OAI-SearchBotOpenAISearch index8.7%
ChatGPT-UserOpenAIUser-triggered fetchNot in headline measurement
ClaudeBotAnthropicTraining18.7%
Claude-SearchBotAnthropicSearch indexNot in headline measurement
Claude-UserAnthropicUser-triggered fetchNot in headline measurement
PerplexityBotPerplexitySearch index8.8%
Perplexity-UserPerplexityUser-triggered fetchNot in headline measurement
Google-ExtendedGoogleTraining and grounding opt-out for Gemini, a robots.txt token2.6% (robots.txt only, n=3,469)
CCBotCommon CrawlOpen web archive widely used for trainingNot in headline measurement
BytespiderByteDanceTrainingNot in headline measurement
AmazonbotAmazonCrawling for Amazon servicesNot in headline measurement
Applebot-ExtendedAppleOpt-out token for Apple AI trainingNot in headline measurement
anthropic-aiAnthropicLegacy training tokenNot in headline measurement

Block shares are robots.txt or server, n=3,590, unless noted. Bots marked “not in headline measurement” were not reported as a share.


Why Do 22 Sites Share One robots.txt?

One identical robots.txt appears on 22 domains, 18 operators and 4 affiliates, mostly white-label sites for the Indian market. It disallows dozens of AI bots, allows Googlebot by name and disallows everything else. A white-label vendor shipping one template would explain it: the operator inherits a policy it may never have read. Five smaller clusters of three to six domains share a robots.txt file in the same way. We do not name any of these sites.


Who Publishes llms.txt?

Crypto and B2B sites more often than casinos. Exchanges lead at 21.0%, then crypto projects at 19.4%, vendors at 18.5%, affiliates at 12.3% and operators at 3.0% (46 of 1,513). Across all 4,322 analysed sites, 11.1% serve a valid file and 1.4% serve llms-full.txt.

Bar chart of valid llms.txt files by segment: crypto exchanges 21%, crypto projects 19%, iGaming vendors 19%, casino affiliates 12% and casino operators 3%. Lines show 95% confidence intervals.
Only 3% of casino operators publish an llms.txt file, against 12% of affiliates and 19-21% of crypto and vendor sites.

We count a file as valid when it returns a 200, is not HTML, and holds a markdown heading or a list of links. Empty and templated files count. Whether llms.txt changes anything for AI search is untested here. It costs little, so treat it as optional.


Did Blocking Cost These Sites Their AI Citations?

Ahrefs AI responses for 935 sites (ChatGPT, Perplexity, AI Overviews, AI Mode, Gemini, Copilot and Grok) show no clear link between blocking and citations per 1,000 organic visits.

Dot chart of median AI citations per 1,000 organic visits, sites that block AI crawlers divided by open sites: casino affiliates 0.99x, crypto exchanges 0.87x, crypto projects 0.80x, vendors 0.60x, operators 0.08x. Every interval crosses 1x.
No segment shows a clear difference in AI citations between blocking and open sites, and every interval includes 1x.

The median ratios of blocking to open sites run from 0.60x to 0.99x in four segments, with every 95% interval crossing 1x. Operators show 0.08x, but half of blocking operators (50%) and 40% of open ones have zero citations, so the median rests on near-zero values.

Raw medians mislead because blocking sites differ from open ones. We controlled for traffic, Domain Rating and segment. The blocking coefficient is +0.099 for all platforms (SE 0.105), +0.152 for ChatGPT (SE 0.090) and +0.023 for Perplexity (SE 0.107). None is distinguishable from zero. Matched checks agree. OAI-SearchBot blocking against ChatGPT citations gives +0.025 (SE 0.113), and PerplexityBot blocking against Perplexity citations gives +0.013 (SE 0.141).

Google-Extended does not control AI Overviews, yet blocking it correlates with more AI Overviews citations (+0.388, SE 0.197). That is what a placebo test is meant to catch, and it is why we make no causal claims. Blocking sites appear to be bigger and more advanced than open ones, so any raw comparison is confounded.

This does not show that blocking costs citations, and it does not show that blocking is free. Three explanations are possible, and we tested none of them. Real bots arrive from verified IP addresses and may pass rules that stop a spoofed user agent. Citations in the Ahrefs panel may predate the blocks, since Cloudflare’s default began in July 2025. And answer engines also draw on other indexes, such as Bing’s.


What Should an Operator or Affiliate Do?

Check what your CDN does, then set separate policies for training and search. Five steps:

  1. Test the server as well as robots.txt. Request your homepage with a browser user agent and with an AI user agent, and compare the status codes.
  2. Decide training and search separately. Opting out of GPTBot is a statement about model training. OAI-SearchBot and ChatGPT-User concern live retrieval. We take no side on training.
  3. Write the policy down. The snippet below is an example of a split policy, not a recommendation.
  4. Expect robots.txt to lose to the WAF. If your CDN returns a 403 or a challenge, a permissive robots.txt changes nothing.
  5. Treat llms.txt as optional. It is cheap and unproven.
# Same page, two user agents. Compare the status codes.
curl -s -o /dev/null -w "%{http_code}\n" \
  -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.3; +https://openai.com/gptbot" \
  https://example.com/
curl -s -o /dev/null -w "%{http_code}\n" \
  -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0 Safari/537.36" \
  https://example.com/
# Example split policy: opt out of training, allow search and user fetch
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: PerplexityBot
User-agent: Claude-SearchBot
Allow: /

A spoofed test from your own network shows what a fake bot gets. Real bots from verified IP ranges can get a different answer, so confirm with log file analysis of what actually arrives. Full bot lists, user-agent strings and a copy-ready robots.txt are in our AI crawlers list. The GPTBot guide covers GPTBot against OAI-SearchBot in detail.


What Does This Study Not Show?

It shows how servers answer a spoofed user agent from one location. Real AI crawlers may get different answers elsewhere.

  • One vantage point. Every request came from a French datacenter IP. France bans online casinos, which is why 34% of operators sat behind a wall. A US vantage point could change the wall share, though not necessarily the block shares among measurable sites.
  • Spoofed user agents. Our IPs do not belong to OpenAI or Anthropic. We measure whether a server refuses the GPTBot user-agent string, not whether GPTBot can get in. The neutral-bot control rules out blanket bot blocking and does not rule out blocking of fake bots.
  • Segment labels come from search signals. A manual check of 50 domains found about 6% junk, which we removed. Some brand clones are labeled as operators, and 533 segment conflicts are flagged in our data.
  • The citation panel is not visibility. Ahrefs AI responses come from its own prompt panel, on a stratified subsample of 935 sites. Ratios of medians are unstable when many sites have zero citations.
  • Homepage and robots.txt only. We did not test deeper pages, crawl budget, JavaScript rendering or SSR.
  • The sample is 4,689 domains, not 5,000.

How Did We Run the Crawl?

We sampled from Ahrefs SERPs and brand queries in 15 countries (GB, US, CA, AU, NZ, DE, NL, SE, FI, IT, ES, BR, MX, IN and JP), CoinGecko’s top 1,000 coins, CoinGecko’s exchange list and iGaming vendor lists from our own audits. We dropped unreachable domains, 190 off-topic ones and junk found in manual checks. That left 4,322 analysed domains and 3,590 with measurable access.

For each domain we parsed robots.txt to RFC 9309, requested /llms.txt, and fetched the homepage with nine user agents: a browser and a neutral bot as controls, plus GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot and Perplexity-User. All requests used a Chrome TLS fingerprint. We ran two passes 30 hours apart and counted a server block only if it repeated. Ahrefs AI responses covered a stratified 935-site subsample. Intervals are Wilson 95% confidence intervals.


Check Which AI Crawlers Reach Your Site

The free AI Visibility checker shows whether ChatGPT, Perplexity, Claude and Gemini name your brand today. If a robots.txt or WAF review turns up a block you did not intend, our AI SEO for iGaming team can help you set the policy. We do generative engine optimization for casinos, affiliates and crypto projects.


Frequently Asked Questions

18.4% of measurable casino and betting operators (210 of 1,142, 95% CI 16.2-20.7%) block at least one of GPTBot, ClaudeBot, PerplexityBot or Google-Extended. Casino affiliates block more, at 27.4%.

Our data does not settle it. GPTBot is a training crawler, so blocking it is a decision about model training. Decide separately about OAI-SearchBot and ChatGPT-User, then check that your CDN enforces what your robots.txt says.

We found no clear link either way. After controlling for traffic and Domain Rating, the ChatGPT coefficient was +0.152 (SE 0.090), not distinguishable from zero. This is correlation on 935 sites.

GPTBot collects pages for training OpenAI’s models. OAI-SearchBot builds the search index behind ChatGPT’s answers. GPTBot is effectively blocked on 17.9% of the sites we measured and OAI-SearchBot on 8.7%.

Share with

Rate the article

4.7/5 - (20 votes)