AThe AI Visibility Index
← Guides

Should you block AI crawlers — or allow them?

If you want AI assistants to recommend your business, allow at least the retrieval crawlers — blocking them makes you uncitable in live-search answers. Blocking is a defensible choice for content businesses protecting their product; for most brands and shops it silently trades away visibility for nothing.

Updated 22 July 2026

Every site now chooses, knowingly or by default, whether the AI systems that answer buyers' questions are allowed to read it. The choice is made in robots.txt, it is legible to anyone, and both options are defensible — for different kinds of site. What is not defensible is making it by accident.

Know which crawlers you're deciding about

They are not one thing. Three species matter:

  • Training crawlers (GPTBot, Google-Extended, ClaudeBot, CCBot, Applebot-Extended…) gather text for future model training. Blocking them affects what the next generation of models remembers — slowly, and unverifiably.
  • Retrieval crawlers and user-fetchers (OAI-SearchBot, ChatGPT-User, Perplexity-User, Claude-User…) fetch pages at answer time, when an assistant is actively composing a response to a live question. Blocking these makes you uncitable today.
  • Ordinary search crawlers (Googlebot, Bingbot) now also feed AI features — Google's AI Overviews and ChatGPT search's Bing-backed retrieval — and blocking those has classic-SEO consequences well beyond AI.

Most "should we block AI?" debates conflate all three. The practical stakes are concentrated in the second group.

The case for blocking — and who it belongs to

If your content is your product — a publisher, a paywalled archive, a dataset business — letting models train on it or quote it wholesale may genuinely cannibalise what you sell. That is a real position with real trade-offs, and several major publishers hold it, typically while negotiating licensing.

Notice who that argument serves. It does not serve a business whose site exists to be found: a shop, a service, a venue, a brand. For them the content is marketing, not product — and blocking the systems buyers now ask for recommendations is blocking the shop window.

What blocking actually costs

We measure this. A site that refuses AI crawlers cannot be read, so it cannot be cited, so in retrieval-led answers it effectively does not exist — and the penalty in AI answers is always omission, not criticism. The brands it competes with remain readable. Nothing announces the loss; there is no error message for the recommendation that went to a rival because your pages could not be fetched.

A sane default policy

  1. Allow retrieval and user-fetch crawlers unless your content is your product.
  2. Decide training crawlers on principle, knowing the effect is slow either way — and that a brand absent from training data is leaning entirely on retrieval to exist in answers.
  3. Never block by accident. Audit robots.txt and your CDN's bot rules; plenty of sites block AI crawlers through a security vendor's default toggle without ever choosing to.
  4. State your policy where machines read it — robots.txt, and optionally llms.txt.

Run a free scan to see exactly which AI crawlers your site currently allows — it is checked in the first seconds, because nothing else matters if the answer is "none".

Common questions

Does blocking GPTBot stop ChatGPT recommending my business?
Not directly — GPTBot gathers training data. But ChatGPT's live search retrieves through its own fetchers and Bing's index, so blocking OAI-SearchBot, ChatGPT-User or Bingbot does stop it reading and citing your pages at answer time. Many sites block all of them together without realising the second kind is the one costing recommendations.
Can I allow AI crawlers on some pages and not others?
Yes — robots.txt rules are per-path, so you can expose your buying guides and product pages while keeping, say, a members' area closed. For most brands the pages worth exposing are exactly the ones that answer a buyer's question: what you sell, to where, at what terms.
My security/CDN provider blocks AI bots by default — does that matter?
It is one of the most common silent visibility losses we see. Bot-protection defaults frequently classify AI crawlers as unwanted automation, so a site's owner believes they are open while every AI fetch gets a 403. Check the CDN's bot rules, not just robots.txt — then verify with a scan.

Read next

See how AI reads your brand

Get your AVI Score out of 100 — and a clear plan to rank higher in AI answers.