Can the AI systems that answer questions actually fetch your pages?
Three kinds of AI bots visit a website, and blocking the wrong kind removes you from answers without any warning.
A single line in a file most owners have never opened can decide whether ChatGPT is allowed to show your website at all. OpenAI states it plainly: sites that are opted out of OAI-SearchBot "will not be shown in ChatGPT search answers." No email arrives when that happens. The business simply stops being a candidate.
Access is the first gate of AI discoverability, and it is the only one that is a clean yes or no. The AI crawlers lexicon entry defines the bots. This page is the practical version: which ones matter to a local business, what each block costs, and how to check your own setup in about fifteen minutes.
Three jobs, not one
The major AI companies now document separate bots for separate purposes. Grouped by job, from their own pages:
Search bots: they decide whether you can be found
OAI-SearchBot (OpenAI) surfaces websites in ChatGPT's search features. PerplexityBot is, in Perplexity's words, designed to surface and link websites in its search results. Claude-SearchBot, per Anthropic's crawler page, indexes content for search, and blocking it "may reduce your site's visibility and accuracy in user search results." For Google's AI Overviews and AI Mode, the relevant crawler is ordinary Google Search: Google says a page only needs to be indexed and snippet eligible.
User fetchers: they visit when someone asks
ChatGPT-User, Perplexity-User and Claude-User fetch a page because a person asked a question. Anthropic says disabling Claude-User "prevents our system from retrieving your content in response to a user query." OpenAI notes that for ChatGPT-User, because the action is user initiated, "robots.txt rules may not apply," and Perplexity says Perplexity-User "generally ignores robots.txt rules."
Training bots: they feed future models
GPTBot crawls content that "may be used in training" OpenAI's models. ClaudeBot collects content that could contribute to Anthropic's training. Google-Extended is a token you name in robots.txt that controls use of your content for training future Gemini models, and Google states it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal."
What each block actually costs
- Blocking a search bot removes you from that assistant's search results. For a plumber in Oxnard who wants calls, this is almost never the intended outcome.
- Blocking a training bot keeps your future content out of that company's training data. It does not remove you from its search. This is a legitimate choice, and it is yours to make.
- Blocking a user fetcher can stop an assistant reading your page at the moment a customer asks about you, where the fetcher honors robots.txt.
The common mistake is a blanket rule written to stop "AI scraping" that sweeps out the search bots along with the training ones.
The block you did not set yourself
If your site sits behind Cloudflare, perhaps set up by a web designer you no longer work with, the rules that matter may not be in robots.txt at all. Cloudflare's legacy Block AI bots setting blocks verified bots classified as crawling for AI training, and it is being replaced by AI bot policies that separate Search, Agent and Training. Cloudflare's changelog says that from September 15, 2026, new domains onboarding to Cloudflare block Training and Agent bots on pages that display ads, while Search stays allowed. It also notes that crawlers combining search and training are affected by settings that block training.
None of that is wrong in itself. The point is that the effective rules for your site may live in a dashboard you have never logged into, and security plugins or hosting firewalls can filter bots the same way.
How to check, in order
- Read your robots.txt. Open yourdomain.com/robots.txt in a browser. Look for the agent names above followed by
Disallow: /, and for aUser-agent: *group that disallows everything. - Check the edge. If you use Cloudflare, look under Security settings for the AI bot options, and in AI Crawl Control, where the Crawlers tab lists each AI crawler with an Allow or Block action.
- Ask whoever runs the site. One question covers the rest: "Do we block any bots by user agent anywhere, including plugins and the host's firewall?"
- Look for evidence. If you can see server logs, search them for OAI-SearchBot, PerplexityBot and Claude-SearchBot. Visits with a normal response mean the door is open.
A reasonable default for a business that wants customers: allow every search bot, allow the user fetchers, and decide about training deliberately rather than by accident. Once the door is open, the next question is whether the bots can read what they fetch, covered in content AI cannot see. The full AI discoverability guide covers every other signal, and you can run the free AIOInsights check to see which named AI agents your robots.txt currently allows.
Questions
What happens if a website blocks OAI-SearchBot?
OpenAI's crawler documentation says that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links. Blocking GPTBot is different: it tells OpenAI the content should not be used to train its foundation models, and OpenAI documents it as a separate bot from the one ChatGPT search uses.
Can a business block AI training but still appear in AI answers?
Yes. OpenAI, Anthropic and Google each document training separately from search: GPTBot, ClaudeBot and the Google-Extended token cover training, while OAI-SearchBot, Claude-SearchBot and Google Search cover search. A small business can disallow the training agents in robots.txt and keep the search agents allowed, so its pages can still be found and cited.
More on AI discoverability
This page is part of AI Discoverability, the AIOInsights guide to whether AI systems can find, read and name a business.
- How do AI assistants find a business to name in an answer?
- Which content on your site is invisible to AI retrieval?
- Why one sentence with your name, city and service does so much work
- What happens when your business details disagree across the web?
- The sites AI reads about you, and why your own site is often not the one it cites
- Does structured data help AI find your business?
- How to measure AI discoverability without fooling yourself
- A first week plan for AI discoverability
- Where AI discoverability work falls short