HVAC

Is my website blocking ChatGPT and other AI from reading it?

Somebody set up your website years ago, and a file called robots.txt tells every automated visitor where it may go. Two lines in it can decide whether an AI search tool ever sees your service pages.

A heating and air company's metal shop open for business, its bay door rolled up on the van and boxed equipment, behind a chain-link fence whose gate is shut, with a notice on the gate showing the AI sparkle struck out, while the sparkle waits outside

Maybe. The answer depends on which AI company's crawler you name and what that crawler does. Most AI companies now run two or three separate robots: one that collects pages for training models, one that builds a search index, and one that fetches a page when a person asks a question. Blocking the first is a decision about your content. Blocking the second can take you out of that company's search answers. The fastest way to know is to open yourdomain.com/robots.txt in a browser and read it.

The crawlers, one company at a time

Each company documents its own robots, and the names are what go in your file:

  • OpenAI. Its crawler documentation lists OAI-SearchBot, used "to surface websites in search results in ChatGPT's search features", and says sites that disallow it will not be shown in ChatGPT search answers. GPTBot collects content that may be used for training; OpenAI says blocking it does not change search. ChatGPT-User visits a page when a person asks, and OpenAI says robots.txt rules may not apply to it.
  • Anthropic. Its help article separates ClaudeBot (training), Claude-SearchBot (search quality) and Claude-User (fetching for a person's question), and says blocking the search and user agents may reduce visibility.
  • Perplexity. Its bot guide describes PerplexityBot as the crawler that surfaces and links sites in Perplexity results, and says the separate Perplexity-User fetcher generally ignores robots.txt because a person asked.
  • Google. Its list of common crawlers says Google-Extended controls whether content may train future Gemini models and "does not impact a site's inclusion in Google Search". AI Overviews read what Googlebot indexed, so blocking Googlebot is how you would really disappear.

Three patterns to look for in the file

The blanket block. User-agent: * followed by Disallow: /. Usually left over from a staging site that went live. It shuts out everyone who obeys the file, search engines included, and it is the most expensive two lines a contractor can own.

The copied list. A long list of AI user agents pasted from an article about protecting content from AI, which blocks the search crawlers along with the training ones. For a news publisher that may be a considered choice. For a heating and air company whose whole goal is to be found, it rarely is.

The empty or missing file. No file means no instructions, and crawlers that obey the standard treat everything as allowed. That is fine, as long as nothing else is blocking them.

The block you did not write

Robots.txt is a request. Some hosting and security services enforce blocks at the network edge instead, where your file never enters into it. Cloudflare announced in July 2025 that it would block AI crawlers by default for new domains and ask each new customer whether to allow them. If your web person moved your site onto a service like that, your AI access may have been decided by a signup screen. Ask whoever manages your DNS and hosting to check the AI crawler settings and tell you, in writing, which bots are allowed.

What an HVAC company should usually allow

There is no single right file, and whether your content trains models is your call. The practical line for a company that wants to be named when someone asks for AC repair nearby: allow Googlebot and Bingbot, allow the search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot), decide on the training crawlers separately, and keep your service, service area and license pages out of any disallowed folder. If a page should not show up anywhere, a noindex tag on that page says so more precisely than a robots rule, because Google's AI features documentation names noindex and the snippet controls as the way to limit what its AI answers show.

A ten minute check

  • Open /robots.txt on your own domain and read every User-agent line.
  • Search the file for Disallow: / under any agent you want to be found by.
  • Ask your host or security provider whether AI crawler blocking is switched on.
  • Run the free check, which reads your robots.txt along with your homepage, sitemap and llms.txt.

Crawler access is the floor, not the win. The rest of what an assistant reads about a heating and air company is on the HVAC guide.

Not legal or tax advice. This describes how AI systems can read a heating and air company in public; licensing, permit, refrigerant, pricing and incentive rules are set by federal, state and local agencies, and they change.

Questions

If I block GPTBot, will my HVAC company disappear from ChatGPT?

Not according to OpenAI. Its crawler documentation says GPTBot is used for content that may train its models, and that disallowing it does not affect ChatGPT search. The crawler that decides whether a site can appear in ChatGPT search answers is OAI-SearchBot, and OpenAI says sites that disallow it will not be shown in those answers.

Does blocking Google-Extended remove me from Google AI Overviews?

Google says Google-Extended does not affect a site's inclusion in Google Search and is not a ranking signal. Its AI features documentation says a page must be indexed and eligible for a snippet to be a supporting link in AI Overviews or AI Mode, and the controls that limit this are nosnippet, data-nosnippet, max-snippet and noindex.