All articles
AI Search

Should Your Business Block AI Crawlers? A Practical Answer

August 17, 2026 6 min readVanguard Media
Should Your Business Block AI Crawlers? A Practical Answer

Somewhere in the settings of your website sits a plain text file called robots.txt, and right now it is making a decision on your behalf: whether the AI systems your future customers use can read your site. Some businesses blocked AI crawlers years ago on a developer's recommendation and forgot about it. Others have never looked. With assistants like ChatGPT and Gemini now fielding questions like "who installs heat pumps near me," that forgotten file has become a business decision worth revisiting. Here is a balanced look at who the crawlers are, what blocking actually does, and how to check where you stand.

Who are the AI crawlers visiting your site?

Several distinct bots matter, and they serve different purposes.

GPTBot is OpenAI's crawler. Content it collects can inform future ChatGPT models, and OpenAI operates related agents for real-time browsing when users ask questions that need current information.

ClaudeBot performs the equivalent role for Anthropic's Claude assistant, gathering web content that shapes what Claude knows about businesses and topics.

PerplexityBot feeds Perplexity, a search product that answers questions with cited sources. Because Perplexity links to the pages it draws from, being readable there can send referral visitors directly to your site.

Google-Extended is the one businesses most often misunderstand. It is not Googlebot. Blocking Google-Extended tells Google not to use your content for Gemini and related AI training, while your ordinary search rankings continue to depend on Googlebot, which keeps crawling as usual. Note the boundary, though: Google has said AI Overviews in Search are part of ordinary search, so opting out of Google-Extended does not remove you from those either. The two systems are documented in Google's crawler documentation.

Each of these respects robots.txt directives, which is precisely why the file deserves your attention: the rules in it are actually followed.

What does blocking mean for a publisher?

For businesses whose product is the content itself, blocking has a coherent logic. A news outlet, a research firm, or a site monetized by page views faces a real trade: if an AI assistant absorbs and summarizes their articles, readers may get the value without ever visiting, and the publisher captures nothing. Licensing negotiations between publishers and AI companies exist precisely because that content has standalone value.

If your website sells information, whether journalism, original research, paid courses, or proprietary data, restricting AI crawlers is a defensible position while the economics get sorted out. The content is the asset, and controlling access to an asset is ordinary business practice.

What does blocking mean for a local service business?

The calculation inverts almost completely. A plumber's website is not the product; it is an advertisement for the product. Nobody reads a furnace repair service page for its literary value. The page exists so that when a homeowner in Burnaby needs help, they find the business, trust it, and call.

Seen that way, an AI assistant summarizing your service page is not theft of value; it is distribution. When ChatGPT tells a user "this company handles emergency plumbing in New Westminster and the surrounding areas," it is doing exactly what you built the website to do. Blocking the crawler cuts off that channel. The assistant cannot recommend a business it cannot read, so the recommendation goes to a competitor whose site was open.

This is why the default advice for local service businesses differs from the advice for publishers. The publisher risks losing the value of their content. The service business risks losing the customer entirely.

Why do most local businesses benefit from allowing AI crawlers?

Three reasons stack up.

First, AI answers increasingly sit where buying decisions start. Questions that used to produce ten blue links now often produce a short, generated answer naming two or three businesses. Being ineligible for that shortlist is a growing cost, and it compounds as more customers adopt assistants.

Second, allowing crawlers costs a local business essentially nothing. Your service pages contain no proprietary secrets. Your prices, if listed, are public anyway. The information an AI would extract is the information you actively want circulated.

Third, openness pairs with the rest of your visibility work. Crawl access is one ingredient among several, alongside entity consistency, reviews, and structured data, that determine whether assistants mention you. Our AI search optimization guide covers the full set, and our post on llms.txt and schema for AI search goes deeper on the technical files involved. Crawl access is the gate in front of all of it: get the gate wrong and the rest of the work goes unread.

How do you check your robots.txt?

You can do this in two minutes without any tools. Type your domain into a browser followed by /robots.txt, for example yourbusiness.ca/robots.txt. The file that appears is plain text and readable by anyone.

Look for blocks like this:

User-agent: GPTBot
Disallow: /

That pair of lines tells OpenAI's crawler it may not read any page on your site. Repeat the scan for ClaudeBot, PerplexityBot, Google-Extended, and a catch-all like User-agent: * followed by Disallow: /, which blocks everything including ordinary search engines. If you find that last pattern on a live business site, fix it urgently, because it suppresses your normal Google visibility too. You can verify how Google sees your site in Google Search Console.

Two common surprises. Some website platforms and security plugins ship with AI-blocking rules enabled by default, so you may be blocking crawlers without anyone having decided to. And some firewall services block bots at the network level, which robots.txt will not show; if assistants seem unable to see content you know is open, the firewall configuration is the next place to look.

Is selective blocking a reasonable middle ground?

Yes, and it is underused. Robots.txt rules can target specific bots and specific directories, so the choice is not all-or-nothing.

A business with a mixed site can open its service pages, location pages, and FAQ content to every AI crawler while disallowing a members-only area, internal documents, or a premium resource library. A consultancy that sells research reports can expose everything except the reports directory. The syntax is simply a Disallow line scoped to a path rather than to the whole site.

You can also distinguish between bots. Some businesses allow crawlers attached to products that cite and link sources while restricting pure training crawlers. That is a finer-grained call. What matters is that the configuration reflects a decision rather than a default nobody remembers making.

If reviewing crawler directives, schema, and site structure feels outside your comfort zone, this is standard ground for our SEO service, and it is also part of how we approach AI integration work, where making your business legible to AI systems is treated as infrastructure rather than an afterthought.

Frequently asked questions

Will blocking GPTBot hurt my Google rankings?

No. Google's ordinary search crawler is Googlebot, and rules aimed at GPTBot, ClaudeBot, or PerplexityBot have no effect on it. Even Google-Extended, Google's own AI-related token, is separate from ranking. The risk of blocking AI crawlers is to your presence in AI assistants, not to your classic search positions.

If I unblock AI crawlers today, when will assistants see my site?

There is no fixed schedule. Each AI company crawls and refreshes on its own cadence, and some assistants also browse live when users ask questions. Treat unblocking as opening a door: necessary immediately, rewarded gradually as the various systems revisit your site.

Can AI companies ignore my robots.txt?

Robots.txt is a convention rather than an enforcement mechanism, and the major crawlers named in this post publicly commit to honouring it. Bad actors exist on the web regardless of what your file says. The practical decision is about the mainstream assistants your customers actually use, and those respect the rules.

Should I block AI crawlers on my staging or development site?

Yes, and block every crawler there, not just AI ones. Unfinished sites leaking into search results or AI answers creates confusion about your real business information. Keep development environments fully disallowed or behind authentication, and keep your production site deliberately open.

Not sure what your robots.txt is currently telling AI crawlers? Request a free SEO audit from Vanguard Media and we will check your crawl directives, along with the rest of your search foundation, and tell you exactly where you stand.

Want results like these for your business?

Book a free strategy session and we will map out your plan.

Book a Strategy Session
Call now Book a Call