Free Site Review
Back to insights
AI & SearchAugust 2026·6 min read·By Dave Collins

Is your website blocking the AI engines? How to check in five minutes

Some agencies go unnamed in AI answers for a dull reason: their own website turns the engines away at the door. Nobody decided it. Here is how to check yours in five minutes, and the change coming on 15 September that starts blocking Googlebot for sites that did nothing but tick a box.

We spend a lot of time measuring how agencies show up in AI answers, and most of the reasons an agency goes unnamed are the ones you would expect. Thin reviews. A Google Business Profile nobody has touched in three years. Rivals who have simply been recommended more often.

Then there is the dull one. Sometimes the engines never named the agency because they were never allowed to read the site. Nobody decided this. It arrived in a plugin, or a hosting default, or a tick box somebody pressed when blocking AI was in the news, and it has been quietly working ever since. It is worth five minutes of your morning to rule out.

The two places it hides

The first is a file called robots.txt, which sits at the root of your website and tells automated visitors what they may read. Yours is public, and so is everybody else's. Type your domain followed by /robots.txt into a browser and you will see it. What you are looking for is any line naming an AI crawler followed by Disallow. The names worth knowing, as at August 2026: OpenAI runs GPTBot for training, OAI-SearchBot for search, and ChatGPT-User when somebody in ChatGPT asks a question and it goes and fetches a page. Anthropic runs ClaudeBot, Claude-SearchBot and Claude-User, split the same way. Perplexity runs PerplexityBot for indexing and Perplexity-User for retrieval. Google runs Google-Extended, which is not the same thing as Googlebot and does not do what most people think.

The split matters more than the names. Blocking ClaudeBot stops Anthropic training on your content. It does not stop Claude-SearchBot indexing you, and it does not stop Claude-User fetching your valuation page for somebody who is asking about you right now. A site that blocks all three has removed itself from the answer. A site that blocks only the training bot has made a defensible choice and kept its seat. The second place is your hosting or your CDN, which is where this gets less obvious, because that setting is not in your robots.txt and you cannot see it by looking at your website.

The 15 September change worth knowing about

Cloudflare sits in front of a large share of the web, very possibly yours, and it has offered a one-click "Block AI bots" option for a couple of years. Plenty of people pressed it. It felt like the responsible thing to do at the time. On 1 July 2026 Cloudflare published its new AI traffic options and set a date of 15 September 2026 for two changes. The first is a new default, and it is narrower than the headlines suggested: for new domains joining Cloudflare, training and agent crawlers will be blocked by default on the pages that display ads, while search crawlers stay allowed. Estate agency websites do not usually carry advertising, so for most agents that part is a non-event.

The second change is the one to pay attention to. From the same date, crawlers that do more than one job are judged on all of their behaviours, and the most restrictive rule wins. In Cloudflare's own words, multi-purpose crawlers "such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service)". Read that again with your own site in mind. If anyone at your agency, or at your web company, ever ticked "Block AI bots", then from 15 September Googlebot is on the blocked list too, because Googlebot also crawls for training. Not your AI visibility. Your Google visibility. Cloudflare has said owners can opt out of the new defaults in their security settings at any point before 15 September, and that it will keep notifying customers meanwhile. Those notifications go to the email address on the Cloudflare account, which in a lot of agencies is a developer who left, or an address nobody reads.

So, the five minutes

One. Open yourdomain.co.uk/robots.txt and read it. If you see GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Claude-SearchBot, CCBot or anything similar sitting above a Disallow line, somebody made that decision. Find out who, and whether they would make it again today. Two. Ask whoever hosts your site one question: are we behind Cloudflare, and is "Block AI bots" or any AI bot blocking switched on? If the answer is yes, ask them to review it before 15 September rather than after. Three. If you have a security plugin on a WordPress site, check whether it has an AI or bot blocking feature and what it is set to, because several enable something by default and describe it as protection. Four. While you are in there, check you are not blocking the fetch bots. Those are the ones that turn up when a real person has asked a real question about you, which is the closest thing to a warm lead in this whole subject.

The Google one that confuses everybody

Google-Extended is not a search setting. It governs whether your content can be used for training and grounding in Gemini apps and the Gemini API, and Google's own documentation says it does not affect inclusion in Search and is not a ranking signal. Which means blocking it does not remove you from AI Overviews or AI Mode, because those are features inside Google Search rather than separate products, and they draw on the ordinary index. The only levers that affect what they can show are the snippet controls, nosnippet, data-nosnippet, max-snippet and noindex, and every one of those costs you something in ordinary search results too.

The practical upshot for an agent is unglamorous but clear. You cannot hide from Google's AI answers and stay in Google. The choice on the table is not whether to appear, it is whether what appears is accurate and current, which is a content and structured data question rather than a blocking one.

Where we land on it

Our own position, for what it is worth, is that search and fetch crawlers should be allowed to read an estate agency website, because those are the ones that put your name in front of somebody who is asking. Training crawlers are a fair debate, and we do not think an agency is wrong to block them. What is worth avoiding is blocking all of it by accident, then paying somebody to work out why the engines never mention you. If the engines are quiet about you and your robots.txt is clean, the reasons are usually elsewhere, and we have written up the five that come up most.

Worth passing on?

Our articles are drafted with the help of AI tools that we regularly use. Each one is measured, edited and approved by real people who stand by it.

Before 15 September

We will tell you what the engines can actually see.

Send us your domain and we will check your robots.txt, your crawler access and what the engines currently say about you, then tell you which of the three is the problem. No obligation.

Get your site checked

No obligation. Mon to Fri, 10am to 5.30pm