You can now refuse AI training and stay in Google
For two years, an estate agency that did not want its website used to train AI models had to accept a cost somewhere in search. On 15 September Cloudflare separated the two, published the conditions an AI company has to meet to be let through, and named the first ones that meet them. One switch still takes Google away, and it is the one people reach for.
Our piece on whether your website is blocking the AI engines carried a warning about a date. The date was 15 September, and the thing to worry about was a tick box labelled "Block AI bots" that looked like it would take Googlebot with it.
That date has been and gone. What actually happened is more interesting than the warning was.
The choice that used to be impossible
For two years the AI question put site owners in a corner. You could let the AI companies read your website, which meant letting them train on it. Or you could block them, which meant losing search traffic that arrived through the same crawlers. Google, Apple and Microsoft all run crawlers that do more than one job, so a rule aimed at training landed on search as well. There was no way to say yes to one and no to the other, and every honest write-up of the subject, ours included, ended with a version of "pick your loss".
On 15 September Cloudflare split the decision into three. Search, Training and Agent are now separate controls and they act independently. There is a new Training option called Disallow AI Training, which publishes your preference into robots.txt and leaves crawlers that do both jobs free to carry on indexing you for search.
So the sentence at the top of this article is now true, and it was not true at any point before 15 September. If your site sits behind Cloudflare, you can refuse to be training material and stay exactly where you are in Google.
What decides who gets let through
The part worth reading properly is the test, because it is published and it is specific. Cloudflare has set out four conditions an operator has to meet to be treated as what it calls Accountable, and a crawler that does both search and training keeps its search access only if its operator meets all four.
The operator has to give site owners a clear way to opt out of AI training through robots.txt or a comparable standard. It has to let site owners opt out of AI generated search summaries. It has to provide URL level visibility into how a site's content is being used for both search and training. And it has to publicly confirm that opting out of training will not affect a site's ranking in ordinary search.
Three operators are named as having met that test with crawlers that do more than one job: Apple with Applebot, Google with Googlebot, and Microsoft with Bingbot. Cloudflare also treats as Accountable a set of companies that run separate crawlers for search and training in the first place, naming Amazon, Anthropic, Meta and OpenAI, on the basis that splitting the crawler achieves what splitting the setting achieves.
That closes a question we left open in August. When Cloudflare first published the four conditions it named nobody against them, and we said in print that we did not know whether Googlebot would qualify. It does.
What happened to everyone who had already ticked the box
This is the part that has changed our advice rather than confirmed it.
The warning was that a legacy "Block AI bots" selection, carried forward into the new controls, would from 15 September catch Googlebot along with everything else. That was not speculation. It was Cloudflare's own published position in July, naming Googlebot, Applebot and BingBot as the crawlers it would apply to for anyone who had selected to block training.
What Cloudflare did on the day was gentler. Existing settings were migrated rather than enforced. In its words, previous Training selections of Block, or Block on pages with ads, migrate to Disallow AI Training. Asked what customers need to do about any of it, its answer is "Nothing, in almost every case. Your current settings carry over on their own." An agency that ticked that box two years ago and never thought about it again has, in most cases, ended up with the sensible setting by accident.
That does not make the check pointless. It makes it quicker.
The switch that still costs you Google
Block has not gone anywhere. It is still an option on the Training control and it still does what it always did. Cloudflare's own description of choosing it is that it will stop Applebot, Bingbot and Googlebot reaching your site, and it says plainly that this includes search.
So the trap is still in the product. What has changed is that it is no longer the outcome of an old decision nobody remembers making. It is a live choice, sitting one line away from the option almost everybody actually wants. Disallow AI Training and Block are adjacent on the same control, they sound similar to anybody who is not steeped in this, and only one of them leaves you in Google.
If somebody at your agency opens that setting having read a headline about AI companies helping themselves to content, there is a fair chance they pick the stronger sounding one. That is the sentence worth passing to whoever looks after your website.
New sites arrive with different defaults
Domains joining Cloudflare from now on get one of two presets. A site that carries no advertising arrives with Search, Training and Agent all set to Allow. A site that does carry advertising arrives with Search on Allow, Training on Disallow AI Training, and Agent set to block on pages that show ads.
Almost every estate agency website falls into the first group, which means a newly built site arrives with nothing blocked at all. If refusing training is a position your agency wants to take, it is now something you go and choose rather than something you inherit from a setup screen.
One honest limit
The second of those four conditions, opting out of AI generated summaries, is currently something a site owner arranges with each operator directly. Cloudflare has said it is aiming to bring that into its own dashboard in early 2027. So having it both ways is real on the training question today, and still a job of work on the summaries question. Anybody telling you the whole thing is one click has not read past the headline.
What we would actually do
Three things, none of which take long.
First, find out whether your site is behind Cloudflare at all. A good number of agents genuinely do not know either way, and if it is not, none of this applies to you and your robots.txt is the whole story.
Second, look at the Training control and check it says Disallow AI Training rather than Block. That is a ten second job for whoever holds the login, and it is the only thing on this page that can cost you Google.
Third, read your own robots.txt, at yourdomain.co.uk/robots.txt. Cloudflare now writes AI crawler lines into that file from your dashboard settings, which means a line somebody edited by hand can reappear, and a line you are reading may have been generated rather than chosen.
The longer version of that check, including the crawlers that turn up when a real person has asked a real question about your agency, is in our five minute walk through, updated to say what Cloudflare did rather than what it said it would do.
A closing note on what this is, because it is easy to read it as bigger than it is. It is a setting at one company. It is not a law and it is not a standard. It covers the share of the web that sits behind Cloudflare, which is large but nothing like all of it, and it binds the operators that have agreed to be bound. It is the best answer anybody has offered yet to a question site owners were being asked to decide with no good options. It is not the last word on it.
Our articles are drafted with the help of AI tools that we regularly use. Each one is measured, edited and approved by real people who stand by it.
You can now opt out of Google's AI answers. Should you?
The setting went live for every website in the world on 31 August, and Britain got it first because of a competition regulator. What it does, the three things it does not do, and why we would leave it alone.
The AI crawler that asked for a private key
Everybody counts AI crawler hits. We asked a duller question on a client site: which pages were they actually taking? The answer was not pages at all, and it led to a check almost nobody runs, and to a number we can now publish.
Why doesn't ChatGPT mention my agency?
You asked ChatGPT who the best agent in your town is and it named your rivals. The five reasons that happens, drawn from live answers we capture and measure, and what actually fixes each one.
We will tell you what your site is actually set to.
Send us your domain and we will check your crawler settings, your robots.txt and what the engines currently say about you, then tell you plainly which of the three is the problem.
Get your site checkedNo obligation. Mon to Fri, 10am to 5.30pm

