Since 1 July 2026, Cloudflare has managed AI crawlers by three purposes – Search, Agent and Training – for all customers, including the free tier. That part is documented on Cloudflare’s official blog. But a quieter second part of the announcement is the one that matters for SEO.
- From 15 September 2026, multi-purpose crawlers are evaluated by their most restrictive applicable behaviour. If you block Training – even through the old “Block AI bots” toggle – Cloudflare says you can also end up blocking Googlebot, Applebot and Bingbot.
- The new default blocks on ad pages, by contrast, only apply to domains newly onboarding to Cloudflare. Existing zones are not covered – a distinction that news coverage keeps blurring.
- Your move: before the deadline, check your Security settings for active Training blocking and decide deliberately – set the opt-out, differentiate, or accept the new logic.
In its own announcement from 1 July 2026, Cloudflare names three bots that its own customers can lock out from 15 September onward: Googlebot, Applebot and Bingbot. Not attackers. Not competitors. The site owners themselves – through a toggle many of them flipped months ago and long since forgot.
The background: Cloudflare has fundamentally rebuilt how AI crawlers are controlled. Instead of the old blanket “Block AI bots” option, there are now three categories – Search, Agent and Training. And from 15 September 2026, per Cloudflare, multi-purpose crawlers that collect both for search and for AI training are subject to the most restrictive applicable rule. Cloudflare itself names Googlebot as an example of such a multi-purpose crawler.
Every time a new AI bot shows up on the web, a client sits down with me shortly after and asks whether we should wall it off – since 2024 that question has been part of every second technical audit I run at SEO Kreativ. The answer used to be fairly low-risk: the AI block hit GPTBot, ClaudeBot and the like, while Google Search stayed untouched. That clean separation is exactly what Cloudflare is now partly removing. If you flipped a blanket “block AI” switch on a Cloudflare customer, you should know before 15 September what that switch means from then on.
In this article I break the announcement down against the primary source: what exactly is new, who is affected (and who is not – a point early coverage easily blurs), how you check your setup in ten minutes, and what the new use= parameter in the Content Signals means for your robots.txt.
What Cloudflare changed on 1 July 2026
The announcement ran under the “Content Independence Day” label, Cloudflare’s annual date for crawler policy, one year after the introduction of the one-click “Block AI Bots” switch and the Pay-per-Crawl programme. The core message in the official blog post (Jin-Hee Lee and Bryan Becker, 01/07/2026): not every AI use is the same, so it needs differentiated controls.
The three categories, as Cloudflare defines them:
| Category | Behaviour per Cloudflare | Examples |
|---|---|---|
| Search | Collects and indexes content to answer questions about it later – classic search and AI search. Cloudflare expects referral traffic in return here. | Search engine crawlers, AI search indexers |
| Agent | Acts in real time on behalf of a user – fetches pages, fills in forms, gets things done. Often a human is waiting at the other end. | ChatGPT fetch, browser agents such as Gemini or Claude in the browser |
| Training | Takes content to train or fine-tune a model. The data is permanently absorbed into the model architecture. | Training crawlers of the model providers |
Important for context: this is a behaviour taxonomy, not a company taxonomy. A bot can fall into several categories at once – and that detail becomes relevant in September. Cloudflare classifies further behaviours as well, but the three AI categories are the ones controllable by all customers for now.
In parallel, Cloudflare has redefined what “Verified Bots” means: being verified now only means a bot’s identity is confirmed – not that it is automatically let through. Whether a verified bot gets access is decided by its category plus your settings.
15 September 2026: new defaults are not the same as a new rule for everyone – the Googlebot trap
This is where a close look at the primary source pays off, because the two changes have completely different reach:
Change 1 – the new defaults (new domains only): For all domains onboarding to Cloudflare from 15 September 2026, the Training and Agent categories are blocked by default on pages with ads, while Search stays allowed. Cloudflare’s reasoning: ads signal that a human is meant to see the page – bots that replace that human attention are kept out there. Existing zones with unchanged settings are not covered by this defaults part according to the blog post; the wording is explicitly “for all new domains onboarding to Cloudflare”.
Change 2 – the multi-purpose rule (everyone who blocks Training): From the same deadline, crawlers that combine Search and Training are evaluated by all of their behaviours – and the most restrictive applicable rule wins. Cloudflare writes verbatim that multi-purpose crawlers such as Googlebot, Applebot and Bingbot will be blocked by customers who have selected to block Training – whether through the new category options or through the legacy “Block AI bots” switch.
That is the real news for SEOs. The old one-click switch used to be the convenient, low-risk recommendation: AI training bots out, search untouched. From 15 September the same switch changes its effect – without you touching it. “Blocks GPTBot and co.” potentially becomes “also blocks Googlebot on the configured pages”. Search Engine Journal highlights exactly this point as the central SEO consequence. Just how granularly Cloudflare classifies Googlebot in the end, and how hard the block actually bites – for instance only on certain pages or signals – remains to be seen in practice from the deadline onward.
One open flank remains – and here I separate cleanly by evidence status: it is documented that Cloudflare urges bot operators to separate their crawlers by purpose. It is speculation whether Google will do so. Per its own documentation, Google has always crawled for search and AI features primarily with Googlebot; the separate Google-Extended token, per Google’s docs, controls only Gemini training, not search. Whether Google splits its crawler architecture because of Cloudflare is open – back with the Content Signals in September 2025, Google had, according to Search Engine Land, not committed to respecting the signals.
Who is affected – and who is not?
The matrix I am currently working through for my own managed Cloudflare zones:
| Setup | Affected from 15/09/2026? | What to do |
|---|---|---|
| Cloudflare + “Block AI bots” (legacy) active | Yes – multi-purpose crawlers incl. Googlebot can, per Cloudflare, fall under the block | Decide before the deadline: set the opt-out or switch to the new categories |
| Cloudflare + new “Training” category set to block | Yes – same multi-purpose logic | Configure deliberately; optionally block Training on ad pages only |
| Cloudflare, no AI blocking active | No (existing zone, defaults unchanged) | Nothing required; a chance to set the new options deliberately |
| New domain, onboarding after 15/09 | Yes – new defaults apply (Training + Agent blocked on ad pages) | Check and adjust defaults after onboarding |
| No Cloudflare | No | Only the robots.txt layer is relevant (see Content Signals below) |
One pattern becomes dangerous here that I know from my work as a Product Developer at iGaming.com just as well as from small WordPress setups: security settings get configured once during setup – often by IT, not by the SEO team – and are never touched again. The “Block AI bots” switch was a no-brainer in many setups in 2025 because it felt risk-free. Those exact setups are now the candidates for unintended Googlebot blocks. If SEO and security responsibility are split in your company: this deadline is the occasion to bring both to the same table.
On scale: Cloudflare states it sits in front of more than 20 percent of all web domains (per Cloudflare’s own figures, not independently verified). Even if only a fraction of those have Training blocking active, we are talking about a meaningful set of domains whose Google crawling can change on 15 September.
Check in 10 minutes: does your setup lock out Google?
Step 1 – dashboard check. Per the Cloudflare docs, you find the setting in the Application Security dashboard under Security settings, filter “Bot traffic”, entry “Block AI bots” (per current Cloudflare docs; the exact labels may change in the dashboard). There you see one of several configurations, including “Only block on hostnames with ads”, “Block on all pages” or the off state. If it shows anything other than “off”, you are subject to the multi-purpose rule from 15 September. Also check the new category presets for Search, Agent and Training, available in the same settings since 1 July.
Step 2 – make a decision. You have three sensible options, depending on your content strategy:
- Option A – set the opt-out: You want to keep blocking training bots but leave Googlebot and co. untouched. Then set the opt-out in the Security settings before 15 September. This confirms, per Cloudflare, that nothing changes for training crawlers with a search function.
- Option B – deliberately accept the new logic: If your content monetises primarily through other channels and you want to strengthen your negotiating position against Google, you can let the multi-purpose rule take effect – fully aware that Google crawling of the configured pages is then blocked. For most publisher and e-commerce setups I look after, that is currently not a realistic option.
- Option C – differentiate: Block Training on ad-funded pages only (“Only block on hostnames with ads”) and watch how crawling develops.
Step 3 – cross-check robots.txt. If you use Cloudflare’s Managed robots.txt, Cloudflare injects directives before your own content. Fetch your robots.txt fresh once (with a cache buster) and check what is actually served. How the directive logic works in general and which mistakes to avoid, I have written up in detail in my robots.txt guide.
Step 4 – measure after the deadline. From mid-September, watch the crawl stats in Search Console (Settings, Crawl stats) and your server logs for Googlebot activity. A sudden drop in crawl requests with unchanged content is the warning sign. For the AI visibility side, it is worth looking in parallel at the new GenAI performance reports in Search Console, if your property already has the rollout.
Content Signals: the new use= parameter in robots.txt
use describes what a crawler may keep and reuse after access – in three levels: immediate, reference, full. Important: this is a preference statement, not a technical block.
Alongside the WAF layer (actual blocking), Cloudflare is expanding the signal layer. The Content Signals, introduced in September 2025, previously had three fields: search, ai-input and ai-train. New is the use field with three levels, from restrictive to permissive:
use=immediate– interact, but store and reuse nothing.use=reference– index, show excerpts, link back (the default).use=full– summarising and reproducing allowed.
Customers with the Managed robots.txt enabled get the new preference added automatically, per Cloudflare. The concrete output can vary depending on configuration; per current Cloudflare documentation the served directive looks roughly like this:
# Cloudflare Managed content with the new content-use signal
User-agent: *
Content-Signal: search=yes,ai-train=no,use=reference
Allow: /
Two points of context. First: Content Signals are preferences, not enforcement – crawlers can ignore them. Cloudflare therefore ties them to an incentive system: behaviour is tracked, and verified bots that disregard declared usage rules or reproduce content in full are meant to lose their verified status – and with it access to a substantial part of the Cloudflare network. Whether that lever disciplines anyone in practice remains to be seen.
Second, a practical detail: Google Search Console can occasionally report “syntax not understood” for Content Signals and newer directives. Per Cloudflare, no impact on crawl rates or SEO has been observed so far – but that is a vendor’s own statement, not an independent measurement. If you meet the notice in GSC: assess first, then act.
Where this fits strategically – between robots.txt, llms.txt and the AI visibility disciplines – I covered in detail in the pillar on SEO, AIO, GEO and LLMO. Short version: the signal layer is growing, but the enforcement layer remains WAF and bot management.
Strategic context: why Cloudflare is doing this
Cloudflare’s argument rests on its own network data: per Cloudflare’s July 2025 analysis, the ratio of crawls to returned visitors was roughly 14:1 for Google, about 1,700:1 for OpenAI and around 73,000:1 for Anthropic (Cloudflare’s own measurement, July 2025). At those magnitudes, the old “crawling for traffic” deal clearly no longer holds for AI answers. The multi-purpose rule is Cloudflare’s attempt to renegotiate that deal: any bot operator who wants continued access should be transparent about what it crawls for – ideally with separate crawlers per purpose.
The calculus behind it is one Cloudflare states openly: bundling search and training in the same bot gives established search providers an advantage, because site owners cannot sacrifice search in order to refuse training. That exact bind is what the multi-purpose rule is meant to break. Whether Google plays along is the open question – I think a short-term crawler split is unlikely (speculation), because Google would be giving up a central piece of leverage. More realistic, to me, is a longer stalemate in which individual publisher groups test the block and Google sits out the effects at first.
My take: “blocking AI” is still treated far too often as a purely defensive security decision. From now on that falls short – it is a visibility and monetisation decision. Whoever blocks forgoes AI citations and potentially, from the deadline, Google crawling of the affected pages. Whoever opens up supplies training data without direct compensation, as long as systems like Cloudflare’s Pay-per-Crawl approach are not broadly in effect. There is no default answer here, only a deliberate trade-off per business model.
Frequently asked questions (FAQ)
Does Cloudflare block Googlebot automatically now?
No. Cloudflare does not block Googlebot on its own. From 15 September 2026, the multi-purpose rule only applies to zones whose owners have activated Training blocking – through the new category options or the old “Block AI bots” switch. Anyone without active AI blocking is not affected by this rule.
I activated “Block AI bots” months ago – what happens on 15 September?
From the deadline, multi-purpose crawlers that combine Search and Training – Cloudflare names Googlebot, Applebot and Bingbot – can fall under your existing block, to the extent of your configuration (all pages or only pages with ads). If you do not want that, set the opt-out in the Security settings before 15 September.
Do the new defaults also apply to my existing Cloudflare zone?
The new ad-page defaults (Training and Agent blocked) apply, per Cloudflare’s blog post, to domains onboarding to Cloudflare. For existing zones the defaults do not change – what is relevant for you is the multi-purpose rule, if you have Training blocking active.
Is robots.txt enough to stop AI crawlers?
No. robots.txt and Content Signals are preference statements – cooperative crawlers respect them, others ignore them. Actual blocking at Cloudflare happens at the WAF layer. The sensible move is the combination: signals for the legal position and cooperative bots, WAF rules for enforcement.
What does the new use= parameter in robots.txt mean?
The use field extends Cloudflare’s Content Signals with a usage preference after access: immediate (store nothing), reference (index, cite, link – the default in the Managed robots.txt) and full (summarising and reproducing allowed). It is a signal, not a technical block – but Cloudflare ties compliance to the verified status of bots.
Can I block Training and still stay visible in ChatGPT and co.?
In principle yes – that is exactly what the category separation is for: block Training, allow Search (and Agent depending on strategy). The catch is the multi-purpose rule: bots that bundle Training and Search in the same crawler still fall under the block. With providers that have cleanly separated crawlers, the differentiation works as intended.
Conclusion: decide deliberately instead of accepting the default
The direction is right: being able to control Search, Agent and Training separately is exactly the granularity my audits have been missing for two years. That Cloudflare rolls the controls down to the free tier makes them the new standard tool – even for small sites that previously only had robots.txt. At the same time, 15 September is a deadline you should not leave to chance. The change to the multi-purpose logic happens server-side at Cloudflare; your switch stays the same, its effect does not.
My approach for the coming weeks: document a baseline, run a dashboard check on every managed Cloudflare zone, record the decision per project – and from mid-September watch the crawl stats closely. Should it turn out that Google reacts to the purpose separation, that would be one of the larger shifts in the crawler ecosystem in years. Until then: signal preferences, configure enforcement, measure the effect.
As of July 2026. Despite careful research, all information is provided without guarantee as to timeliness and completeness. This content is for general informational orientation only and does not constitute individual legal or professional advice.


