15 September 2026: Cloudflare Changes the Crawler Defaults
Search, Agent, Training — Cloudflare re-sorts bots and blocks two categories by default on ad-bearing pages. Why that can take Googlebot down with it.
by Jean Pierre Kolb ·
My own site does not sit behind Cloudflare — it runs on its own Apache, and my Content-Signal line is exactly what it claims to be: a request. Which is why I look at what Cloudflare announced on 1 July 2026 with a certain detachment. Because there the request turns into a block, and it lands on 15 September 2026 — for more than 20 percent of all web domains, which sit behind that network. If yours is one of them, the next few weeks decide which crawlers still get to see you. And one detail in the evaluation logic can cost you Googlebot of all things.
What changes on 15 September
Cloudflare is switching the default settings for bot access. The new rule in one sentence: on pages that display ads, crawlers in the Training and Agent categories will be blocked by default — Search stays allowed.
The reasoning behind it is economic and fairly straightforward: an ad on a page is the signal that a human was meant to land there and see something. On those pages Cloudflare treats human attention as the actual goal and keeps away the bots that consume it without giving anything back. Search, by contrast, traditionally funnels visitors back — so it stays open.
Who it hits:
- All domains newly onboarded to Cloudflare get the new defaults automatically.
- Existing free customers who have not touched their settings by 15 September do too — that is stated in the press release.
- Anyone who does not want the switch can record that any time up to 15 September in their security settings.
One honest note on status: Cloudflare has said it will finalize the defaults and classifications ahead of the deadline, using ecosystem feedback and its own tests. The direction is set; individual details may still move.
The new taxonomy: Search, Agent, Training
The more interesting part is the sorting underneath. Until now the debate boiled down to "AI bot, yes or no" — a question that gets less meaningful every year, because by now there is AI in practically everything. So Cloudflare no longer asks whether a bot is "AI", but what it does on your site. Three categories are now available to all customers, free tier included:
| Category | What the bot does |
|---|---|
| Search | Proactively builds an index of your content to answer questions about it later — with referral traffic or other equitable compensation as the expectation |
| Agent | Acts in real time on behalf of a human: chat fetch bots such as ChatGPT-User, but also browser agents driving Chrome |
| Training | Takes your content to train or fine-tune a model — the content is permanently absorbed into the model architecture |
Beneath that sits a broader classification with eleven behaviors: alongside the three above there are Transact (checkout actions on behalf of users), Data Collection, Security Testing, SEO, Ads Verification, Social / Link Preview, Feed Fetching and Monitoring & Operations. A bot is explicitly tagged with all applicable categories, not just one — and that is exactly where the problem in the next section comes from.
This also changes what "Verified" means. Until now a verified bot was allowed by default. From now on the label only makes a bot allowable — what lets it through is the category you have opened up. Non-verified bots stay blocked by default as before.
The Googlebot trap
This is the point I consider the most dangerous in practice. From 15 September, multi-purpose crawlers are evaluated by all of their behaviors, and the most restrictive applicable rule wins.
Googlebot, Applebot and BingBot crawl for search and for training. So anyone blocking training — through the new options or through the old "Block AI bots" switch — shuts those three out entirely. Classic search indexing included.
That is not a side effect but the intent: Cloudflare wants bot operators to separate their crawlers by purpose instead of bundling search and training into the same user agent. The pressure for that is deliberately routed through site owners and thus onto the crawler operators. For you it still means one thing: a switch you may have flipped two years ago for good reasons can cost you Google visibility from mid-September.
And unlike a robots.txt line, this is not an appeal. The block takes effect at the network level, before the request reaches your server — a crawler cannot ignore it, because it never gets through in the first place.
What to check before 15 September
If your site sits behind Cloudflare, these are the four moves I consider worthwhile — in this order:
- Check whether "Block AI bots" is active for you. That is the legacy switch which translates into the new Training category. If it is on, the Googlebot question concerns you directly.
- Make a deliberate call on Agent. This category is new and easily overlooked. An agent acts for a real human who is waiting for an answer right now — that is closer to a visitor than to a scraper. Block agents wholesale and you cut yourself off from a channel that is only just emerging.
- Check whether your pages even count as "with ads". The new defaults hinge on exactly that. On an ad-free company or documentation site, less changes than the headlines suggest.
- Keep your
robots.txtand your Content Signals consistent. The preference layer stays in place, and a contradiction between "AI welcome" in theContent-Signalline and a hard block at the CDN is precisely the kind of inconsistency that helps nobody. The SEO & GEO Analyzer flags that conflict as an informational note.
If your site does not sit behind Cloudflare, nothing changes for you immediately — robots.txt and Content-Signal remain a declaration of intent. The details are in Content Signals & C2PA, including the new use field Cloudflare introduced into the directive on the same day.
The economic background
The switch does not come out of nowhere. Cloudflare's own anniversary numbers paint a clear picture of why the old trade — we crawl you, you get visitors — no longer holds:
- More than 50 percent of internet traffic is now non-human.
- 52 percent of all crawler requests went to training as of June 2026 — in spring 2025 it was 22 percent.
- More than 36 percent of activity comes from multi-purpose crawlers blending search, agent use and training.
- The ratio of crawls to visitors sent back ranged from 118:1 to nearly 50,000:1 for major AI crawlers.
- A 2025 Pew study found that when Google shows an AI summary, users click a traditional result in only 8 percent of visits — without a summary it is 15 percent, almost twice as often. They click a link inside the summary in 1 percent of visits.
Cloudflare's answer amounts to more than blocking. The existing Pay Per Crawl is being reshaped into Pay Per Use: what gets paid for should no longer be the fetch but the actual value — a page may be crawled once and then cited thousands of times, or fetched a hundred times and never used. First partners are Ceramic.ai with a pay-per-query model and You.com, where an agent pays on demand for a single premium document. On top of that come a new Attribution Business Insights dashboard for bot management customers and a research program on crawl efficiency — according to Cloudflare, more than 50 percent of good bots' crawl traffic goes to pages that have not changed at all.
Whether a working market comes out of this is open; Cloudflare calls it an experiment itself. For practical purposes, the deadline is what counts first.
FAQ
Does this affect me if I am not on Cloudflare?
Not directly. The new defaults are a configuration in the Cloudflare network and apply only to domains that sit there. Indirectly it still concerns you, because crawler behavior as a whole shifts: if a meaningful share of the web blocks training and agents, operators will change their crawlers — and the new, separated bots will show up in your logs too.
Will I lose Google rankings if I block training?
Possibly, and that is precisely the risk. Because from 15 September the most restrictive rule applies and Googlebot crawls for both search and training, a training block shuts it out completely. No access, no indexing; no indexing, no ranking. If search visibility matters to you, blocking training is no longer a risk-free default.
What is the difference between Agent and Search?
Timing and purpose. A search crawler builds an index in advance so it can answer questions later. An agent shows up because right now a human wants something from your page — it reads a page, fills in a form or fetches a piece of information somebody is waiting for. Economically those are two entirely different things: one promises future traffic, the other already is the visit, just without a browser.
Do I need to change my robots.txt now?
Not strictly. The robots.txt remains the preference layer and is untouched by the Cloudflare switch. It is still worth aligning both layers: what you block at the CDN should not be explicitly invited in your Content-Signal line. If you want to use the occasion, add the new use field while you are there.
Further reading
The preference layer with all signals, the new use field and the path-specific rules is covered in Content Signals & C2PA. Technical crawler management and the rest of the GEO foundation are in Structured Data & Technical GEO; the frame comes from the pillar What is GEO?. How to measure any of this is in Measuring GEO — and the Outlook on Agents. Check your own crawler rules with the SEO & GEO Analyzer.
Cloudflare's original sources: the post "Your site, your rules" on the taxonomy and defaults, and the agentic internet bot report for the numbers.