A new toggle showed up in a lot of Cloudflare dashboards on September 15: “Disallow AI Training.” No explanation banner, no walkthrough — just a setting sitting there, waiting to be clicked or ignored.
So people did what people do. They clicked it, then wondered what they’d just done. Would it stop AI companies from training on their content? Sure, probably. Would it also quietly knock them out of Google search results? That part was less clear, and a few early blog posts made it sound scarier than it actually is.
Here’s the real answer: this setting lets you say no to AI training crawlers specifically, while your normal Google and Apple search crawling keeps working exactly as it did the day before. That’s a genuinely new option, not a rebrand of an old one — confirmed directly in Cloudflare’s own announcement and independently reported by outlets like Search Engine Journal. There’s still one gap worth knowing about, though, and we’ll get to it.
The Choice This Replaces
Before this update, blocking a crawler was an all-or-nothing move. Turn it off, and you turned off everything that crawler did — no exceptions.
That was a problem, because a lot of AI companies run two crawlers under one name, doing two different jobs. One crawls your site to power search results. The other scrapes the same site to train an AI model. Historically, they traveled together. Block the crawler to stop the training use, and you blocked the search-indexing one right along with it — taking your search visibility down with it, whether you meant to or not.
Plenty of businesses were fine having Google or Apple index their pages. Far fewer were comfortable having years of blog posts, product pages, and original research quietly folded into someone else’s model with no say in the matter. But until now, you couldn’t pick one without risking the other.
What’s Actually Happening Under the Hood
Every website can carry a small text file called robots.txt. It’s less a lock than a sign at the front gate — a set of house rules crawlers are expected to read and respect. Most reputable ones do.
Cloudflare’s new setting works by updating that sign automatically. Turn on “Disallow AI Training,” and Cloudflare adds a Disallow instruction to your robots.txt aimed at AI-training use specifically. Cloudflare groups Amazon, Anthropic, Meta, and OpenAI together as companies that run entirely separate search and training crawlers, so blocking the training one never touches the search one. Google, Apple, and Microsoft work a little differently, and it’s worth getting this part right: their main crawler does the actual crawling, full stop. Google-Extended and Applebot-Extended aren’t second crawlers making their own visits to your site — they’re permission tokens (think of a permission token as a separate checkbox for one specific use, not a second visitor showing up at your door) that control whether the content Googlebot or Applebot already collected can be reused for AI training. Same crawl, different rulebook for what happens to the data afterward.
That’s the split that makes this whole setting possible. Disallow the training token, and the crawl itself doesn’t change at all — Googlebot and Applebot keep visiting and indexing exactly as before. What changes is downstream: that company’s AI-training pipeline is no longer allowed (per the signal, if it’s respected) to pull from what got crawled.
This setting replaces the older “Block” option, which used to shut a crawler down completely, search role included. If you’d already picked “Block” for one of these crawlers before September 15, Cloudflare migrated you over to “Disallow AI Training” automatically. According to Cloudflare, most customers don’t need to touch anything.
There’s also a default worth knowing about if you’re setting up a new domain on Cloudflare rather than managing an existing one. New sites that monetize through advertising default to “Disallow AI Training” out of the gate. Sites that don’t monetize that way default to “Allow” instead. Either way, it’s a starting point, not a fixed rule — you can change it in your dashboard whenever you want.
Worth saying plainly, though: a robots.txt line is a preference, not a padlock. It works because major crawlers choose to honor it. It isn’t a legal guarantee, and it can’t physically stop a crawler that decides to ignore it. Think of it as a strong, industry-standard signal — not an ironclad barrier.
So Will This Actually Hurt My Google Rankings?
No. And this is the part causing most of the confusion online right now.
Googlebot does the crawling and indexing your rankings actually depend on. Google-Extended isn’t a second crawler — it’s a separate permission token that governs whether the content Googlebot already collected can be reused to train AI models. Cloudflare’s new setting targets that token only. Googlebot’s own crawling behavior never changes.
Google has stated directly that Google-Extended doesn’t affect a site’s inclusion in Search or its ranking. Disallow it, and you opt out of the training reuse without touching indexing at all.
Same story with Apple. Applebot-Extended governs AI-training reuse; the regular Applebot does the actual crawling for search and, by extension, things like Siri and Spotlight suggestions. Cloudflare’s language here is direct too: Applebot-Extended “doesn’t crawl pages and isn’t considered in search ranking” — confirming it’s a permission signal, not a bot making its own requests.
One thing this setting does not touch, and it’s worth being precise about: whether your content shows up in Google’s AI Overviews or AI Mode. That’s a separate visibility layer, controlled through its own setting inside Google Search Console — nothing to do with this Cloudflare toggle. Don’t assume flipping this switch changes your AI Overview presence one way or the other. It won’t.
Bing Is the Exception — And You Should Plan Around It
Here’s the caveat every write-up on this topic should lead with, but a few skip.
Microsoft hasn’t added matching robots.txt-level support yet. Turn on “Disallow AI Training” through Cloudflare today, and that preference simply doesn’t reach Bing’s crawlers the way it reaches Google’s or Apple’s.
Bing still relies on an older, page-level tool instead: the NOARCHIVE meta tag. Different mechanism, different scope — it has to be applied page by page rather than flipped once at the domain level.
Microsoft has reportedly said early 2027 is the target for full robots.txt support. Fair enough — but treat that as a stated intention, not a locked date. Timelines like this slip more often than they land on schedule.
Worth adding: this same gap applies more broadly than just Bing. Cloudflare’s own transparency notes mention Google is still rolling out URL-level detail for how Google-Extended opt-outs get applied, and Apple’s version of that same tooling is reportedly further out, into 2027. None of that changes whether your search rankings are safe today — it just means the newer, more granular reporting layer everyone eventually wants is still being built, crawler by crawler.
Practically, that leaves three quick questions worth asking before you decide anything:
- Do you publish a lot of original content you’d rather not see quietly absorbed into someone’s training data? If so, turning this on costs you nothing on the search side, so there’s little reason not to.
- Does Bing traffic actually move the needle for your business? If yes, this Cloudflare setting alone won’t cover you — you’ll need the NOARCHIVE tag too, at least until Microsoft ships its own version.
- Are you looking for a guarantee, or a reasonable industry-standard signal? If your goal is closer to airtight legal protection, this setting isn’t that, and that’s a different conversation entirely — one for your legal or compliance team, not your CMS settings panel.
For most businesses publishing content and caring about both search visibility and some say over AI training, flipping this on is close to a no-downside decision on the Google/Apple side. It’s just not a complete solution on the Bing side yet.
Checking the Setting on Your Own Site
Quick version, since dashboard menus shift over time and screenshots go stale fast:
Log into Cloudflare. Find AI Crawl Control or Bot Management — Cloudflare has been folding crawler-related settings into these two areas. Check what’s selected per named crawler: Allow, Disallow AI Training, or Block. Cross-reference against Cloudflare’s own current documentation for the exact click path, since that’s the source that’ll stay accurate longer than any guide written today, including this one.
Frequently Asked Questions
Will turning on “Disallow AI Training” hurt my Google or Apple search rankings?
No. Google-Extended and Applebot-Extended are permission tokens, not separate crawlers — they control whether already-crawled content can be reused for AI training, while Googlebot and Applebot keep crawling and indexing for search exactly as before. This setting only targets the training-reuse token, and both companies state that disallowing it doesn’t affect search inclusion or ranking.
Does this setting stop my content from appearing in Google’s AI Overviews?
No. AI Overviews and AI Mode visibility run through a separate Google Search Console setting, unrelated to this Cloudflare toggle.
Does Bing respect this setting yet?
Not at the robots.txt level, no. Microsoft hasn’t added matching support as of this writing, so this setting doesn’t currently reach Bing’s crawlers the way it does Google’s or Apple’s. Bing’s current opt-out route is the separate NOARCHIVE meta tag. Microsoft has reportedly said early 2027 for full support, though that isn’t a confirmed date.
Is “Disallow AI Training” the same as blocking a crawler completely?
No. The older “Block” option stops a crawler entirely, search role included. “Disallow AI Training” is narrower — it only targets AI-training reuse, either by blocking a fully separate training crawler (Amazon, Anthropic, Meta, OpenAI) or by disallowing the training-specific permission token (Google, Apple), depending on how that company splits the two.
Do I need to do anything if I already had crawlers set to “Block” on Cloudflare?
Probably not. Cloudflare says existing “Block” selections were migrated to “Disallow AI Training” automatically, and most customers won’t need to change anything manually. Still worth a quick check on your own dashboard if you want certainty for your specific domain.
One Setting, Two Separate Decisions
Blocking AI training used to mean risking your search visibility right along with it. That’s the tradeoff Cloudflare’s new setting breaks — at least for Google and Apple, both of which already separate the training-reuse decision from the crawling itself. Flip it on, and you opt out of training without touching search.
Bing is the piece still catching up. Worth watching, not worth panicking over, especially if you can lean on NOARCHIVE in the meantime. And it’s worth remembering, always, that a robots.txt signal works because good actors choose to respect it — not because it’s enforced by anything stronger than reputation.
If you publish content you’d rather not see feeding someone else’s model for free, this is a low-risk setting to turn on. Small move. But it’s a decent sign of where things are heading: search visibility and AI training are slowly becoming separate decisions, instead of one switch that controlled both.