A year after giving publishers a one click block AI bots switch, Cloudflare rewrites bot rules around behavior, with a Sept 15 default that makes blocking training cost search visibility.
Cloudflare's network carries a large slice of internet traffic. The company just rewrote the question website owners ask about bots, from "is it AI?" to "what is it doing, storing, and resharing?" In doing so, it redrew the boundary between blocking AI training and being findable on the open web at all.
The update lands a year after Cloudflare's first "Content Independence Day," which gave publishers a one-click switch to block AI bots and laid the groundwork for a Pay-Per-Crawl marketplace. The new rules swap the old identity-based taxonomy, which asked whether a bot counted as "AI," for a behavioral one: what the bot is doing on the site, what it is storing, and how it plans to reshare the content.
"Instead of defining a bot primarily as 'AI' or not, our updated approach to classification will ask deeper questions about bot or agent behavior," Cloudflare writes in the announcement. "What are they doing on my site? What are they storing? And how will they reshare my content?"
The practical bite comes on September 15. After that date, multi-purpose crawlers like Googlebot, Applebot, and BingBot will be governed by all of their behaviors, with defaults enforced by the most restrictive applicable rule. A site that blocks "Training" will, by default, also block those crawlers' search and indexing work, because the same bots do both jobs. The lever the publisher thought controlled one thing now cuts off three at once.
That makes the trade-off Cloudflare's own post calls "Faustian" structurally sharper than before. Smaller publishers can still let crawlers in to stay visible in search results, but doing so means accepting that the same visit may also feed model training. The third option, Pay-Per-Crawl, lets a publisher set a price for crawlers that want to use the content for AI; Cloudflare sits between buyer and site, sets the rate, and handles the handshake so individual publishers do not have to negotiate with every model lab.
Incumbents get an asymmetric benefit from the new default. Google, Apple, and Microsoft run crawlers that double as search indexers and AI training sources, so any publisher who blocks one behavior effectively loses the others. A site that wants to stay in Google search while refusing Google training has to whitelist the bot and rely on the crawler's own opt-out, rather than flip a Cloudflare switch. The result is that the firms with the largest existing crawlers can keep training on publishers who need to remain in search, while smaller AI labs and startups that only crawl for training can be cleanly turned away.
Cloudflare frames the prior "we crawl you, you get referrals" deal as broken, citing Google's move from a list of blue links to a full answer engine that returns answers directly rather than sending clicks back. Community reaction on Hacker News reads the shift as overdue for some publishers and risky for others, with commenters split on whether the new default favors large crawlers or finally makes the cost of crawling visible.
The open questions are enforcement and adoption. Cloudflare says it will identify bots by behavior rather than self-declared user-agent strings, but the company has not published the detection method in detail. A publisher who opts out of training will want to know whether a crawler that claims to only index will also train, and whether Cloudflare can prove it isn't. The Pay-Per-Crawl marketplace has been live since July 2025; its price discovery across thousands of small sites is the test case for whether the third option is a working market or a label.
September 15 is the date to watch. That is when the behavioral defaults become the rule rather than the opt-in, and when the gap between "block AI" and "disappear from search" stops being a setting and starts being a bill.