Cloudflare Just Made AI Companies Choose: Pay Up or Get Blocked
Cloudflare Just Made AI Companies Choose: Pay Up or Get Blocked
I was on my second coffee when the email arrived. The subject line said something about AI crawlers and a September deadline. I almost archived it.
The bots won.
That was how Cloudflare CEO Matthew Prince put it in the announcement that came out July 1. Non-human traffic crossed over human traffic on the internet sometime this year — earlier than anyone expected. And Cloudflare, which sits in front of a huge chunk of that traffic, just drew a line in the sand.
Starting September 15, 2026, Cloudflare’s default settings will block any crawler that cannot tell the difference between searching the web and training on it. If a bot blends search, agentic use, and training together — the way most major AI companies currently operate — it gets blocked from any site running ads, unless the site owner manually opts in. No asking. No negotiation. Just a date and a default.
The publisher side of this is not hard to understand. Most sites running ads make money from human visitors. They are not making money from having their investigative reporting scraped into a training dataset that gets used to answer questions without linking back. The exposure argument only works up to a point. When the exposure stops converting to human readers, it stops being exposure and starts being something else.
The mechanism behind this is also worth understanding. Cloudflare sits in front of roughly 20 percent of all web traffic, which means it has unique visibility into what is actually crawling a site and how. That position is what lets them enforce this distinction in the first place. Smaller CDN providers or registrars do not have this leverage. It is a specific company using a specific market position to set a new industry norm. Whether that is good for the internet or just good for Cloudflare depends on who you ask, and the answer usually correlates pretty directly with who is answering.
The mechanism behind this is also worth understanding. Cloudflare sits in front of roughly 20 percent of all web traffic, which means it has unique visibility into what is actually crawling a site and how. That position is what lets them enforce this distinction in the first place. Smaller CDN providers or registrars do not have this leverage. It is a specific company using a specific market position to set a new industry norm. Whether that is good for the internet or just good for Cloudflare depends on who you ask, and the answer usually correlates pretty directly with who is answering.
I have talked to enough small publishers over the past year to know the frustration runs deep. One editor at a trade publication told me they watched their entire archive get ingested by a major AI company and then saw AI-generated summaries of their stories appear on other platforms with no attribution and no traffic coming back. The follow-up was cordial and completely unproductive. There was no mechanism to do anything about it.
What changed is the scale. Prince put a number on it in the announcement: the world’s largest search engine — clearly referring to Google, though never named directly — has access to roughly twice as much information as other AI companies. The reason, according to Cloudflare, is that Google makes it difficult for site owners to stay discoverable in search without simultaneously being used for AI training. You cannot easily separate the two. Google disputes this. The company points to Google Extended, a bot that lets publishers opt out of having their content used for AI training and AI products. Using it does not affect search rankings. Cloudflare’s point is about the defaults, not the options. Defaults matter because most site owners never change them.
Here is what Cloudflare is actually doing. Mixed-use crawlers will be blocked by default on any page with ads. Site owners who want to allow them can change their settings. This applies to new Cloudflare customers immediately, new sites added by existing customers, and every existing customer on the free tier. The paid tiers get additional controls.

The payment infrastructure is the second piece. Cloudflare already runs a marketplace called Pay Per Crawl, where sites can charge bots for access. It is now expanding into Pay Per Use — not just charging when content gets fetched, but when it creates value. The initial partners are Ceramic.ai and You.com, both paying publishers when their content appears in AI-powered search results or premium access products. Other AI companies can build on the same model if they want access.
The third piece is efficiency. Cloudflare’s own data shows that over 50 percent of crawl traffic from AI bots is spent re-fetching pages that have not changed. Separate crawlers with proper caching signals fix that. The defaults Cloudflare is changing are partly about economics, not just ethics.
September 15 is the date to watch. By then, the defaults will have shifted for a large portion of Cloudflare’s free tier customers — which, in practice, is most of the internet. AI companies that have not separated their crawlers will start seeing gaps in their training data and degraded performance in agentic products that rely on real-time web access.

The question is whether this changes behavior or just creates a new negotiating position. AI companies could build separate crawlers. They could pay for access through Cloudflare’s marketplace. They could do nothing and lose access to a significant portion of the web. History suggests they will do some combination of all three, depending on how much they need each publisher’s content.
The companies most exposed are the ones that built their data strategies around the assumption that the internet was free to scrape. That assumption was always provisional. Cloudflare just made it expire faster.
The footnote here is that Google actually offers the opt-out mechanism that Cloudflare is implicitly demanding. The problem is that it is opt-out, not opt-in. Defaults matter. Cloudflare is betting that most publishers, given a real choice and a real payment mechanism, will not choose free.
For enterprise security and data teams, the practical question is what happens to training datasets that are already built. Models trained on data scraped before September 15 are not retroactively blocked. The change affects future access. Companies that have already ingested large portions of the web may find their competitive advantage actually increases — they paid the implicit price already, in compute and electricity, and now their competitors have to make different choices. That is not an ethical point. It is just a description of how infrastructure transitions usually work.
For enterprise security and data teams, the practical question is what happens to training datasets that are already built. Models trained on data scraped before September 15 are not retroactively blocked. The change affects future access. Companies that have already ingested large portions of the web may find their competitive advantage actually increases — they paid the implicit price already, in compute and electricity, and now their competitors have to make different choices. That is not an ethical point. It is just a description of how infrastructure transitions usually work.