Policy

Cloudflare's September 15 Deadline Forces AI Companies to Choose: Search or Training

Cloudflare will block multipurpose web scrapers from ad-supported sites by default, pushing AI firms to separate search bots from training crawlers and pay publishers for content access.

Last verified:

Cloudflare is reshaping the economics of AI content licensing by enforcing a hard technical boundary between discovery and extraction. According to TechCrunch AI, the infrastructure company will begin blocking dual-purpose web crawlers—those serving both search indexing and model training—from sites with advertising on September 15, 2026. The restriction applies to new Cloudflare customers, newly registered domains under existing accounts, and all free-tier users. Paid customers with existing configurations retain their current rules unless they actively update them.

The Crawler Separation Mandate

The core mechanic is straightforward: multipurpose bots will face blockage by default on ad-supported pages unless the site operator explicitly permits them. This forces AI model providers and infrastructure companies to declare their intent upfront—either a crawler indexes for search, or it ingests for training, but not both under a single user agent string. TechCrunch reports that Cloudflare frames this as addressing publisher concerns about giving away intellectual property without compensation while remaining discoverable through legitimate search engines.

The policy implicitly targets Google’s dual advantage. Cloudflare notes that the search giant gains access to roughly twice the content surface area available to competitors because site owners face a trade-off: remain invisible to Google Search, or accept that Googlebot will feed data into Gemini and AI Overviews. Google has countered this framing by highlighting Google Extended, a separate bot that site owners can block independently without losing Search visibility. However, the company’s main Googlebot continues to serve both search results and AI-powered features.

Pay Per Use: Monetizing Content Extraction

Cloudflare is evolving its marketplace tooling beyond flat-rate scraping fees. The company is shifting toward “Pay Per Use” pricing, which charges AI companies based on the value extracted from publisher content—not just the act of accessing it. This creates a new cost lever: publishers can price based on how frequently their content appears in AI training sets or agent outputs, rather than metering crawl volume.

Why This Matters

This policy reorders the power dynamic in AI content economics. Publishers gain a practical tool to monetize what was previously free extraction. Smaller AI companies face higher compliance costs—they must now negotiate separately for search and training access, whereas incumbents with direct publisher relationships can absorb licensing fees. The September 15 enforcement date is only 10 weeks away, signaling Cloudflare’s intent to shift internet norms before bot traffic consolidates further. If enforcement holds, AI developers will need to redesign their crawl infrastructure and budget for publisher licensing as a line item, not a technical detail.

Frequently Asked Questions

Who does Cloudflare's new policy affect?

All new customers, new domains added by existing customers, and all free-tier users. Existing paid customers retain current settings unless they change them.

Can AI companies still access web content after September 15?

Yes, but only if site owners explicitly allow it. Publishers can now negotiate licensing fees through Cloudflare's Pay Per Crawl and Pay Per Use tools.

What's the difference between search crawlers and AI training crawlers?

Search crawlers index pages for discoverability (Google Search, Bing). Training and agent crawlers extract data to build models or power agentic features. Cloudflare now requires these to operate under separate bot identities.

#content licensing #web scraping #AI training #publisher rights #Cloudflare