Servers Online195+ Locations
Cloudflare Will Block AI Crawlers by Default on September 15, 2026 — Here's What Actually Changes

Cloudflare Will Block AI Crawlers by Default on September 15, 2026 — Here's What Actually Changes

On September 15, 2026, Cloudflare starts blocking AI training and agent crawlers by default on ad-supported pages. Here's exactly what changes, who is affected, and how data teams and site owners should prepare.

The short version: on July 1, 2026, Cloudflare announced that starting September 15, 2026, it will block AI training and AI agent crawlers by default on pages that carry ads. Search crawlers stay allowed. Mixed-use crawlers that blend search, training, and agent behavior — including Googlebot, Bingbot, and Applebot — get blocked wherever a site owner has chosen to block training. The new defaults apply to domains that onboard to Cloudflare from that date, new sites created by existing customers, and existing free-tier zones. If you collect web data, run agents, or operate a site behind Cloudflare, the rules of access are about to change on a hard deadline.

This guide covers what was announced, who is affected, and what to do before September 15 — whether you run crawlers or run a website.

Key dates at a glance

  • July 1, 2026 — Cloudflare announces the new crawler taxonomy and defaults (its second "Content Independence Day"; the first, July 1, 2025, introduced the original one-click AI-bot block).
  • Now → September 14, 2026 — the new AI traffic controls are already live in every Cloudflare dashboard; existing customers can set their own preferences before defaults kick in.
  • September 15, 2026 — training and agent crawlers are blocked by default on ad-supported pages for new zones and free-tier zones.

In this article

What Cloudflare announced on July 1, 2026

Cloudflare's announcement is a response to a milestone the company itself measured: bots now generate the majority of web traffic, and AI crawlers are the fastest-growing slice of it. CEO Matthew Prince framed the change bluntly: "Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge." Cloudflare also noted that more than half of all AI crawler traffic is wasted effort — re-fetching pages that haven't changed since the last visit.

The core of the announcement is a shift from a single on/off switch ("block AI bots") to a purpose-based system. Cloudflare no longer asks whether a visitor is a bot — it asks what the bot is for, and applies different default rules to each purpose. Site owners get granular controls; crawler operators get an ultimatum: separate your crawlers by purpose and declare that purpose honestly, or get treated as the most-restricted category you blend in.

The new crawler taxonomy: Search, Agent, Training

Cloudflare now sorts AI crawlers into three primary categories, each with its own default treatment:

Category Cloudflare's definition Default from September 15
Search "Collects or indexes your content, so it can answer questions about it later" — site owners should expect referral traffic or compensation in return. Allowed on all pages
Agent "Automated behavior that is acting, usually in real time, on a person's behalf, to get something done right now" — e.g. ChatGPT-User, browser-use agents. Blocked on ad-supported pages
Training "A crawler taking your content to train or fine-tune a model" — content is "permanently absorbed into the underlying architecture of the AI." Blocked on ad-supported pages

Why "ad-supported pages"? Cloudflare's reasoning: "An ad is a signal that a website owner meant for a person to land there and see it." If a page is monetized by human attention, a bot that consumes the content without delivering a human is, by default, unwelcome.

The taxonomy goes further than the big three. Cloudflare now also tracks Transact, Data Collection, Security Testing, SEO, Ads Verification, Social/Link Preview, Feed Fetching, and Monitoring & Operations as distinct bot purposes — meaning legitimate price monitoring, ad verification, and uptime tooling now have named lanes of their own rather than being lumped in with AI training.

The mixed-use problem: why this hits Googlebot too

The most consequential detail is the treatment of mixed-use crawlers — bots that perform search indexing and model training with the same user agent. Googlebot, Applebot, and Bingbot are Cloudflare's own examples. Under the new rules, a multi-purpose crawler "will be blocked by customers who have selected to block Training" — the crawler inherits the treatment of its most-restricted purpose. Google disputes the framing, pointing to its Google-Extended token that lets sites opt out of training use without affecting Search rankings, but Cloudflare's system puts the burden on the crawler operator to run separate, honestly-labeled crawlers per purpose.

Who is affected on September 15

The new defaults do not flip every existing Cloudflare zone overnight. Per Cloudflare's announcement and TechCrunch's reporting, the September 15 defaults apply to:

  • New domains onboarding to Cloudflare from that date;
  • New sites created by existing customers;
  • Existing free-tier zones.

Existing customers can review and change their AI traffic settings today — the controls are live in the Security settings of every dashboard, and paid customers who do nothing keep their current behavior. But the direction of travel is clear: because Cloudflare fronts roughly a fifth of the web, and because free-tier zones are included in the default flip, a very large share of the long-tail web becomes block-by-default for training and agent traffic in one day.

Content Signals: the robots.txt extension with teeth

Alongside the blocking defaults, Cloudflare extended its managed robots.txt with a machine-readable preference line:

Content-Signal: search=yes,ai-train=no,use=reference

The new use field declares one of three permitted content-use levels:

  • immediate — a bot may interact with the content but store and reuse nothing;
  • reference (the default) — a bot may index, excerpt, and link back;
  • full — a bot may summarize and reproduce the content.

Two things make this more than a polite suggestion. First, Cloudflare ties it to enforcement: a bot that reproduces content in full cannot achieve Verified status, and unverified bots fall into the default-blocked pool. Second, the legal context has shifted — under the EU AI Act, whose general-purpose AI obligations become enforceable on August 2, 2026, machine-readable opt-outs carry real legal weight for anyone training models on scraped data. A robots.txt line is no longer just etiquette; ignoring it is becoming a documented, timestamped liability. (For the broader legal picture, see our guide to proxy and web data compliance.)

Verified bots and Web Bot Auth: identity over IP

The verification system changed in a subtle but important way. Previously, a Verified bot was allowed by default — verification was the permission. Now verification and permission are decoupled: "Verified" means a bot is allowable within its relevant category. A verified training crawler is still blocked on a page where training is blocked. To earn and keep Verified status, operators must demonstrate honest self-representation and no abuse of access.

Cloudflare also proposed a mechanism for transitive trust — carrying a bot's identity and declared purpose through intermediaries like proxy layers and agent infrastructure, using the standard Forwarded header from RFC 7239:

Forwarded: for="openai";use="reference"

This sits on top of Web Bot Auth, the cryptographic bot-identity scheme (HTTP Message Signatures) that Cloudflare moved into production earlier this year. The strategic signal for the data industry is hard to miss: the web's gatekeepers are migrating from IP reputation ("where does this request come from?") to declared, verifiable identity ("who is this and what will they do with the content?"). IP quality still decides whether undeclared traffic survives first contact — but for declared bots, identity is becoming the passport.

Pay Per Crawl becomes Pay Per Use

Cloudflare's Pay Per Crawl marketplace — which lets sites charge crawlers per request — is evolving into Pay Per Use, which lets publishers charge when their content actually creates value downstream, not just when it is fetched. Launch partners are Ceramic.ai and You.com: publishers earn when their content surfaces in Ceramic's AI search results or when You.com accesses premium content. For data teams, the takeaway is that "blocked" and "free" are no longer the only two states a page can be in — licensed access at a price is becoming a first-class option, and it will increasingly be the compliant path to content that the defaults now wall off.

What data teams should do before September 15

If you operate crawlers, agents, or any automated data collection, here is a concrete pre-deadline audit:

  1. Inventory your exposure. Identify which of your target domains sit behind Cloudflare and how much of your collection volume they represent. That is the traffic the September 15 defaults can touch.
  2. Classify your own traffic against the taxonomy. Is each workload search, agent, training, data collection, SEO, ads verification, or monitoring? If one crawler does several of these under one user agent, you now have the Googlebot problem — plan to split it into separate, honestly-labeled crawlers per purpose.
  3. Check whether your targets are ad-supported. The default block applies to pages carrying ads. Content APIs, docs, and ad-free pages are outside the default (though owners can still block manually).
  4. Pick a lane: declared or undeclared. If your use case can live with the rules — referral-friendly indexing, agent traffic acting for a real user — apply for Verified status, implement Web Bot Auth signatures, and honor Content Signals. If you stay undeclared, understand that you are choosing the adversarial pool, where detection is multi-layered and IP reputation is the first filter.
  5. Parse and log Content-Signal lines now. Even before enforcement, a stored record of the signals you saw — and honored — is cheap insurance, especially with the EU AI Act's August 2 enforcement date creating legal weight behind machine-readable opt-outs.
  6. Monitor block rates around the deadline. Expect managed challenges rather than clean 403s. Baseline your success rates per domain now so you can see the September 15 step-change and separate policy blocks from ordinary rotation and rate-limit issues.
  7. Budget for licensed access. Where a target's data genuinely matters to your product, Pay Per Use may end up cheaper — and far more defensible — than an arms race.

What site owners should do

If your site is behind Cloudflare, the controls are already in your dashboard — you don't need to wait for September 15:

  • Open your zone's Security settings and review the new AI traffic options; decide per category (search / agent / training) what you want to allow.
  • If you use Cloudflare's managed robots.txt, set your Content-Signal preferences so compliant crawlers get a machine-readable answer.
  • If you monetize content, look at Pay Per Use as a third option between "open" and "blocked."
  • If you rely on AI-search referrals for discovery, be careful about blanket-blocking Search — under the new taxonomy it remains the one default-allowed category for a reason.

Where proxies fit after September 15

An honest assessment, since this is a proxy provider's blog: Cloudflare's new system governs declared bots — crawlers that identify themselves and want to be let in legitimately. It does not change the physics of undeclared automation, but it does raise the stakes around it. As default blocking pushes more AI-adjacent collection out of the declared lane, more of it flows through headless browsers on residential and ISP IPs, which in turn sharpens the premium on IP quality: datacenter ranges are the first thing purpose-based filtering writes off, while residential proxies and static ISP proxies keep working precisely because they are indistinguishable from the human traffic that ad-supported pages exist to serve. (If you're weighing the trade-offs, start with our datacenter vs residential comparison.)

Two practical notes. First, the Forwarded: for=...;use=... transitive-trust proposal explicitly contemplates bots operating through proxy layers — declared agent traffic routed via proxies is part of the design, not a loophole. Second, whichever lane you choose, verify your infrastructure before the deadline rather than during it: our free proxy checker will tell you what your pool actually looks like from the outside.

Frequently asked questions

Does Cloudflare block all AI crawlers on September 15, 2026?

No. Search crawlers remain allowed by default everywhere. Only training and agent crawlers are blocked by default, and only on ad-supported pages, for new Cloudflare zones and existing free-tier zones. Site owners can loosen or tighten any of it.

Does the change affect existing paid Cloudflare customers?

Not automatically. Existing paid zones keep their current settings; the new defaults apply to domains onboarding from September 15, new sites created by existing customers, and free-tier zones. The controls are live for everyone now.

Will Googlebot really get blocked?

On sites that choose to block training, yes — Cloudflare states that multi-purpose crawlers like Googlebot, Applebot, and Bingbot "will be blocked by customers who have selected to block Training" until they separate crawling purposes. Google disputes the characterization, citing its Google-Extended opt-out. How this standoff resolves is the single biggest open question of the change.

Is robots.txt still enough to signal my preferences?

It's becoming more powerful, not less. Cloudflare's Content-Signal extension makes robots.txt express purpose-level preferences (search vs. training vs. reuse level), and the EU AI Act — enforceable for general-purpose AI from August 2, 2026 — gives machine-readable opt-outs legal significance for model training.

Do proxies bypass the September 15 block?

The block targets declared crawler identities, so it doesn't apply to traffic that presents as a regular browser session — which is why the change is expected to shift more collection toward browser-based setups on residential and ISP IPs. But Cloudflare's bot detection is multi-signal (fingerprinting, behavior, IP reputation together), and the compliant path — verification, Content Signals, licensed access — is getting more viable at the same time as the adversarial path gets more expensive. Choose deliberately, and get legal review for anything at scale.


Sources

Published July 20, 2026. We'll update this post when the defaults take effect on September 15, 2026.

novaproxy