Feature

Cloudflare Does Not Block Every Mixed-Use AI Crawler

Cloudflare’s September 2026 crawler controls explained: new-domain presets, existing-site migrations, search risk and what publishers should verify.

Impetuous · · 4 Min Read

Cloudflare does not default-block every mixed-use AI crawler. Since September 15, 2026, its recommended preset for a new ad-supported domain is Search: Allow, Training: Disallow AI Training, and Agent: Block on pages with ads. New domains without ads are offered Allow for all three categories, while existing domains are migrated according to their previous settings.

Choose your domain’s situation to see Cloudflare’s preset or migration result.

Cloudflare Preset And Migration Check

New-domain values are recommended onboarding presets. Owners can change them.

Recommended preset: new domain with ads
ControlResult
SearchAllow
TrainingDisallow AI Training
AgentBlock on pages with ads
Bot Preference SyncEnabled

This publishes a no-training preference while retaining search access for mixed-use crawlers Cloudflare designates Accountable.

Source: Cloudflare’s September 2026 policy and migration tables. An em dash means the cited migration summary does not specify that value.

For an ad-funded publisher that depends on search discovery, Cloudflare’s recommended preset is a sensible starting point: allow search, publish a no-training preference, and block agents where ads appear. That is different from indiscriminately blocking every crawler that serves both search and training purposes.

The setting that creates the clearest search risk is Training: Block or Training: Block on pages with ads. Cloudflare now applies those enforcement choices to mixed-use crawlers, potentially affecting search access by Googlebot, Bingbot, and Applebot.

The Final Preset Differs From the July Announcement

Cloudflare’s July announcement said that, for new domains, Training and Agent crawlers would be blocked on pages displaying ads while Search remained allowed. It also said mixed-purpose crawlers would be subject to all applicable behavior rules, with the most restrictive rule winning (Cloudflare’s July announcement). Contemporary reporting consequently described mixed-use crawlers as being blocked from ad-supported pages by default (TechCrunch).

Cloudflare changed that configuration before implementation. Its September announcement gives new domains these recommended onboarding presets:

Control With ads Without ads
Search Allow Allow
Training Disallow AI Training Allow
Agent Block on ad pages Allow
Bot Preference Sync Enabled Enabled

These are recommended presets, not immutable rules. A domain owner can change each setting during onboarding or later. Cloudflare says the controls are available on all plans at the domain level under Security Settings (Cloudflare’s September policy).

Disallow AI Training Is Not an Access Block

The distinction matters for mixed-use crawlers such as Googlebot, Bingbot, and Applebot.

Disallow AI Training publishes the applicable no-training preference through Bot Preference Sync. Cloudflare allows mixed-use crawlers it designates “Accountable” to continue crawling for search. Other training crawlers are blocked. This option is available only for the Training category.

Block on pages with ads enforces a block on pages Cloudflare detects as serving ads. It cannot be expressed as a stable robots.txt rule because the affected page set can be large and change frequently.

Block prevents the selected crawler category from reaching the whole domain. Because mixed-use crawlers are now included in Training blocking, selecting Training: Block can stop Googlebot, Bingbot, and Applebot for search as well as training. Cloudflare recommends Disallow AI Training, rather than Block, when the objective is to state a no-training preference without sacrificing search access.

There is an important platform-specific limit. As of September 26, 2026, Cloudflare said Disallow AI Training did not automatically convey a no-training preference to Bing through robots.txt. Microsoft supported NOARCHIVE and was targeting robots-based support for early 2027. Publishers concerned about Bing therefore need to review Microsoft’s current controls rather than assume the Cloudflare setting alone completes the job.

More generally, a robots.txt preference is not access control: compliance is voluntary. Cloudflare’s managed robots.txt documentation says it can prepend managed directives to an origin’s existing file, but the file itself does not technically prevent requests (Cloudflare’s managed robots.txt documentation). Preference publication and network enforcement serve different purposes.

Existing Domains Keep Different Outcomes

Existing settings migrate automatically, but the outcome depends on the domain’s previous configuration.

For a domain that used only the legacy Block AI Bots control:

  • If the legacy control was disabled, Search, Training, and Agent remain Allow.
  • If the legacy setting was Block, Search becomes Allow, Training becomes Disallow AI Training, and Agent becomes Block on pages with ads.
  • If the legacy setting was Block on pages with ads, the same three-setting configuration applies.

For a domain that already used granular controls, Cloudflare preserves the practical effect where possible. Search and Agent selections carry over. A previous Training selection of Block or Block on pages with ads becomes Disallow AI Training. Cloudflare deprecated the single legacy toggle in favor of separate Search, Training, and Agent controls (Cloudflare’s migration tables).

“Automatic migration” therefore does not mean automatic blocking for every existing site. A site that previously allowed AI bots remains permissive unless its operator changes the controls.

Verify Search, Training, and Agent Behavior Separately

Do not treat the migration banner or preset label as sufficient verification.

  1. Record the effective settings. In the domain’s Security Settings, capture the Search, Training, Agent, and Bot Preference Sync values.
  2. Check the public robots.txt. Confirm that the response contains the intended no-training directives and retains necessary origin rules. This verifies the published preference, not crawler compliance.
  3. Review representative URL classes. Include pages with ads, pages without ads, and hostnames with different monetization. Use Cloudflare security events and origin logs to check the observable access behavior for each class.
  4. Watch search crawling after enforcement changes. If Training is changed to Block or Block on pages with ads, monitor requests from major search crawlers along with priority-page discovery and indexing. The crawl and indexing workflow provides a measurement structure.
  5. Retain evidence for each objective. For a policy of “search allowed, training preference published, agents blocked on ad pages,” verify each observable clause separately. Access logs cannot prove whether an operator later uses retrieved content for training. Use a post-remediation verification pass rather than assuming one successful or blocked request proves domain-wide behavior.

For publishers dependent on search discovery, the recommended Disallow AI Training preset is not the high-risk choice. The risk comes from applying a broad Training enforcement block that now reaches mixed-use search crawlers.