Skip to content
Searcle Book a demo

Keep AI Search Access Without Opening Content to Training

Nina Okonkwo

Set Cloudflare’s Search, Agent and Training controls for B2B discovery. Compare Disallow with Block, then check robots.txt, crawler support and WAF rules.

For a B2B website that wants buyers to find its public content while refusing model-training use, start with Search: Allow, Training: Disallow AI Training, and Agent: Allow for intended public-page use, with Bot Preference Sync enabled. Review agent access separately if automated visitors create abuse or operational risk. This policy preserves discovery while expressing a training restriction; it does not guarantee that every operator will honor that restriction.

Choose your search, training and agent policies; check the access consequences beside the controls.

Check Your Crawler Policy

Discovery-Preserving Starting Policy

Search
Allow
Training
Disallow AI Training
Agent
Allow

Qualifying Accountable mixed-use crawlers retain search access; other training crawlers are blocked.

Enable Bot Preference Sync to publish applicable no-training directives. Operator support still matters.

Agent Allow adds no category block for intended public-page interactions. Review abuse risk separately.

Allow is not a bypass for other security rules. This tool checks category-policy consequences, not your live configuration.

Limits Behind These Results

Training Block includes mixed-use crawlers such as Googlebot, Bingbot and Applebot, even when Search is Allow. Block on pages with ads includes relevant mixed-use crawlers on pages Cloudflare detects as serving ads.

Disallow AI Training is not a universal guarantee against training use. Cloudflare says Bing robots.txt support for the no-training preference is targeted for early 2027. Google-Extended controls specified Gemini training and grounding uses, not every AI feature.

Bot Preference Sync does not translate complex individual custom rules into matching directives. WAF rules can block an allowed crawler, and skip rules can undermine an intended block.

Source: Cloudflare’s September 15, 2026 policy announcement; operator and WAF limits are sourced in the article.

The critical distinction is Disallow AI Training versus Block. Under Cloudflare’s September 15, 2026 changes, blocking Training can also block mixed-use search crawlers, including Googlebot and Bingbot. Disallow AI Training instead preserves access for qualifying mixed-use crawlers while publishing training preferences and blocking other training crawlers. Operator support still matters. Cloudflare’s September announcement explains the behavior and remaining gaps.

Search, Agent and Training Describe Different Uses

Cloudflare classifies automation by what it does, not simply whether its operator is an AI company. A single bot can have multiple behaviors. These three controls are available across all plans. Cloudflare’s bot documentation defines them as follows:

Behavior Purpose B2B Starting Policy
Search Collects or indexes content to answer questions later Allow public service pages, comparisons and educational content
Agent Acts in real time on a person’s behalf, such as fetching a page or using a browser Allow intended buyer interactions; review abuse risk
Training Collects content to train or fine-tune models Decide separately from search discovery

An assistant retrieving your implementation guide for a buyer is different from a crawler collecting that guide for model training. Blocking both as “AI traffic” can remove a useful discovery channel without making the training decision any clearer.

Verified does not mean you must allow it. Cloudflare verification concerns honest identification and non-abusive behavior, including respect for crawl directives and reasonable request rates. It also recognizes intermediaries operated by many end users: trusting the operator does not mean trusting every person directing it. Verified-bot documentation describes these distinctions.

Training Block Can Override Search Allow

As of October 3, 2026, Cloudflare’s announced controls distinguish four actions, applied to the relevant behavior at domain level. The September announcement supersedes the earlier three-option launch description.

Action Effect Main Caveat
Allow Adds no block for that behavior Other settings and security rules can still deny access
Disallow AI Training Publishes applicable no-training directives through Bot Preference Sync; blocks other training crawlers Training only; Accountable mixed-use crawlers retain search access
Block on pages with ads Blocks relevant crawlers on pages Cloudflare detects as serving ads Includes mixed-use crawlers
Block Blocks relevant crawlers across the domain Includes mixed-use crawlers

Cloudflare identifies Googlebot, Bingbot and Applebot as mixed-use crawlers. Setting Search to Allow does not override a Training Block. For a lead-generation site, a domain-wide Training Block can therefore conflict directly with the goal of keeping public pages discoverable.

Disallow AI Training is the more targeted starting point when you want search access but not training use. It combines published preferences with blocking for other training crawlers; it is not equivalent to a universal technical barrier against model training.

Operator Support Limits the Training Preference

Cloudflare says Bing’s robots.txt support for the no-training preference is targeted for early 2027. Until then, Disallow AI Training does not automatically convey that preference to Bing through robots.txt. “Accountable” includes commitments as well as capabilities already delivered, as described in Cloudflare’s September announcement.

For Google, the usage-control distinction is independently supported: Google-Extended is a robots.txt token, not a separate HTTP crawler identity, and restricting it does not affect Google Search inclusion or ranking. Its scope includes specified Gemini training and grounding uses—not every AI feature. Google’s crawler documentation explains that scope.

These distinctions matter when interpreting a successful configuration. A published directive confirms what your site is asking an operator to do. It does not, by itself, establish which downstream uses have stopped.

Verify the Domain Settings and Served Robots.txt

Record Policies Before Changing Them

In the domain’s Security Settings, review the Search, Training and Agent policies. Record individual crawler actions and custom security rules before changing them, so you can distinguish a category-policy change from an existing exception.

Do not infer current settings from an old “Block AI Bots” screenshot. Cloudflare announced migration of legacy blocking choices to Disallow AI Training, with separate Agent settings. Verify what your domain actually shows; the September announcement explains the migration mappings.

Inspect the Public Robots.txt Response

Confirm Bot Preference Sync is enabled if you want Cloudflare to publish your category-level preferences. It generates or prepends directives to the public /robots.txt response, so inspect that response—not only the file stored in your CMS.

Sync does not translate complex individual custom rules into matching directives. If you disable it for a custom policy, maintain the corresponding directives yourself. Cloudflare’s sync explanation describes those limits.

OpenAI provides a concrete example of separate purposes: allowing OAI-SearchBot supports ChatGPT search eligibility, while disallowing GPTBot expresses a training opt-out. OpenAI says the settings are independent and search systems may take approximately 24 hours to reflect robots.txt changes. Access does not guarantee inclusion. OpenAI’s crawler documentation also notes that robots.txt may not apply to user-initiated ChatGPT-User requests.

That exception is a reason to review Agent behavior separately rather than assuming a training directive governs every request from an AI product.

WAF Rules Can Change the Effective Policy

An Allow selection is not a universal bypass. Cloudflare says upstream WAF rules can still block an allowed crawler; skip rules can also undermine an intended crawler block. Review rule order and matching conditions rather than toggling the category policies repeatedly. WAF integration guidance covers both cases.

When a crawler cannot reach a public page, compare its category setting with the individual crawler action and matching security rules. A Search Allow setting alone is not evidence that the request reached your content.

Keep confidential information behind authentication. Crawler preferences are not access controls for private documents.

Measure Discovery Separately From Crawler Activity

Use AI Crawl Control to inspect crawler activity and robots.txt compliance. Its request and enforcement data do not establish whether content was used for training, whether a buyer saw it, or whether that buyer converted. Cloudflare’s overview describes the monitoring capabilities.

Also check WAF events when troubleshooting: upstream blocks may not appear in AI Crawl Control analytics, according to Cloudflare’s WAF guidance. A quiet crawler report can therefore require investigation outside that dashboard.

After a change, track crawler blocks, search visibility, AI referrals and qualified enquiries separately. Use a consistent buyer-query sample for AI search monitoring. Judge the policy by preserved discovery and useful buyer activity, not by a higher crawler-request count.