Obenan

Obenan Briefing · Signal

On July 1, 2026, Cloudflare split AI traffic into Search, Agent, and Training.

A major edge platform stopped treating AI access as one on-or-off switch. AI traffic is now three separate operating decisions: stay discoverable, govern real-time agents, and limit training crawls, with a crawl-to-referral ratio to test whether that traffic sends anything back.

Published July 8, 2026

The one line

AI visibility governance is becoming a question of who may read, for what purpose, and what they send back, not a single block-AI-bots toggle. On September 15, 2026 Cloudflare announced a setting to refuse AI training while staying in search; whether your own domain is set that way is a separate check.

Published
July 8, 2026
Format
Signal briefing
Sources
7 primary, public
Coverage
Cloudflare releases, July 1 and September 15, 2026

Public sources only. Observed facts and our interpretation are kept separate.

The 60-second read

Three questions, three answers.

What Cloudflare changed at the edge, why governing AI traffic by behavior matters, and what it still does not decide.

Q1What changed

AI traffic became three separate decisions.

On July 1, 2026, Cloudflare replaced its single Block AI bots preset with behavior-level controls that separate AI crawlers into Search, Agent, and Training. In July, every customer, including the Free tier, could allow a behavior, block it on all pages, or block it only on pages that display ads. On September 15, 2026, Cloudflare added a fourth setting, Disallow AI Training, which is available only for Training.

Q2Why it matters

A site can stay discoverable while being stricter with agents and training.

Because the three behaviors are governed separately, a business no longer has to make one all-or-nothing choice. It can keep Search access that supports discoverability while treating real-time agent traffic and extractive training crawls differently. On September 15, 2026, Cloudflare announced a Disallow AI Training setting that, in its words, "lets you easily stay indexed for search while refusing to let that same crawler train on your content." From that date, a new domain is offered one of two recommended presets when it is set up, and the owner can change them at any time.

Q3What is still missing

Whether the fact behind the page is true, current, and complete.

Controlling who may crawl does not make the answer an AI gives about a business correct. A crawl-to-referral ratio can show whether a bot returns any visitors, but it does not prove recommendation quality, and it does not fix the facts a business publishes.

What to do with this

Six moves for the people governing AI access.

Split by what a site owner controls and what the platform controls.

  1. 01

    Treat AI access as three decisions, not one (operator-controlled)

    Decide separately whether you allow Search discoverability, Agent activity, and Training crawls. A single block-everything setting is now the blunt option, not the only one.

  2. 02

    Keep Search open if discoverability matters (operator-controlled)

    For most local and multi-location businesses, being reachable by search-oriented crawlers supports being found. To refuse AI training and stay in search, Cloudflare's September 15, 2026 post points to Disallow AI Training. Block is stricter than it sounds: Cloudflare says Block on Training now also stops mixed-use crawlers such as Googlebot, Bingbot, and Applebot, which crawl for both search and training, so it can cost you search discoverability too. Cloudflare also says that until Microsoft's support arrives, targeted for early 2027, Disallow AI Training does not pass the no-training preference to Bing through robots.txt.

  3. 03

    Read crawl-to-referral as pressure, not proof (shared)

    The ratio tests whether a bot operator sends visitors back for the content it takes. Use it to question extractive crawlers, not to claim conversion, recommendation, or revenue.

  4. 04

    Edge controls are platform-owned, not merchant truth (platform-controlled)

    Cloudflare governs crawler access and classification at the edge. It does not author the hours, address, services, and policy an AI repeats about a business. That truth stays yours.

  5. 05

    Separate access from recommendation when you report (shared)

    Keep two columns. Did crawlers reach the site, and separately, is the fact the AI repeats correct. A green access setting is not a correct answer.

  6. 06

    Verify your own zone before you believe a default (operator-controlled)

    A vendor announcement tells you what a product now offers. It does not tell you what your own domain is set to. Before you report a policy internally, open the Cloudflare zone that serves your site and read the Search, Training, and Agent settings it actually has. Then fetch your own robots.txt, the short public file at the root of a website that tells automated visitors what they may read (for a restaurant, yourrestaurant.com/robots.txt), and check what it now publishes and whether an older Managed Robots.txt preference has moved. Announced and in force are two different facts, and only one of them is about you.

The evidence, dated

Two releases, six moving parts.

The launch wave on July 1, 2026, then Cloudflare's September 15, 2026 release on the date it had set in July, shown in the order Cloudflare published them.

01Cloudflare · July 1

Behavior-level AI traffic controls launch for all customers

Cloudflare announced that all customers, including the Free tier, can now manage AI traffic by Search, Agent, and Training, because not all AI traffic serves the same purpose.

02Cloudflare · July 1

Each behavior gets its own allow or block option

The changelog says each behavior can be left unblocked, blocked on all pages, or blocked only on pages that display ads. It also said that, starting September 15, 2026, new domains would default to blocking Agent and Training on ad pages while Search stays allowed. Cloudflare's September 15 post describes that change differently: a new domain is offered one of two recommended presets when it is set up, and the owner can change them; for a site that earns money from ads, the preset sets Training to Disallow AI Training rather than blocking it.

03Cloudflare · July 1

Attribution Business Insights adds a value signal

For Bot Management Enterprise customers, a new dashboard shows bot traffic to content pages and site-wide and per-operator crawl-to-referral ratios over 24 hours, 7 days, or 30 days. Cloudflare describes it as visibility for a wider set of stakeholders, not a new control plane.

04Cloudflare · July 1

BotBase exposes per-bot behavior and detection IDs

BotBase shows how Cloudflare classifies each bot by behavior, including Search, Agent, and Training, exposes a detection ID for each, and publishes every tracked bot in Cloudflare Radar's public bots and agents directory.

05Cloudflare · September 15

The promised date arrives with a named training refusal

Cloudflare announced a new Disallow AI Training setting that, in its words, "lets you easily stay indexed for search while refusing to let that same crawler train on your content". Cloudflare says accountable mixed-use crawlers remain allowed for search, and lists four requirements a crawler operator must meet, or commit to meeting, to be called Accountable. One of them is "Assurance that opting out of AI training will not affect traditional search results".

06Cloudflare · September 15

Managed Robots.txt is set to give way to Bot Preference Sync

Cloudflare says "Managed Robots.txt will be deprecated in favor of Bot Preference Sync. Customers who enabled Managed Robots.txt will migrate to the new system." Its Bot Preference Sync post, published August 21, 2026 and updated September 15, 2026, states the intent as "the preference you set is the preference you publish", and says the feature is available to all customers from the Free tier to Enterprise.

In July, AI-traffic policy went from a single switch to a set of behavior-specific decisions with a business-readable value signal attached. In September, Cloudflare announced a setting to refuse AI training while staying in search, and said Managed Robots.txt will give way to Bot Preference Sync.

The distinction that matters

A traffic control is not a true answer.

Governing which crawlers may reach a site decides whether an AI can read it. It does not decide whether what the AI then says about the business is accurate, current, and complete.

For a static page the gap is small. For a local business whose hours, availability, services, and policy change, the fact behind the page is where most AI answers go wrong, and no crawl setting fixes that.

Accurate identity

Name, address, and category that match reality across the surfaces an agent reads.

Current hours and availability

Open hours, special hours, and real availability, not a list that is quietly out of date.

Complete policy

Booking, deposit, cancellation, and service terms an agent can act on without guessing.

Services and attributes

What the business actually offers, described the way a customer would ask for it.

Freshness

Facts that are current at the moment of the answer, not last quarter's truth.

Evidence

Reviews and signals that corroborate the claim, so the answer holds up.

Where it breaks today

When the access policy is right and the answer is wrong.

Each of these can happen on a site whose AI traffic controls are set exactly as intended.

01

A site allows Search and blocks Training, and an AI still repeats hours that were changed three weeks ago.

02

An operator sees a healthy crawl-to-referral ratio for one bot and assumes it is recommending the business, when the ratio only counts referred visits, not recommendations.

03

A business blocks Agent traffic to protect content, then finds that a real customer's assistant cannot complete a task it could have handled.

04

A crawler is allowed and cleanly classified, and an answer engine still cites an address the business corrected months ago.

05

A multi-location operator reads that Managed Robots.txt customers will migrate to Bot Preference Sync, tells the marketing team the training refusal is now in force, and never opens the zone to confirm which setting the domain is actually on.

In each case the access policy was set correctly. The failure was in the truth behind the page, or in reading a traffic metric as a recommendation.

Where the work sits

Three lanes of AI traffic, governed separately.

Cloudflare's release splits AI traffic into three behaviors that can now be allowed or blocked on their own terms, instead of one Block AI bots switch.

The point is not the labels. It is that discoverability, real-time agent activity, and model training are now separate operating decisions instead of one undifferentiated bucket.

01Search

Stay discoverable

Crawlers that support being found and indexed. Most businesses keep these allowed so they remain reachable.

Decision: keep discoverability

02Agent

Govern real-time activity

Autonomous agents acting on a person's behalf in real time. Allow, limit, or block depending on whether agent activity helps or extracts.

Decision: govern agent access

03Training

Limit extractive crawls

Crawls that gather content to train models. Since September 15, 2026, Cloudflare offers Disallow AI Training to refuse training while staying in search. Block is the stricter choice and also stops crawlers that serve search.

Decision: limit training

AI traffic as three governance decisions, not one switch.

Obenan's view

An edge can decide who reads your site. Only the merchant can make the answer true.

The Cloudflare release is a real step. Treating AI traffic as three decisions is more useful than a single block-AI-bots switch, and a crawl-to-referral ratio gives operators a concrete way to question extractive crawlers.

Obenan works on the layer a traffic control cannot reach: whether the hours, address, services, availability, and policy an AI repeats about a business are accurate, current, and complete. Access is necessary. A correct answer is the goal.

How to read this

Observed, inferred, and watched. Kept separate on purpose.

This is a Signal briefing. We report what the public record shows, then our interpretation, then what we are still watching.

Observed

Cloudflare's July 1, 2026 launch, changelog, and docs describe the Search, Agent, and Training controls, the three per-behavior options offered in July, Attribution Business Insights for Bot Management Enterprise, and BotBase classifications published to Radar. Cloudflare's September 15, 2026 post announces a Disallow AI Training setting for Training only, says Block and Block on pages with ads now also stop mixed-use crawlers such as Googlebot and Bingbot, and describes two recommended presets offered to new domains, which the owner can change, in place of the default the July changelog described. It defines an Accountable designation that, in its words, recognizes "capabilities available today and concrete commitments to deliver them", says Microsoft's robots.txt no-training support is targeted for early 2027, and says Managed Robots.txt will be deprecated in favor of Bot Preference Sync. Re-checked September 22, 2026; the September 15 post and the Bot Preference Sync post were re-read October 5, 2026.

Inferred

That AI traffic governance is becoming three separate operating decisions, and that crawl-to-referral is a useful business-pressure metric, is Obenan's reading of the public record, not a Cloudflare claim.

What we are watching

Whether behavior-segmented controls and crawl-to-referral reporting become standard operator practice, whether what a site owner finds in their own zone matches what the announcement describes, and whether any of it begins to touch the accuracy of the facts AI systems repeat about a business.

What this briefing does not claim

The boundaries, stated plainly.

Precision protects the reader and the companies named. These are the limits of the claim.

  1. 01

    This briefing does not claim Cloudflare's controls improve rankings or recommendations in ChatGPT, Gemini, Perplexity, Google AI Mode, or any AI system.

  2. 02

    Attribution Business Insights is available only to Bot Management Enterprise customers, not to all Cloudflare customers.

  3. 03

    Obenan is not affiliated with, partnered with, integrated with, or endorsed by Cloudflare, OpenAI, Microsoft, Google, or any named source. They are public source subjects only.

  4. 04

    The crawl-to-referral ratio measures crawler volume against referred visitors. It is not conversion lift, recommendation quality, citation quality, revenue, or ROI.

  5. 05

    Cloudflare's classifications do not perfectly capture every AI request or every operator intent, and are not a merchant-truth or recommendation signal.

  6. 06

    Obenan is the guide, not the hero: it does not operate edge infrastructure, crawl control, or bot classification, and this Signal reports a Cloudflare change, not a Obenan product launch.

  7. 07

    Obenan reports what Cloudflare announced on July 1 and September 15, 2026. No authenticated Cloudflare zone, customer domain, generated robots.txt response, or migration state was checked for this briefing, so nothing here describes the effective policy of any real site.

If you govern AI access, who fixes the answer?

Cloudflare can decide which crawlers reach your site. Obenan works on the merchant-truth layer a traffic control cannot reach: the facts an AI repeats about your business.

Public research. No partner status is claimed or implied.

Sources

All primary and public. Each can be inspected directly. Seven official Cloudflare pages covering the July 1 and September 15, 2026 releases, each marked with the date it was last checked.

Primary sources

  1. 1.
    Cloudflare: Your site, your rules, new AI traffic options for all customersblog.cloudflare.com · July 1, 2026 · checked September 22, 2026

    Primary. Launch of the Search, Agent, and Training controls, all-customer availability including the Free tier, and the new-domain default Cloudflare then planned for September 15, 2026. Source 6 describes what Cloudflare offered on that date instead: two recommended presets the owner can change.

  2. 2.
    Cloudflare: New options to manage AI traffic (changelog)developers.cloudflare.com · July 1, 2026 · checked September 22, 2026

    Primary. Per-behavior allow or block options and the new-domain default the changelog then planned for September 15, 2026. Source 6 reports what Cloudflare offered on that date instead.

  3. 3.
    Cloudflare: Unmasking the crawls with Attribution Business Insightsblog.cloudflare.com · July 1, 2026 · checked September 22, 2026

    Primary. Dashboard purpose, crawl-to-referral ratios, top-bot views, and business-facing reporting scope.

  4. 4.
    Cloudflare: Attribution Business Insights (docs)developers.cloudflare.com · July 1, 2026 · checked July 8, 2026

    Primary. Bot Management Enterprise availability, the crawl-to-referral ratio definition, and the visibility-not-a-control-plane boundary.

  5. 5.
    Cloudflare: More visibility into bot traffic with BotBase and Attribution Business Insights (changelog)developers.cloudflare.com · July 1, 2026 · checked July 8, 2026

    Primary. BotBase behavior classifications, per-bot detection IDs, and publication to Cloudflare Radar's public bots directory.

  6. 6.
    Cloudflare: Have it both ways: stay discoverable in search while disallowing AI trainingblog.cloudflare.com · September 15, 2026 · checked September 22, 2026

    Primary. Announcement of the Disallow AI Training setting, the four requirements for an Accountable crawler operator, the presets recommended to new domains from September 15 for ad-supported and non-ad-supported sites, and the planned deprecation of Managed Robots.txt in favor of Bot Preference Sync.

  7. 7.
    Cloudflare: Say it once: introducing Bot Preference Syncblog.cloudflare.com · August 21, 2026, updated September 15, 2026 · checked September 22, 2026

    Primary. How a configured preference is published into robots.txt, the stated intent that the preference you set is the preference you publish, and availability to all customers from the Free tier to Enterprise.

Obenan Briefings. AI Visibility and Discoverability lane. This is a Signal briefing: it reports a public development and Obenan's interpretation, separated on purpose.