Robert Hu
GEO & SEO

AI Visibility Is Becoming a Permission Stack

Robert Hu··6 min read
Four separate permissions layered over one web page: search indexing, model training, AI summaries and agent access, each granted or withheld independently

For two years the question was binary. Can AI systems read my site. Blocking the crawler meant disappearing from search, so almost nobody did, and the question stopped there.

On September 15, Cloudflare took that question apart.

What actually changed

Cloudflare now classifies crawler behavior into three controls: Search, crawling to build a search index; Training, crawling to train or fine-tune a model; and Agent, user-directed agents visiting a page on behalf of a person.

The problem it fixed was the mixed-use crawler, one crawler doing both search and training. Refuse one and you refused the other. A new Disallow AI Training setting publishes a no-training preference in robots.txt while leaving qualifying mixed-use crawlers free to keep crawling for search.

Qualifying means what Cloudflare calls Accountable: the operator must offer opt-outs from AI training and AI summaries, give URL-level visibility into which pages were made available for training, and assure that opting out of training will not affect search results. Cloudflare says Apple, Google and Microsoft meet those requirements or have committed to timelines for the rest. It also calls the relevant Amazon, Anthropic, Meta and OpenAI crawlers Accountable, since those companies run separate search and training crawlers, so the training one can be blocked without touching search.

Two numbers frame the whole thing. Fewer than 1% of Cloudflare sites choose to block search bots. Seventeen percent use some mechanism to block training. Almost nobody refuses discovery; a meaningful minority already refuses one use of what gets discovered.

Newsletter

Follow the research

Research notes and analysis on how AI, digital transformation, product discovery, and customer behavior are changing commerce.

Subscribe to Hu's Weekly Hoot

The part worth checking yourself

That claim is worth checking at the source rather than taking from a vendor.

Google's own crawler documentation says Google-Extended manages whether crawled content may be used for training future Gemini models and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI. It then says plainly that Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search."

Worth reading precisely. That token's stated scope is Gemini apps and Vertex grounding, not generative features inside Search itself, and Cloudflare notes Google offers a separate webmaster-portal toggle for generative search results. Anyone assuming Google-Extended pulls them out of AI Overviews is assuming something Google's documentation does not say.

On Cloudflare's account, Microsoft is further back. Cloudflare says Bing currently expresses training preferences through the NOARCHIVE meta tag, that Microsoft is building support for a no-training preference in robots.txt targeted for early 2027, and that until then selecting Disallow AI Training does not convey that preference to Bing. I found no current Microsoft documentation saying this in Microsoft's words, so treat it as Cloudflare's description.

Four permissions, four different levels of maturity

Search controls are old and universally understood. Training controls just became usable without sacrificing search, at least for the operators Cloudflare certifies. Summaries are earlier: an opt-out is an Accountable requirement, and Cloudflare says its goal is to let owners control how much content is included, set in one place, by early next year. Agents are earliest. Cloudflare ships no Disallow setting for agents, because the Internet "does not yet have a well-established directive for expressing Disallow preferences to agents," pending standards like ai-prefs.

So four permissions exist as concepts. Search and training have clearer mechanisms today, while summary-specific and agent-specific controls remain fragmented and less standardized. The permission stack is more mature as a business concept than as a control panel.

Where this sits in the AI visibility picture

I have written about retrievability, whether an AI system can find and understand your product, about agent evaluation and purchase, and about what Search Console's generative AI report can and cannot tell you.

Permission is the piece between them. Machine-readable does not automatically mean machine-authorized. A system may technically be able to retrieve a page while the business grants different rights for search, training, summaries or agent activity. That is an operating and technical distinction rather than a legal one, and a preference published to a crawler is a request that cooperating operators honor, not a right that enforces itself.

So AI visibility now involves three separate questions, and they have different owners. Can the system technically retrieve the content. What uses does the business permit. And can the business measure the economic value of what results. Retrievability, permission and measurement is a more honest description of the operating model than any one of them alone. Permission stack is my shorthand for the middle question, not anyone's product name or industry terminology.

The commerce question Cloudflare raises and does not settle

Cloudflare is explicit that the right answer depends on the business model. Its onboarding presets differ for ad-monetized sites, on the logic that ad revenue needs a human to see the page. A publisher funded by advertising may optimize for audience volume, it argues, while a retailer may prefer fewer visitors who are likelier to buy.

It supports that with figures: more than half of consumers read summaries in Search and are over 40% more likely to end their search afterward, while consumers referred by AI search convert at three to five times the rate of traditional search referrals. Cloudflare publishes no source or methodology for either number, and the second sits awkwardly beside Adobe's finding that AI-referred shoppers convert 42% better. Three to five times is a different claim from 42%, and I cannot say which is right. Neither discloses enough to reconcile them: merchant populations, how AI-referred traffic is defined, attribution windows, which AI sources count, journey stage and what counts as lift could each differ. The conclusion is not that one party is wrong. It is that no stable cross-platform benchmark exists yet for the commercial value of AI-referred traffic, which matters because that value is what a permission decision is supposed to weigh.

The policy point survives the uncertainty. Access policy should follow the economics of the interaction rather than a reflexive block-everything or allow-everything position. Nothing here says every retailer should allow summaries and agents, or that any should block them.

What this asks of operators

These are business decisions wearing technical clothing. Search discoverability is settled for almost everyone. Whether your content trains somebody's model, whether it can be summarized in place of a visit, and whether agents may fetch it are open, and they belong with whoever owns the economics of the site rather than defaulting to whoever administers the DNS.

The objections

Most merchants benefit from maximum discovery and may never need any of this. Restricting access today could cost visibility in systems that matter more later, a risk nobody can price.

Declarations also depend on compliance. Cloudflare can identify and block crawlers on its own network, but a published preference is a request. Accountable is Cloudflare's designation, not an industry standard, and Cloudflare defines, certifies and reports on it.

The hardest part is measurement. Deciding whether summaries or agent access are worth their cost means attributing outcomes to each permission separately, which almost nobody can do today.

If you had to write your company's policy on training, summaries and agent access this quarter, who in your organization would actually own that decision?

Follow the research

I publish research notes and analysis on how AI, digital transformation, product discovery, and customer behavior are changing commerce.

Subscribe to Hu's Weekly Hoot