
Cloudflare is changing one of the basic assumptions behind web crawling: allowing an automated crawler onto a website no longer has to mean granting every use of the information it collects.
On September 15, 2026, Cloudflare announced a new Disallow AI Training setting alongside separate controls for Search, Training and Agent traffic. The goal is to let website operators remain discoverable through traditional search while separately refusing the use of their content for AI model training.
Search and AI Training Are Different Uses
A publisher may want a search engine to index its pages because search results can send readers back to the original website. The same publisher may not want those pages collected to train or fine-tune an AI model. Cloudflare says fewer than 1% of sites on its network block Search bots, while 17% enable some mechanism to block Training.
The difficult case is a mixed-use crawler that performs both functions. Cloudflare says its new setting publishes the applicable no-training preference while allowing qualifying mixed-use crawlers to continue accessing the site for search. Training-only crawlers can be handled independently without automatically eliminating search visibility.
Robots.txt Is a Preference, Not a Firewall
For administrators, the technical distinction is important. A robots.txt directive can state what a site owner wants a crawler to do, but the file cannot identify a crawler’s real purpose or physically stop software that ignores the instruction. Network enforcement can add that missing layer by identifying traffic, classifying its behavior and blocking crawlers that violate policy.
Cloudflare is also replacing its older Managed Robots.txt approach with Bot Preference Sync. The system keeps a site’s published robots.txt preferences aligned with the Search, Agent and Training policies configured at the network edge.
This makes crawler administration increasingly similar to ordinary IT access-control policy. Administrators define what a class of automated clients is permitted to do and then try to keep the public declaration and technical enforcement consistent. BitcoinVersus.tech has previously covered related infrastructure controls in its FortiGate Firewall Setup Guide and identity management in Kerberos: Network Authenticator.
Mixed-Use Crawlers Get New Rules
Cloudflare introduced an Accountable designation for crawler operators that meet or make time-bound commitments to transparency requirements. These include a mechanism for opting out of AI training, a mechanism for opting out of AI summaries, URL-level visibility into content made available for training, and assurance that declining AI training will not reduce traditional search results. Cloudflare says Apple, Google and Microsoft meet or have committed to the qualifications.
Cloudflare also says relevant crawlers from Amazon, Anthropic, Meta and OpenAI separate their Search and Training functions, allowing the Training crawler to be blocked without necessarily removing search access.
For IT teams, the broader change is straightforward: crawler management is no longer only an SEO setting. It now intersects with network security, content governance, intellectual-property policy and AI infrastructure.
The administrative question is becoming more precise than “Can this bot access the website?” It is increasingly “What is this bot allowed to do with the information after it accesses the website?”
BitcoinVersus.Tech Editor’s Note:
We volunteer daily to ensure the credibility of the information on this platform is Verifiably True. If you would like to support to help further secure the integrity of our research initiatives, please donate here: 3C9o19EH5HSiwEPyCTmEKzxhNCbo2X6TTb
https://x.com/1BitcoinVersus/status/1937006164555993338
BitcoinVersus.tech is not a financial advisor. This media platform reports on financial subjects purely for informational purposes.
Leave a comment