Cloudflare Introduces 'Disallow AI Training' Setting to Block AI Scrapers While Preserving Search Indexing
Cloudflare has introduced a new "Disallow AI Training" setting that allows web administrators to permit search engine indexing while preventing site content from being used to train AI models. Tech giants including Apple, Google, and Microsoft have pledged to honor this directive. This feature resolves a persistent dilemma for site owners who want to protect their intellectual property from generative AI models without sacrificing search engine visibility and web traffic. By establishing clear standards recognized by major tech companies, it marks a significant step forward in web data governance. Under the new policy, Cloudflare still permits hybrid crawlers that serve both search and AI training if they are flagged as "Accountable," though blocking certain crawlers may impact search rankings. Additionally, existing users with AI blocking enabled are automatically migrated to this setting, while Google, Apple, and Microsoft are rolling out complementary URL-level and robots.txt mechanisms.
## BACKGROUND
Web crawlers traditionally gather online content to build search engine indexes, helping users discover web pages via queries. However, the surge in generative AI led tech companies to harvest vast amounts of web data for model training, often blending search indexing and AI scraping into the same automated systems. As a result, publishers previously faced an all-or-nothing choice between blocking AI scrapers or losing search engine traffic.