Cloudflare launched Bot Preference Sync, a feature that automatically updates a site's robots.txt file to match the AI bot policies (Search, Agent, Training) configured in the Cloudflare dashboard, eliminating the need to maintain a separate static file. The feature is rolling out to all customers, free through enterprise, in the coming week, and will be on by default for new signups. Cloudflare also tightened its Transparency requirements: mixed-use crawlers that both search and train must now provide additional disclosures (respecting no-training preferences, offering opt-outs from AI summaries, providing URL-level visibility, and proving that disallowing training doesn't hurt search rankings) to avoid being blocked when Disallow Training is set. Publishers relying on ad revenue get a new default of Disallow Training while staying searchable.
Table of contents
Copy link New questions facing the InternetCopy link The call for TransparencyCopy link Introducing Bot Preference SyncCopy link What's next?Questions this post answers
What is Cloudflare Bot Preference Sync and how does it update robots.txt?
Bot Preference Sync is a Cloudflare feature that automatically writes your configured AI bot preferences for Search, Agent, and Training categories into your site's robots.txt file, prepending the generated directives while preserving any existing Disallow rules. It removes the need to manually maintain a static robots.txt, keeping stated preferences and edge-enforced blocks aligned. It rolls out to all plans, from Free to Enterprise, and is on by default for new customers. daily.dev helps site owners track platform changes like this before they affect crawler access.
What extra transparency requirements does Cloudflare impose on mixed-use AI crawlers that do both search and training?
Bots that perform both search and training must respect a no-training preference in robots.txt, let site owners opt out of AI summaries, provide URL-level visibility into which pages were used for training plus search metrics, and publicly demonstrate that disallowing training doesn't hurt search rankings. Crawlers that skip these requirements are blocked outright whenever a site sets Disallow Training, with no benefit of the doubt given. Track evolving crawler transparency rules on daily.dev before they change your site's AI visibility.
What robots.txt default does Cloudflare set for ad-supported publisher sites versus other site owners?
Ad-supported or publisher sites that select the option indicating they monetize pages with ads get Training set to Disallow by default, keeping their content out of AI model training while remaining in search indexes so they still receive referral traffic. Other new customers get no blocks or disallows by default for Search, Agent, or Training when onboarding a domain; the choice is left entirely to the customer. Publishers weighing AI training exposure against search visibility can follow this shift on daily.dev.