Cloudflare now writes a site's AI crawler policy into robots.txt for it
Cloudflare has connected two things that used to be separate. The AI bot categories a site owner sets in the Cloudflare dashboard (Search, Agent, Training) are now mirrored automatically into that site's robots.txt. The generated block is prepended to whatever robots.txt already says, wrapped in comment markers reading "BEGIN Cloudflare Bot Preference Sync" and "END", so existing Disallow lines survive. It is available on every plan from Free to Enterprise. Cloudflare's post says that "for all new customers, Bot Preference Sync will be on by default", and existing customers on the older managed robots.txt feature get prompted to move across.
The Training category also gained a softer option. Instead of a hard block, Disallow writes a no-training preference into robots.txt in a form that lets a crawler doing both search and training keep reaching the content for indexing, provided that crawler meets four disclosure conditions: respect no-training preferences in robots.txt, offer site owners an opt-out from AI summaries, provide URL-level visibility into which pages were used for training alongside search metrics, and demonstrate publicly that disallowing training does not damage traditional search results.