Cloudflare Robots.txt for AI Crawlers: What AEO Teams Should Do
Should you let Cloudflare write your robots.txt for AI crawlers? As a starting point, yes, provided you decide what you want from each type of AI system before you accept what it generates. Search Engine Journal's coverage of Bot Preference Sync makes the sharper point: a setting that works by category cannot describe a policy that treats individual bots differently, and most real policies do.
Key takeaways
- Cloudflare's Bot Preference Sync manages AI crawler policy by category, so it can't express per-bot decisions such as allowing GPTBot while blocking Bytespider.
- Decide your stance on three separate jobs first: search indexing, live answer retrieval, and model training. Then map bots to them.
- A generated robots.txt is a draft. Diff it against your intended policy before you accept it.
- Blocking a crawler can remove you from an answer engine's citations, so treat every block as a visibility decision, not only a legal one.
- Measure with server logs and a fixed prompt set, and change one rule at a time so you can see what moved.
What Cloudflare's Bot Preference Sync does
Search Engine Journal, in a piece by Slobodan Manic, reports that Cloudflare's Bot Preference Sync sets AI crawler policy per category rather than per crawler. You express a preference for a class of bot, and the platform turns that into robots.txt rules and, where you've enabled it, enforcement at the edge.
The convenience is real. Anyone who has maintained a robots.txt by hand knows the file rots. New user agents appear, old ones get renamed, and nobody owns the list. A vendor that tracks the bot landscape and writes the file for you removes a chore that most marketing teams never wanted.
The catch, as Search Engine Journal frames it, is granularity. If your actual decision is to allow one company's crawler and refuse another's, there's no setting that says so. You get the category behaviour, and the exceptions are your problem.
That's a fair trade for some sites and a bad one for others. Which one you are depends on how much your business cares about the difference between bots, and that's the question the rest of this post is about.
Why per-category policy clashes with how you actually decide
Nobody decides about crawlers by category. They decide by counterparty.
Take a B2B software company with a documentation site, a blog and a pricing page. The marketing lead wants the blog cited in ChatGPT and Perplexity answers. Legal doesn't want product documentation used to train models. The pricing page changes weekly, and sales would rather a bot never quote a stale number. Those are three different rules for three different content types, and each involves different bots.
A category setting can only see the bot's type. It can't see that you have a commercial relationship with one AI vendor and none with another, or that one crawler has a reputation for ignoring your crawl-delay while another behaves. It also can't see which section of the site you're talking about, unless the generated rules let you scope them by path.
So the first job isn't to pick a Cloudflare setting. It's to write down your intended policy in plain language, one line per content type and per purpose. Only then can you tell whether a category toggle covers it or whether you need hand-written exceptions on top.
How AI answer engines separate crawling, training and citing
The reason category settings exist at all is that AI companies run more than one kind of bot, and the kinds have different consequences for you. Check each vendor's own documentation for the current list of user agents, because names and behaviour change. As a working model, there are three jobs.
- Training crawlers collect content to build or improve models. Blocking them affects future model knowledge, not whether a live answer can cite you today.
- Search or index crawlers build a retrieval index that an answer engine queries later. Block these and you may vanish from that engine's cited results.
- User-triggered fetchers visit a page because a person asked a question that needs it. These behave more like a browser than a crawler, and vendors differ on whether they honour robots.txt for them.
Google adds its own wrinkle. Its documentation distinguishes the crawler that indexes pages for Search from a separate control token for how content is used in its generative products. Blocking one is not blocking the other, and conflating them is a common and expensive mistake.
This is why the citation side matters. Search Engine Journal's coverage of an AI citation test suggests the order in which sources appear matters less than people assume. Whatever you make of that result, it points the same direction as the crawler split: the gate comes first. A page that a retrieval bot can't fetch never enters the pool that gets ranked, ordered or quoted. Everything else in AEO, from passage structure to entity clarity, only counts once the bot can get in.
One caution. Robots.txt is a request, not a lock. Well-behaved crawlers honour it, and others don't. Enforcement, if you want it, happens through bot management rules at the network edge, which is a separate setting with separate consequences.
What to change this week
You don't need a strategy offsite. You need an afternoon and a shared document.
- Export your current robots.txt and the current Cloudflare bot settings. Save both with today's date so you have a rollback.
- Write the intended policy as a table with one row per bot type and content area. The table below is a template.
- Generate the Cloudflare version and diff it against your intent. Look for lines that block a search or retrieval bot you wanted to keep.
- Add explicit exceptions for anything the category setting can't express. Order matters in robots.txt, and more specific rules for a named user agent generally take precedence, so test the result rather than trusting it.
- Check enforcement. Confirm whether the edge is actively blocking bots or only advertising a preference in the file.
- Set an owner and a review date. A policy nobody reviews goes stale in a quarter.
Here's the template, filled in for the software company above.
| Bot job | Content area | Stance | Reasoning | Owner |
|---|---|---|---|---|
| Search and retrieval | Blog, guides | Allow | We want citations in answer engines | Growth lead |
| Search and retrieval | Pricing, plans | Allow, but keep a dated, plain-text version current | Bots will quote whatever they find | Product marketing |
| Training | Blog | Decide explicitly | Visibility upside is unclear; this is a brand and legal call | Legal and marketing |
| Training | Documentation | Block | Contractual and competitive sensitivity | Legal |
| User-triggered fetch | All public pages | Allow | A person asked for it; blocking loses a real reader | Growth lead |
| Unidentified or abusive bots | All | Block at the edge | Robots.txt won't stop them anyway | Security |
The useful thing about this table isn't the specific answers. It's that every cell has a named person, which turns an infrastructure toggle into a decision someone can defend. Growth teams that already run a regular content review can fold this into it; the growth teams solutions page covers how that kind of cross-functional ownership usually gets set up.
The worked example: blocking the wrong bot
A mid-size retailer switches on a broad AI-bot block in a single afternoon because the category setting makes it one click. Nobody reads the resulting file. Six weeks later the content team notices that their buying guides no longer show up when they test questions in an answer engine that used to cite them.
The cause isn't mysterious. The block covered the retrieval crawler along with the training crawlers. The team wanted to opt out of training, and they opted out of visibility as well.
The fix is boring: narrow the rule, wait for the engine to recrawl, and re-test. But there's a lesson underneath. The retailer had no baseline. Without a record of what was cited before the change, they couldn't say when things broke, and they spent two weeks arguing about whether it was a content problem.
This is what a category setting hides. The cost of a coarse rule isn't the rule itself. It's the difficulty of telling afterwards which part of it did the damage.
How to measure whether the policy is working
You need two data sources, and they answer different questions.
Server or edge logs tell you which bots are actually visiting, how often, and which paths they hit. Filter by user agent and group by the job categories above. You're looking for three things: that the bots you allowed are arriving, that the ones you blocked have stopped, and that nothing important is returning errors to bots you wanted.
A fixed prompt set tells you what the answer engines do with the access they have. Pick 30 to 50 questions your buyers really ask. Run them on a schedule in each engine you care about, and record whether your domain is cited, which page, and what the answer says about you. Keep the wording identical from run to run, since small rephrasing changes results.
The two together give you a cause-and-effect view. Logs show the door is open. The prompt set shows whether anyone walks through it. Tools such as MediaPilot are aimed at tracking media and AI-search visibility, which is the second half of that pairing; whatever you use, the discipline is the same.
A simple plan:
| What to track | Source | Cadence | Signal to act on |
|---|---|---|---|
| Bot visits by user agent and path | Edge or server logs | Weekly | An allowed retrieval bot goes quiet |
| Cited domain and page per prompt | Fixed prompt set | Fortnightly | A page drops out after a policy change |
| Error rates served to bots | Logs | Weekly | Spikes in 403 or 5xx to bots you allow |
| Policy versus reality | Diff of robots.txt and edge rules | After every change | Any mismatch |
Change one rule at a time. If you flip three settings on Monday and citations move on Friday, you've learned nothing.
What not to do
A few habits are worth dropping, and some anxiety is worth setting down.
- Don't accept a one-click block as a legal strategy. If your concern is training use, a robots.txt line is a signal and not a licence term. Your counsel should decide what protects you.
- Don't block everything by reflex. Being absent from answer engines has a cost that's easy to miss because nothing visibly breaks. You just stop appearing.
- Don't chase every new user agent. Lists of AI bots grow constantly. Maintain your policy at the level of jobs and named counterparties, and let the platform handle the long tail.
- Don't assume the file is the whole story. A robots.txt and an edge rule can disagree. Whichever is stricter tends to win in practice, and you should know which one that is.
- Don't treat this as a one-off. Vendors change crawler names and behaviour. A quarterly review beats a heroic annual one.
There's also a broader shift worth watching without overreacting to. Search Engine Journal has separately reported on an experiment in charging AI agents per page. It's one person's test, not a market standard, and nothing suggests it's about to become normal. But it points at where access policy may head: away from a yes-or-no gate and towards conditions. If that happens, teams that already have a written policy per bot will adapt faster than teams that inherited whatever a toggle produced.
One more thing to ignore: any claim that a specific robots.txt tweak will lift your citations. Access is a precondition, not a ranking factor. Once bots can reach you, what gets cited depends on the content, its clarity, and how well it answers the question. If your page structure and entity data need work, the platform overview shows how the pieces of a marketing operation fit together, and the plumbing described here is only the first of them.
Frequently asked questions
Does blocking AI crawlers in robots.txt remove me from AI answers?
It can, if you block the retrieval or search crawler that an answer engine uses to find pages. Blocking only training crawlers is a different decision with different effects. Check each vendor's documentation to see which user agent does which job before you write the rule.
Is Cloudflare's generated robots.txt safe to publish as is?
Treat it as a draft. Compare it with your written policy, and look specifically for rules that cover a bot you wanted to allow. Publish only after that diff, and keep the previous file so you can roll back.
Do all AI bots respect robots.txt?
No. Reputable vendors say they do, and user-triggered fetchers are handled differently by different companies. If you need enforcement rather than a request, that's an edge or firewall rule, and it needs its own testing.
How long until a policy change shows up in AI answers?
There's no reliable number, and it varies by engine and by how often each one recrawls your site. Make one change at a time, keep your prompt set fixed, and compare over several weeks rather than several days.
Should I allow AI training crawlers?
That's a business and legal decision, not a technical one, and the visibility benefit is unclear. Decide it explicitly per content area, with legal in the room, instead of letting a default make the call for you.
Sources
- https://www.searchenginejournal.com/cloudflare-will-write-your-robots-txt-and-it-has-a-point/589262/
- https://www.searchenginejournal.com/i-made-my-website-charge-ai-agents-a-penny-per-page-then-i-watched-claude-pay-it/589447/
- https://www.searchenginejournal.com/ai-citation-test-finds-source-order-matters-less-than-it-looks/589806/