What changed on September 15
Cloudflare’s announcement sorts AI crawler traffic into three categories, search, agent and training, and sets new defaults: for domains onboarding from 15 September 2026, training and agent crawlers are blocked by default on pages that display ads, while search crawlers remain allowed. The controls behind it are available on every plan, including Free.
The ad-pages condition is the interesting design choice. It targets exactly the sites whose business model is visits, on the logic that a bot which reads the page and answers the user elsewhere consumes the content and starves the model that paid for it.
Existing zones keep their configured behavior. Which sounds reassuring and is actually the reason to go look: what your zone does today is a mix of decisions made at different times under different defaults, and the September change is a good excuse to find out what they add up to.
Search, agent and training are three different deals
The category split matters more than the vendor names, because each category is a different trade:
- Search crawlers index your pages so AI search products can cite and link them. OpenAI’s OAI-SearchBot and Anthropic’s Claude-SearchBot live here. Blocking them removes you from cited answers, the closest thing AI products have to sending you traffic.
- Agent crawlers fetch pages on a user’s behalf in the moment, ChatGPT-User and Claude-User style, when someone asks an assistant to read a page or complete a task on it. Whether that is a welcome visitor or a scraper with better manners is genuinely context-dependent.
- Training crawlers collect text for future model training, GPTBot, ClaudeBot and friends. Nothing comes back on any timescale you can measure, and what they take persists in models for years.
Most of the loud “block all AI” advice from the last two years predates this split and treats all three as one. The defaults Cloudflare picked, search yes, training no, are a defensible middle position, but they are still a default, not your decision.
The crawl-to-refer math
Cloudflare publishes the ratio that makes the whole conflict legible: pages crawled per visitor referred back. On Radar’s July 2026 numbers, Google search sat near 5 to 1, five pages fetched per visit sent. OpenAI stood at roughly 250 to 1, Perplexity near 290 to 1, Anthropic near 1,900 to 1 and Mistral around 3,400 to 1.
The trend inside those numbers is as telling as their size. A year earlier Anthropic’s ratio had been measured in the tens of thousands to one. Then Claude’s web search launched with clickable citations, referrals existed at all, and the ratio collapsed by orders of magnitude. The economics improve exactly when the products start linking out, which is worth remembering when deciding whether to block the search category: cited answers are the part of this ecosystem that pays anything back.
Checking what your zone actually does
Fifteen minutes, three places:
- The AI crawler controls in your Cloudflare dashboard show the current allow-and-block state per category and per bot, and the bot analytics show who has been hitting you, how often, and what happened to them.
- Your robots.txt, because rules accumulated there over the years may say something entirely different from what the network layer enforces, and the mismatch confuses both audits and crawlers.
- A live check from outside: fetch a page with a blocked bot’s user agent and confirm you get the block you expect, and fetch with an allowed one and confirm you don’t.
Write down what you found before changing anything. The point of the exercise is that allow and block are now per-category decisions, and the record of what you decided is what saves the next person from guessing.
What blocking costs, what it protects
The honest framing is a trade with unknowns on both sides. Blocking training bots protects content from uncompensated model-building, at the cost of absence from whatever visibility future models provide. Blocking search bots is more immediate: AI search products stop citing you, and for sites whose audience increasingly asks assistants instead of search engines, that is a real channel dying quietly.
For this site we keep the search category open, and the price of that choice is plain: answer engines get content they sometimes answer with directly, no click involved. Nothing in our analytics can separate the visitors a citation brought from the visitors a paraphrased answer replaced, and anyone claiming certainty on that balance is selling something.
Pay-per-crawl, in one breath
Cloudflare’s marketplace lets a site answer a crawler with a price instead of a yes or no: pay the configured rate per fetch or receive HTTP 402, the status code that waited thirty years for a job. Whether meaningful money flows through it yet is unclear from outside, but it reframes the question usefully, from moral argument to price discovery.
Where robots.txt still fits
Network enforcement did not retire robots.txt, it clarified the file’s role: robots.txt is the published policy, readable by every crawler and every auditor, and the network layer is the lock on the door. Keeping the two consistent is the actual maintenance task, and a common failure is a blanket disallow written in 2024 that now contradicts the search-allowed stance configured upstream.
User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Disallow: / # training and citations both gone
User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Allow: / # no training, keep the citations
Our robots.txt generator knows the current AI crawler names by category, and the robots.txt tester answers the question that actually matters after any edit: which rule wins for which bot on which URL, before a crawler finds out for you.
AI crawler control questions
Does Cloudflare block AI crawlers by default?
For domains onboarded since 15 September 2026, partly yes: training and agent crawlers are blocked by default on pages that display ads, while search crawlers stay allowed. Existing zones keep whatever they had configured, and the controls are available on every plan including Free.
How do I see which AI bots are crawling my site?
Cloudflare’s bot analytics list them per zone by name and category, and without Cloudflare your server logs do: the relevant crawlers identify themselves in the user agent as GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and so on. Cloudflare Radar also publishes ecosystem-wide crawler traffic if you want the big picture.
What is the difference between GPTBot and OAI-SearchBot?
GPTBot collects training data for OpenAI’s models. OAI-SearchBot indexes pages so ChatGPT search can cite and link them. Blocking the first keeps you out of future training runs, blocking the second removes you from answers that would have linked to you, which is why the two deserve different decisions.
What is pay-per-crawl?
A Cloudflare marketplace where a site sets a per-page price and participating AI crawlers either pay it or receive HTTP 402.
Will blocking AI crawlers remove my site from ChatGPT answers?
From answers based on live search, yes, over time: if OAI-SearchBot cannot fetch you, ChatGPT cannot cite you. What models memorized in past training runs is already baked in and does not disappear by blocking anything now.
Does blocking AI bots hurt Google rankings?
No. Classic Googlebot is a search crawler and stays untouched by AI-bot rules, and Google’s AI training opt-out runs through the separate Google-Extended token.
What is a crawl-to-refer ratio?
How many pages a company’s bots fetch per visitor its products send back. Cloudflare Radar computed it at around 5:1 for Google search but in the hundreds or thousands to one for AI companies.