AI crawlers are the bots that read your website on behalf of Google, ChatGPT, Claude, and Perplexity — and in the last ninety days, the rules governing them changed more than they had in the previous decade. Cloudflare announced it will block an entire class of AI crawlers by default. UK regulators forced Google to build an opt-out for its AI answers. A $1.5 billion copyright settlement got final approval. And some of the biggest publishers in the country started openly debating whether to shut Google out entirely.
At The Digital Hall, we read all of that the same way: the open web is quietly becoming a permissioned web. That is not a reason to panic. It is a reason to get deliberate about which AI crawlers you let in, what they find when they arrive, and whether your content is actually worth citing once they get there.
Here is what really happened this summer, what it means for your SEO, AEO, and GEO strategy, and the five moves we are making for clients right now.
What is inside this AI crawler update
- What are AI crawlers, and why they suddenly matter
- Cloudflare blocks mixed-use AI crawlers on September 15
- Google has to let publishers opt out of AI answers
- Google’s AI answers are sending more clicks than they did in spring
- Licensing is quietly replacing scraping
- Publishers are threatening to walk away from Google
- What this means if you are not a publisher
- Five smart moves to make before September 15
- AI crawler FAQs
What are AI crawlers, and why do they suddenly matter?
AI crawlers are automated bots that fetch web pages for artificial intelligence systems. They do three very different jobs: indexing pages for a search results list, gathering text to train a large language model, and retrieving live sources so an assistant can answer a question right now. GPTBot, ClaudeBot, PerplexityBot, Google-Extended, and OAI-SearchBot are all AI crawlers, and they do not all want the same thing from your site.
For years, those three jobs traveled together under a single user agent, which meant one decision — allow or block — covered all of them. Say yes and you fed the training data. Say no and you disappeared from the answer. That all-or-nothing bargain is what 2026 is dismantling, and it is why AI crawlers moved from a server-log curiosity to a boardroom conversation.
If you have been following how AI reads your website, this is the next chapter. The reading part is settled. The permission part is not.
1. Cloudflare will block mixed-use AI crawlers by default on September 15
This is the item with a deadline on it, so we are putting it first.
On July 1, 2026, TechCrunch reported that Cloudflare will change its default settings to block “mixed-use” AI crawlers from ad-supported pages beginning September 15, 2026. A mixed-use crawler is one bot doing search indexing, AI agent work, and model training all at once. Cloudflare wants those purposes split into separate, declared bots so site owners can say yes to one and no to another.
Two numbers from that reporting stuck with us. Bots now outnumber humans on the internet. And more than half of AI crawler traffic is spent re-fetching pages that have not changed. CEO Matthew Prince framed the urgency plainly: “we must go further and act faster.”
Why this one matters to you specifically
The new default applies to new Cloudflare customers, new sites created by existing customers, and every existing free-tier customer. An enormous number of small business websites sit behind Cloudflare’s free tier — often because a developer set it up years ago and nobody thought about it again. If that is you, and your site runs ads or affiliate content, AI crawlers you currently welcome could be turned away automatically next month. You would not get an email about it. You would just slowly stop showing up in AI answers.
Go look at your bot settings before September 15. That is genuinely the whole assignment.
2. Google now has to let publishers opt out of AI answers
On June 3, 2026, the UK’s Competition and Markets Authority secured what it called a world first: Google must give site owners a way to keep their content out of generative AI search features. TechCrunch reported the control arrives as a toggle inside Google Search Console, covering AI Overviews, AI Mode, and AI Overviews in Discover.
The critical detail is the one nobody expected Google to concede: opting out will not be used as a ranking signal for traditional search. You can leave the AI summary without leaving the results page. Testing starts with UK publishers before a global rollout, so expect that toggle to land in your Search Console eventually.
Here is our honest read, and it may not be the popular one. Most small and mid-sized businesses should not touch that toggle. Publishers opt out because their revenue depends on the pageview. Your revenue depends on the phone ringing. When Google’s AI names your practice, your firm, or your shop as the answer, that summary is doing sales work a blue link never did. Opting out of it means opting out of the Google Zero era entirely, and that era is not going to wait for you.
3. Google’s AI answers are sending more clicks than they did in the spring
This is the update almost nobody in our industry is talking about, and it is the most encouraging news of the summer.
TechCrunch, reporting Similarweb data on July 27, 2026, put hard numbers on the shift:
- AI Overviews now appear in 43% of Google searches, up from 15% a year earlier.
- AI Mode visits grew from 126 million in June 2025 to 279 million by May 2026.
- Citations inside AI responses increased more than fivefold in a year.
- After a May 7 Google update, the share of AI-answer sessions that reached an actual webpage climbed from about 25% in March to nearly 60% by May 30.
- Meanwhile, only 6.8% of US desktop ChatGPT queries included citations as of May 2026.
Read those last two together, because that is the strategy. The click is not dead — it became conditional. Google is putting links back into its AI answers, which means citation-worthy content is recovering real value after a brutal 2025. ChatGPT, on the other hand, still cites almost nothing. Betting your visibility on one engine has never looked worse, which is exactly why we treat SEO, AEO, and GEO as one program instead of three. If AI Overviews have been eating your traffic, the second half of 2026 is your window to win some of it back.
4. Licensing is quietly replacing scraping
On July 21, 2026, a federal judge granted final approval to Anthropic’s $1.5 billion copyright settlement — roughly $3,000 each across about 500,000 works, and the largest copyright settlement on record in the United States. As TechCrunch explained, the case turned on how the books were acquired — downloaded from pirate libraries — not on whether training a model is fair use.
A month earlier, Getty Images signed a display partnership with OpenAI, and Forbes noted that licensed images now surface inside ChatGPT answers — while pointing out the deal does not hand users commercial rights to those images.
Two lessons for the rest of us. First, content with clean ownership, clear authorship, and traceable provenance is becoming a licensable asset rather than free fuel. Second, please do not assume an image an AI assistant shows you is yours to drop into a campaign. We wrote about the widening legal edges of this in our piece on AI agent lawsuits, and the same principle runs through our position on AI ethics in digital marketing: use it, but use it honestly.
5. Publishers are threatening to walk away from Google
On July 22, 2026, Nieman Lab summarized Wall Street Journal reporting that several major publications lost more than 40% of their search traffic between June 2025 and June 2026 — while The Guardian and the BBC actually gained.
The quotes are what make it real. USA Today Co. CEO Mike Reed: “It’s time to take a stand and say enough is enough.” People Inc. CEO Neil Vogel said blocking Google entirely is “100% on the table.” And Cloudflare’s Stephanie Cohen described the goal that ties this whole summer together: to be “discoverable without having to give your content away.”
That is the sentence we keep coming back to. Discoverable without giving it all away. That is the entire fight over AI crawlers in nine words.
What the new AI crawler rules mean if you are not a publisher
Most of the coverage you will read about AI crawlers is written by publishers, for publishers, about a business model that is not yours. A media company sells attention, so a summary that satisfies the reader is a lost sale. You sell a service, a product, an appointment, a consultation. For you, a summary that names your business is a warm introduction.
So the math flips. Publishers are negotiating for leverage. You should be negotiating for presence. Blocking AI crawlers to protect 900 words of blog content is a trade almost no small business should make — and we say that as people who have watched a Richmond, VA practice get named in an AI answer and book new patients off it the same week.
What you should protect is different: your pricing pages from bad scrapes, your proprietary research from being repackaged, and your bandwidth from bots re-fetching the same unchanged page fifty times a day. That is a settings conversation, not a strategy retreat.
Five smart moves to make before September 15
1. Find out which AI crawlers are actually visiting you
Open your server logs or your Cloudflare bot analytics and look for GPTBot, OAI-SearchBot, ClaudeBot, Google-Extended, PerplexityBot, and Bingbot. You cannot make a good decision about AI crawlers you have never looked at. Most business owners we walk through this are genuinely surprised by what is knocking.
2. Set your crawler policy on purpose, not by default
Your robots.txt should reflect a decision you made, not one a plugin made in 2021. Pair it with an llms.txt file so assistants get a clean map of what you want quoted. We broke the whole thing down in llms.txt explained — it takes about twenty minutes and it is the cheapest AI visibility work available right now.
3. Make your pages machine-readable
Engines cannot quote what they cannot parse. Organization, LocalBusiness, Article, FAQ, and Product schema tell AI crawlers what your entity is and what you are claiming, in a format that leaves no room for guessing. Start with schema engineering and get your entity clarity right before you write another word of content.
4. Write answers, not just articles
Lead with the direct answer, then support it. Define your terms in one sentence. Use real questions as headings. Keep paragraphs tight enough to lift. That structure is what gets a page pulled into a summary, and it is the backbone of our AEO best practices and our full SEO, AEO, and GEO playbook.
5. Measure citations, not just clicks
If the click became conditional, your dashboard needs to catch what happens before the click. We built SERPfinity, our own AI search visibility platform, precisely because no legacy rank tracker could tell a client whether ChatGPT, Google AI Overviews, or Perplexity was naming them. Pair that with the thinking in our post on the AI attribution gap and you stop flying blind.
Where The Digital Hall stands on AI crawlers
We have never been interested in tricking a crawler, human or artificial. White Hat is the first letter of The WRRAP Around Method™ for a reason, and that principle holds whether the reader is a person, Googlebot, or one of the newer AI crawlers reading your site on someone’s behalf at two in the morning.
Our founder, MonicaFaye Hall, says it more bluntly than we do: if your website doesn’t make SEO noise, no one will ever see it. Twenty years in, that has not changed — only the audience has. The businesses winning right now are the ones treating AI crawlers as guests they prepared for, not intruders they are hiding from.
And we still believe small and mid-sized businesses deserve the same caliber of strategy that enterprise brands have always had. That belief is why we publish this instead of saving it for a sales call. If you want to see where your business scores today, our AI Visibility Index lays out the six pillars we measure against.
Free download: the AI SEO Checklist
If you would rather work from a list than a lecture, grab our AI SEO Checklist (free PDF). It walks through the content, technical, and analytics work that makes a site worth citing — which is the same work that makes AI crawlers useful to you instead of expensive to you. No form, no gate. Just download it.
AI crawler FAQs
What are AI crawlers?
AI crawlers are automated bots that fetch web pages for artificial intelligence systems, either to train a model, to index content, or to retrieve live sources for an assistant’s answer. GPTBot, ClaudeBot, PerplexityBot, and Google-Extended are common examples.
Should I block AI crawlers on my website?
For most small and mid-sized businesses, no. Blocking AI crawlers removes you from the AI answers where customers are increasingly making decisions. Blocking makes more sense for publishers whose revenue depends on pageviews, or for sites protecting proprietary research and pricing data.
Will blocking AI crawlers hurt my Google rankings?
Under the UK CMA agreement announced June 3, 2026, opting out of Google’s generative AI features will not be used as a ranking signal for traditional search. Blocking Googlebot itself is a different matter and will remove you from search results entirely.
What is the difference between AI crawlers and search crawlers?
A search crawler indexes pages so they can be ranked and linked. AI crawlers gather content so a model can learn from it or summarize it inside an answer. Cloudflare’s September 2026 policy change specifically targets bots that blur those purposes together.
How do I know if AI crawlers are visiting my site?
Check your server access logs or your CDN’s bot analytics for known user agents such as GPTBot, ClaudeBot, OAI-SearchBot, PerplexityBot, and Google-Extended. If you use Cloudflare, its bot dashboard reports AI crawler activity directly.
Does an llms.txt file stop AI crawlers?
No. An llms.txt file does not block anything, it guides. It gives AI crawlers a curated map of your most important, most quotable pages so assistants pull from the content you actually want represented.
The bottom line on AI crawlers in 2026
The summer of 2026 handed site owners something we have been asking for since AI answers arrived: choices. You can decide which AI crawlers get in. You can stay in search without being summarized. You can license instead of donate. What you cannot do is ignore the settings and hope the defaults are on your side, because on September 15 some of those defaults change without asking you.
Check your bot settings. Publish your llms.txt. Fix your schema. Write pages worth quoting. Then measure whether the engines are actually naming you.
If you would rather not do that alone, that is what we are here for. Book a complimentary consult and we will walk your site’s AI crawler setup, entity clarity, and citation footprint with you — or explore our AEO services and GEO marketing services to see how we run the full program for Richmond, VA businesses and national brands alike.
Don’t just rank. Become the answer.