How to Block AI Crawlers in WordPress: GPTBot, ClaudeBot, Google-Extended and More (2026)

How to Block AI Crawlers in WordPress: GPTBot, ClaudeBot, Google-Extended and More (2026)

Which AI crawlers collect training data and which power AI search, how to block them in WordPress with robots.txt or a firewall, and what llms.txt really does.

AI companies now crawl the web for two different purposes: to collect training data for their models, and to fetch pages live when a user asks a question. Many WordPress site owners want to block the first but stay visible in AI search answers. That is possible, but only if you know which crawler does what.

This guide lists the main AI crawlers in 2026, shows how to block them in WordPress, and is honest about what robots.txt can and cannot do.

The Short Answer

To stop major AI companies from using your content for model training, add Disallow: / rules in robots.txt for their training crawlers or tokens: GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (Google Gemini), Applebot-Extended (Apple) and CCBot (Common Crawl). Leave their search crawlers, such as OAI-SearchBot and Claude-SearchBot, allowed if you want to appear in AI search results.

In WordPress, edit robots.txt through Yoast SEO (Tools > File editor) or a physical file in your site root. Remember that robots.txt is a request, not a lock: reputable crawlers honour it, user-triggered fetches may not, and scrapers ignore it. For enforcement, use a firewall rule, such as Cloudflare’s AI crawler blocking.

The Main AI Crawlers and What They Do

Robots.txt tokenCompanyPurposeBlock if you want to…
GPTBotOpenAITraining dataOpt out of OpenAI model training
OAI-SearchBotOpenAIChatGPT search resultsLeave out of ChatGPT search (not recommended for visibility)
ChatGPT-UserOpenAIFetches a page when a user asksRobots.txt may not apply
ClaudeBotAnthropicTraining dataOpt out of Anthropic model training
Claude-SearchBotAnthropicSearch result qualityLeave out of Claude search
Claude-UserAnthropicFetches a page when a user asksStop user-directed fetches
Google-ExtendedGoogleGemini training and grounding (not Search)Opt out of Gemini use without affecting Google Search
Applebot-ExtendedAppleApple foundation model trainingOpt out of Apple model training
PerplexityBotPerplexityPerplexity search resultsLeave out of Perplexity search
CCBotCommon CrawlOpen web archive widely used for AI trainingStay out of Common Crawl datasets
Sources: OpenAI, Anthropic, Google, Apple, Perplexity, Common Crawl.

Two details matter. First, Google-Extended and Applebot-Extended are control tokens, not separate crawlers: Google says Google-Extended does not affect Google Search, including AI Overviews, which follow your Googlebot rules. Second, OpenAI and Perplexity state that their user-triggered agents (ChatGPT-User and Perplexity-User) may not follow robots.txt, because a person requested the page.

Diagram grouping AI crawlers into model training crawlers to block (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot), AI search crawlers to allow (OAI-SearchBot, Claude-SearchBot, PerplexityBot) and user-triggered fetchers, with a note never to block Googlebot
Training, search and user-triggered AI crawlers at a glance. Diagram by Jackober, based on crawler documentation from OpenAI, Anthropic and Google.
If your priority is…Then…
Keeping your content out of model trainingBlock the training tokens (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot)
Getting cited in AI search toolsAllow OAI-SearchBot, Claude-SearchBot and PerplexityBot
Google Search and AI Overviews trafficNever block Googlebot; blocking Google-Extended is safe for Search
Reducing server load from aggressive botsUse rate limiting or a firewall, not only robots.txt
Protecting paid or members-only contentPut it behind login; robots.txt does not protect private content

There is no universally right answer. Yoast’s own discussion of blocking AI bots lays out the trade-off between protecting content and being visible where people now search.

How to Block AI Crawlers in WordPress

1. Check your current robots.txt

Open https://yourdomain.com/robots.txt. WordPress serves a virtual file by default; if a physical robots.txt exists in the site root, it takes priority. Note what is already there so you do not overwrite existing rules.

2. Add the rules

This example blocks training crawlers while leaving search crawlers and Googlebot untouched:

# Opt out of AI model training
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: CCBot
Disallow: /

Each group needs its own User-agent line. Google’s robots.txt specification explains how crawlers pick the most specific matching group, so a named group overrides your general User-agent: * rules for that crawler.

3. Choose where to edit

  • Yoast SEO: go to Yoast SEO > Tools > File editor and paste the rules. Yoast SEO Premium also has toggles under Settings > Advanced > Crawl optimization to block unwanted bots such as GPTBot, CCBot and Google-Extended.
  • A physical file: upload robots.txt to the site root by SFTP or your hosting file manager.

4. Enforce it at the edge (optional)

If you need actual blocking, not a request, use your CDN or firewall. Cloudflare offers one-click AI bot blocking, including on its free plan, and AI Crawl Control to allow or block individual crawlers. Our guide to setting up a CDN for WordPress covers getting a site behind one.

5. Verify

Reload your robots.txt URL to confirm the rules are live, and clear any page cache if the old version still appears. OpenAI notes it can take about 24 hours for its systems to pick up changes. Check your server logs later for the user agents you blocked.

What About llms.txt?

llms.txt is a proposed Markdown file, published by Answer.AI in September 2024, that offers language models a curated map of a site. It is not a blocking mechanism and not a ratified standard. Google states in its guidance on AI features and your website that you do not need special AI files to appear in Google Search, including its AI features. Adding one is harmless, but do not expect it to change rankings or block anything.

Common Mistakes

  • Blocking Googlebot to stop AI Overviews. It removes you from Google Search entirely. Google-Extended does not control AI Overviews either; snippet controls such as nosnippet are the documented options.
  • Blocking every bot with a “User-agent: *” disallow. This also blocks search engines and can deindex the site.
  • Assuming robots.txt removes content already collected. It only applies to future crawling.
  • Blocking search crawlers by accident. Disallowing OAI-SearchBot or PerplexityBot removes you from those search tools, not just from training.
  • Forgetting the cache. Caching plugins and CDNs can serve an old robots.txt for a while.

For how search crawlers treat WordPress sites more broadly, see our guide to WordPress indexing and crawl behavior and the complete WordPress SEO guide.

FAQ

Does blocking GPTBot remove my site from ChatGPT?

No. GPTBot is OpenAI’s training crawler. ChatGPT search uses OAI-SearchBot, and OpenAI treats the two settings independently, so you can block one and allow the other.

Will blocking Google-Extended hurt my Google rankings?

According to Google’s documentation, no. Google-Extended only controls use in Gemini models and Gemini Apps grounding; Google Search, including AI Overviews, is governed by Googlebot.

Do all AI companies respect robots.txt?

The companies listed above document robots.txt support for their crawlers, with exceptions for user-triggered fetches. Unknown scrapers may ignore it entirely, which is why firewall rules exist.

Can I slow AI crawlers down instead of blocking them?

Some support it. Anthropic, for example, honours the non-standard Crawl-delay directive for ClaudeBot. Google does not support Crawl-delay, so rate limiting at the server or CDN is more dependable.

Is blocking AI crawlers a security measure?

No. Robots.txt is public and voluntary. Protect private areas with authentication and follow standard WordPress security practices.

The Bottom Line

Separate training crawlers from search crawlers. Block GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot if you want to opt out of AI training; keep Googlebot and AI search crawlers allowed if you want the traffic. Edit robots.txt through Yoast or a root file, and add a firewall rule when you need enforcement rather than a polite request.

Sources & Further Reading

Jackober uses AI tools for research, drafting, and editing. Articles are editorially reviewed and factual claims are checked against cited sources. Last reviewed: October 2026.

Leave a Reply

Your email address will not be published. Required fields are marked *

You May Also Like