SEO Pulse Feature

See which AI crawlers
can access your site

A visual grid checks your robots.txt against 20 major crawlers — search engines, AI trainers, live browsing agents — and shows ✓ or ✗ for the current page URL.

20 bots. One grid.
Current page URL.

SEO Pulse reads your robots.txt and checks it against 20 known crawler user-agents — for the specific URL you're currently viewing, not just the root domain. A ✓ means the current URL is accessible to that bot. A ✗ means it's blocked.

Each bot card has a ⓘ tooltip showing the bot's owner, its type (search crawler, AI trainer, live browsing agent, AI training opt-out), and a plain-English explanation of what blocking it actually means for your site.

This matters because the distinctions are often misunderstood. Blocking Google-Extended doesn't affect Google Search. Blocking GPTBot doesn't stop ChatGPT from browsing your site. Blocking PerplexityBot doesn't protect your content from AI training. SEO Pulse explains each one so you can make informed decisions.

✓ Allow
Googlebot
Google · Search
✓ Allow
Bingbot
Microsoft · Search
✓ Allow
DuckDuckBot
DuckDuckGo · Search
✗ Block
GPTBot
OpenAI · AI Training
✗ Block
ClaudeBot
Anthropic · AI Training
✓ Allow
OAI-SearchBot
OpenAI · AI Search
✓ Allow
PerplexityBot
Perplexity · AI Search
✗ Block
CCBot
Common Crawl · Training
✗ Block
Bytespider
ByteDance · Training
✓ Allow
Claude-SearchBot
Anthropic · AI Search
✗ Block
Google-Extended
Google · Gemini Training
✓ Allow
Applebot
Apple · Search

Every crawler, explained

SEO Pulse checks all of these against your robots.txt for the current page URL.

User-agent Owner Type What blocking means
GooglebotGoogleSearch crawlerPrevents Google from crawling your pages. Doesn't remove already-indexed pages but prevents updates.
Google-ExtendedGoogleAI training opt-outOpts your content out of Gemini AI training. Does NOT affect Google Search indexing — Googlebot handles that separately.
AdsBot-GoogleGoogleAd qualityChecks pages linked from Google Ads. Blocking may affect ad quality scores.
BingbotMicrosoftSearch crawlerPrevents Bing from crawling your pages. Also powers Microsoft Copilot answers.
GPTBotOpenAIAI trainingStops OpenAI collecting your content for model training. Does NOT affect ChatGPT live browsing.
ChatGPT-UserOpenAILive browsingFetches pages when a ChatGPT user asks it to browse the web. Not used for training.
OAI-SearchBotOpenAIAI searchIndexes pages for ChatGPT's search feature. Blocking removes your content from ChatGPT search.
ClaudeBotAnthropicAI trainingCrawls for Claude AI training data. Blocking signals exclusion from future Claude training datasets.
Claude-UserAnthropicLive browsingFetches pages for live Claude user queries. Not used for training.
Claude-SearchBotAnthropicAI searchIndexes pages for Claude search answers. Blocking may reduce your site's visibility in Claude.
PerplexityBotPerplexityAI searchCrawls for Perplexity search. Not a training crawler — Perplexity uses OpenAI/Meta models.
ApplebotAppleSearch crawlerCrawls for Siri, Spotlight, and Safari suggestions.
Applebot-ExtendedAppleAI training opt-outOpts your content out of Apple Intelligence training. Does NOT affect Applebot search crawling.
meta-externalagentMetaAI trainingMeta's primary crawler for AI model training and search across Meta's products.
FacebookBotMetaLink previewFetches metadata for Facebook and Instagram link previews. Primarily for previews, not AI training.
CCBotCommon CrawlAI trainingNon-profit web archive widely used to train language models — including many smaller AI labs.
BytespiderByteDanceAI trainingByteDance crawler for TikTok's AI systems. No official documentation. Widely reported to not consistently respect robots.txt.
AmazonbotAmazonAI crawlerUsed to improve Alexa and Amazon AI features. Respects robots.txt.
DuckDuckBotDuckDuckGoSearch crawlerCrawls pages for DuckDuckGo's search index.
YandexBotYandexSearch crawlerCrawls pages for Yandex Search. Relevant if you have Russian-language traffic.

Distinctions that matter

Most robots.txt guides treat all AI bots the same. The reality is more nuanced — and the differences affect what you're actually protecting.

GPTBot ≠ ChatGPT browsing
GPTBot collects training data for OpenAI's language models. ChatGPT-User is what fetches pages when a ChatGPT user asks it to browse the web. Blocking GPTBot stops training data collection. It does not stop ChatGPT from reading your pages on a user's request. Both are separate user-agents with separate functions.
Google-Extended doesn't affect search
Blocking Google-Extended opts your content out of Gemini AI model training only. Googlebot handles Search indexing completely independently. Many sites block Google-Extended as a training opt-out without any impact on their Google Search rankings or visibility.
PerplexityBot is not a training crawler
Perplexity does not train its own AI models — it uses foundation models from OpenAI and Meta. PerplexityBot indexes pages for Perplexity's search answers. Blocking it removes your content from Perplexity results. It does not protect your content from being used in OpenAI or Meta model training.
Applebot-Extended is an opt-out, not a block
Like Google-Extended, Applebot-Extended controls whether your content is used to train Apple Intelligence models. Blocking it has no effect on Applebot's standard search crawling for Siri and Spotlight. It's a separate user-agent specifically for the training opt-out signal.

Common questions

What is an AI bot access checker? +
An AI bot access checker reads your robots.txt and tests it against known AI crawler user-agents to show which ones are allowed or blocked for a specific URL. SEO Pulse checks 20 crawlers — including search bots, AI training crawlers, live browsing agents, and AI training opt-out signals — and shows results as a visual grid for the current page you're viewing.
Does blocking GPTBot stop ChatGPT from browsing my site? +
No. GPTBot and ChatGPT-User are separate user-agents. GPTBot collects training data for OpenAI's AI models. ChatGPT-User is what fetches pages when a ChatGPT user asks it to browse the web in real time. Blocking GPTBot does not affect ChatGPT live browsing — that requires blocking ChatGPT-User separately.
Does blocking Google-Extended affect my Google Search rankings? +
No. Google-Extended controls whether your content is used to train Google's Gemini AI models only. Googlebot handles Google Search crawling and indexing completely separately. Blocking Google-Extended will not affect your Google Search rankings, crawl budget, or indexing in any way.
Is PerplexityBot an AI training crawler? +
No. Perplexity does not train its own AI models — it uses foundation models from OpenAI and Meta. PerplexityBot crawls and indexes pages for Perplexity's AI search results. Blocking it will remove your content from Perplexity search answers. It will not protect your content from being used in OpenAI or Meta model training — to do that you'd need to block GPTBot and meta-externalagent respectively.
Does Bytespider respect robots.txt? +
This is disputed. Bytespider is ByteDance's crawler associated with TikTok's AI systems. There is no official vendor documentation page. It has been widely reported to not consistently respect robots.txt disallow rules. SEO Pulse notes this in the tooltip for Bytespider so you have accurate context when making decisions.
How does SEO Pulse check AI bot access? +
SEO Pulse fetches and parses your robots.txt, then runs the current page URL through the rule set for each of 20 known crawlers. The result is shown as ✓ allowed or ✗ blocked per bot in a visual grid inside the Robots tab. Clicking the Robots tab shows the full parsed robots.txt with rule-by-rule highlighting.

See your AI bot access in seconds

Install SEO Pulse, open any page, check the Robots tab. 20 bots checked instantly.

Add to Chrome — Free All features →

★★★★★ on Chrome Web Store · 100+ users · No data sent to servers