AI Crawler Access Checker

Check whether GPTBot, ClaudeBot, PerplexityBot, Google-Extended and other AI bots may crawl your site.

What is an AI crawler?

AI crawlers are bots run by AI companies such as OpenAI, Anthropic, Google and Perplexity. Some collect pages to train models; others fetch pages to answer users' questions in AI search. Sites control their access with robots.txt.

Three kinds of AI bots

PurposeExamplesIf you block it
TrainingGPTBot, ClaudeBot, Google-Extended, CCBotYour content is not used to train future models
AI searchOAI-SearchBot, Claude-SearchBot, PerplexityBotYour pages may not be cited in AI search answers
User fetchChatGPT-User, Claude-User, Perplexity-UserAssistants cannot open your page when a user asks about it

How the check works

The tool reads your robots.txt and evaluates the rules for each bot exactly as a crawler would — a bot-specific group first, otherwise the User-agent: * group — for the path you enter.

Choosing a policy

Many publishers block training bots but allow search and user bots so they can still be cited and get traffic from AI assistants. Example:

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

Frequently asked questions

Does blocking Google-Extended remove me from Google Search?

No. Google-Extended only controls use of your content for Gemini training. Googlebot, which powers Search and AI Overviews, is separate.

Do AI bots respect robots.txt?

The major companies say their crawlers follow robots.txt. It is a request, not technical enforcement.

Should I block AI crawlers?

It is a business decision. Blocking training bots protects your content; allowing search bots helps you appear in AI answers.