Clovion AI
Free Tool

AI Crawlability Checker

Check if ChatGPT, Claude, Perplexity, Gemini, and Google AI Overviews can crawl your site — see exactly which AI bots your robots.txt blocks.

10 AI bots checkedInstant readoutNo signup
Background

What are AI crawlers?

AI crawlers are the automated user-agents that AI products send to your site so they can read it, index it, or cite it back to a user.

They are not all the same. Some bots collect training data. Some fetch pages live when a user asks a question. Some build the search index that powers AI Overviews. Your robots.txt controls each one independently.

If you want to appear in AI answers, the live-answer and retrieval bots are the ones that matter. The training crawlers are a separate, slower, content-policy decision.

ChatGPT uses GPTBot
GPTBot · ChatGPT-User

OpenAI runs two relevant agents: GPTBot collects training data, ChatGPT-User fetches pages when a ChatGPT user browses live.

Claude uses ClaudeBot family
ClaudeBot · Claude-Web · anthropic-ai

Anthropic ships three: ClaudeBot powers live answers, Claude-Web is the general crawler, anthropic-ai collects training data.

Perplexity uses PerplexityBot
PerplexityBot

Perplexity uses a single agent. Allowing it makes you eligible to appear as a cited source in Perplexity answers.

Gemini uses Google-Extended
Google-Extended

Google-Extended is the opt-out signal for Gemini and Google AI Overviews. Blocking it cuts you out of both.

Frequently Asked Questions

An AI crawler is an automated user-agent that visits webpages so AI products like ChatGPT, Claude, Perplexity, and Google AI Overviews can read, index, or cite their content. Different bots play different roles: some collect training data, some fetch live answers, some index the web for retrieval.

It depends on your visibility goals. If you want your brand to appear in AI answers, you generally want to allow live-answer and retrieval bots like ChatGPT-User, ClaudeBot, PerplexityBot, and Google-Extended. Blocking training-only bots like GPTBot or anthropic-ai is a content-policy choice that has little impact on whether your brand surfaces today.

Edit your robots.txt to remove or scope any Disallow rules that target the bot. A common pattern is a User-agent line for each bot followed by an Allow: / directive. Push the change, wait for the next crawl, and re-run this checker to confirm the new status.

Indeterminate means we could not fetch robots.txt cleanly, the file was missing, or your server returned an unexpected status. The bot is probably allowed by default, but we cannot confirm. Most often it means the URL has a redirect chain, an authentication wall, or no robots.txt at the root.

No. Allowing crawler access is necessary but not sufficient. Citation depends on whether your content is structured, machine-readable, relevant, recent, and authoritative for the prompt. The free checker confirms access only — the full Clovion product also audits content, schema, and competitive position.

Yes, completely. The AI Crawlability Checker runs against your live robots.txt, returns a per-bot allow/block readout in seconds, and never asks for a card or signup. The paid Clovion plan adds daily monitoring, multi-page checks, llms.txt generation, and recommendations tied to your visibility score.

Get your score

Know your full AI visibility score.

Crawler access is one signal. The full score adds mention rate, sentiment, citations, and competitive position across ChatGPT, Claude, Perplexity, and AI Overviews — in about a minute.

Get Free Score