Clovion AI
FREE TOOL

Robots.txt AI Bot Checker.

Paste your robots.txt or enter a URL. See exactly which of 15 AI bots are allowed, blocked, or fall through the cracks of your rules.

The basics

What is robots.txt for AI?

Robots.txt is a 30-year-old text file sitting at the root of your site. It was designed for search-engine spiders, but AI engines inherited the same protocol. Today it is the front door of your domain for every AI crawler that respects standards — which is most of them.

A bot that is disallowed at the door cannot read your pages, which means it cannot quote, summarize, or recommend them in AI answers. A bot that is allowed gets to read — but reading is still only step one of being cited.

The checker walks every line of your robots.txt against a list of 15 known AI agents and tells you, per bot, whether you currently let them in.

User-agent blocks

Each block targets one crawler. "User-agent: *" applies to anything without a more specific block. Most AI bots will pick up their own named block if present.

Allow / Disallow rules

Allow opens a path, Disallow closes one. Modern crawlers honor the longest, most specific matching rule — so order is less important than precision.

AI bots respect robots.txt (mostly)

Major AI crawlers — GPTBot, ClaudeBot, PerplexityBot, Google-Extended — publicly commit to honoring robots.txt. A small minority ignore it; for those, network-level controls are the only reliable answer.

Frequently Asked Questions

robots.txt is a small text file at the root of your site that tells crawlers which paths they may fetch. AI assistants and AI search engines run their own crawlers, and most of them honor the same protocol web crawlers have used for decades. Blocking a bot here means it cannot read those pages — which means it cannot quote, summarize, or recommend them in AI answers either.

It depends on what you want. If you want to appear in answers from ChatGPT search, Claude, Perplexity, and Google AI Overviews, allow their crawlers (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended). Training-only crawlers (anthropic-ai, CCBot) are a separate question — some teams allow them to seed future models, others block them for IP reasons. The checker shows you all 15 so you can decide per-bot.

Find the User-agent block matching that bot (or the wildcard "User-agent: *" block if no specific one exists) and either remove the Disallow line or add an Allow rule for the paths you want crawled. The most permissive matching rule wins for most modern bots. Re-run the checker after deploying robots.txt to confirm the change took effect.

No. Allowing a bot is necessary, not sufficient. AI engines also weigh authority, freshness, structured data, citations from other domains, and whether your content actually answers the prompt. robots.txt opens the door — content quality and external signals decide whether you get cited once you are through it.

Indeterminate means your robots.txt has a rule that could be read multiple ways for that bot — for example, conflicting Allow and Disallow lines under a wildcard, or a pattern that some crawlers honor and others ignore. We flag it so you can audit the rule manually rather than assume the bot will resolve it the way you intended.

Yes. The robots.txt checker is free, with no signup. If you want a deeper picture — your share of voice across AI engines, which prompts surface your brand, sentiment, citations, and improvement recommendations — run a free AI visibility scan on your domain.

Get your score

Know your full AI visibility score.

robots.txt is the front door. The full scan goes inside: which AI engines surface your brand, where competitors win, and the three highest-lift fixes to close the gap — free, no card.

Get Free Score