aeotime

Free tool · no signup

Free AI Crawler Checker

Enter a domain to see which AI bots its robots.txt allows, blocks or partly blocks: GPTBot, ClaudeBot, PerplexityBot, Google-Extended and nine more, with the rule that decides each one.

What this checker tests

AI crawlers vs. robots.txt and JavaScript · illustrative

The tool downloads /robots.txt from the domain you enter and evaluates it for 13 AI user-agents the way the crawlers themselves do, following RFC 9309. A bot obeys the group that names it; if none does, it falls back to User-agent: *. Inside a group the longest matching path wins, and on a tie Allow beats Disallow.

Each bot gets one of three results for your homepage: Allowed (no rule stops it), Partial (the homepage is open but some paths are disallowed) or Blocked (it can't fetch the homepage at all). The "Why" column tells you whether the verdict came from the bot's own group or from the wildcard group — the most common source of accidental blocks.

Training bots, search bots and user fetchers

AI companies now run several agents each, and blocking one is not the same as blocking the company.

  • Training crawlers (GPTBot, ClaudeBot, CCBot, Bytespider, Meta-ExternalAgent) collect pages that may be used to train models. Blocking them is a legitimate choice and doesn't remove you from AI search results.
  • Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) build the index that AI search answers cite. Block these and you are unlikely to be cited or linked by that engine.
  • User fetchers (ChatGPT-User, Claude-User, Perplexity-User) load a page when a person asks about it. OpenAI says robots.txt may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores it.
  • Usage tokens (Google-Extended, Applebot-Extended) aren't crawlers. They tell Google and Apple whether content their regular bots fetched may be used for AI. Google states that Google-Extended does not affect inclusion in Google Search, and AI Overviews are built from Googlebot's crawl.

Why robots.txt isn't the whole story

robots.txt is a request, not a lock. A bot can be allowed here and still get a 403 from your CDN or firewall — several CDNs now offer one-click AI bot blocking that overrides whatever robots.txt says. The reverse also happens: a bot is blocked in robots.txt, but an old rule you forgot about is the reason.

To see what a crawler actually receives, run the same URL through the AI crawler view, which requests the page as GPTBot, ClaudeBot and PerplexityBot and compares the responses with a normal browser. To change your rules on purpose, build a new file with the robots.txt generator.

FAQ

Questions

Does blocking GPTBot remove my site from ChatGPT?+

Not by itself. GPTBot is OpenAI's training crawler. ChatGPT search results come from OAI-SearchBot, and pages users ask about are fetched by ChatGPT-User. To stay citable in ChatGPT search, keep OAI-SearchBot allowed.

Why does a bot show as blocked when I never mention it?+

Because it falls back to the User-agent: * group. If that group contains Disallow: /, every bot without its own group is blocked. Add a named group with Allow: / for the bots you want.

Does Google-Extended control AI Overviews?+

No. Google says Google-Extended does not affect inclusion or ranking in Google Search. AI Overviews and AI Mode use Googlebot's index, so the only way to leave them is to limit Googlebot or use snippet controls such as nosnippet.

How often do AI crawlers re-read robots.txt?+

Each operator decides, but RFC 9309 says crawlers should not cache robots.txt for more than 24 hours in general. Expect changes to take effect within a day or two, not instantly.

Keep reading

Related

Check once. Or track it every week.

aeotime runs your prompts on eight AI engines every week, shows who gets recommended instead of you, and turns it into an action plan.