No accountGooglebot and 21 AI crawlers

robots.txt tester and validator

Check robots.txt for any site: which search and AI bots it lets in, the line that decides each one, and the mistakes that block pages you meant to open.

  • Free, no account.
  • A few seconds.
  • Nothing is stored.

Which AI crawlers your robots.txt lets in

The tester checks Googlebot, four other search engines and 21 AI crawlers. Each one below links to its operator's own documentation.

Search engines

  • GooglebotGoogleGoogle Search, including AI Overviews and AI Moderobots.txt: Yes
  • BingbotMicrosoftBing search, which also feeds Copilotrobots.txt: Yes
  • ApplebotAppleSpotlight, Siri and Safari suggestionsrobots.txt: Yes
  • DuckDuckBotDuckDuckGoDuckDuckGo searchrobots.txt: Yes
  • YandexBotYandexYandex searchrobots.txt: Yes

AI search and answers

  • OAI-SearchBotOpenAIPages shown and cited in ChatGPT searchrobots.txt: Yes
  • Claude-SearchBotAnthropicSearch results for Claude's answersrobots.txt: Yes
  • PerplexityBotPerplexityPerplexity's index for cited answers, not trainingrobots.txt: Yes
  • DuckAssistBotDuckDuckGoSources for DuckAssist answers, not trainingrobots.txt: Yes
  • Amzn-SearchBotAmazonSearch experiences in Amazon products, such as Alexarobots.txt: Yes
  • MistralAI-IndexMistralMistral search, which answers questions in Viberobots.txt: Yes

AI training

  • GPTBotOpenAITraining OpenAI's modelsrobots.txt: Yes
  • ClaudeBotAnthropicTraining Anthropic's modelsrobots.txt: Yes
  • Google-ExtendedGoogleGemini and Vertex AI training; no effect on Search or AI Overviewsrobots.txt: Read by another crawler
  • Applebot-ExtendedAppleApple Intelligence training; Applebot reads the rulerobots.txt: Read by another crawler
  • CCBotCommon CrawlThe open web archive many models train onrobots.txt: Yes
  • meta-externalagentMetaTraining Meta's AI modelsrobots.txt: Yes
  • BytespiderByteDanceTraining ByteDance's modelsrobots.txt: Not documented
  • AmazonbotAmazonAmazon's services, including model trainingrobots.txt: Yes
  • MistralAI-TrainingMistralTraining Mistral's modelsrobots.txt: Yes

Fetches on a user's request

  • ChatGPT-UserOpenAIPages a ChatGPT user asks it to openrobots.txt: Not always
  • Claude-UserAnthropicPages a Claude user asks it to openrobots.txt: Yes
  • Perplexity-UserPerplexityPages a Perplexity user asks it to openrobots.txt: Not always
  • meta-externalfetcherMetaLinks a user asks a Meta AI product to fetchrobots.txt: Not always
  • Amzn-UserAmazonRequests a user starts in Amazon's productsrobots.txt: Not always
  • MistralAI-UserMistralPages a Vibe user's question leads it torobots.txt: Yes

The last column is what each operator says about robots.txt. Tokens checked against their documentation on 29 September 2026.

What blocking each bot costs you

Three kinds of AI crawler, three different prices. Most mistakes come from blocking one when you meant another.

  • Block AI training

    Stays in answers

    GPTBot, ClaudeBot, Google-Extended, CCBot, meta-externalagent

    Your pages stay out of future training sets of the crawlers that obey. AI answers can still find and cite you through the search crawlers. This is the block most publishers choose.

  • Block AI search

    Leaves AI answers

    OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot

    These crawlers fetch the pages AI answers cite. Block them and those answers cannot link to you: OpenAI says a site that blocks OAI-SearchBot does not appear in ChatGPT search.

  • Block user fetches

    Stays in answers

    ChatGPT-User, Claude-User, Perplexity-User

    These open a page because a person asked an assistant to read it. Several say robots.txt may not apply to them, so a Disallow is a request at best. A firewall rule is what actually blocks them.

Google-Extended does not take you out of AI Overviews. It controls Gemini training and grounding. AI Overviews and AI Mode are part of Search and use what Googlebot crawls.

What the robots.txt validator checks

Five groups. Every threshold is Google's or RFC 9309's, and each finding points at its line in the file.

  • Delivery

    The status code and what Google does with it (a 4xx allows everything, a 5xx or 429 stops crawling), redirects past five, the content type, UTF-8, and the 500 KiB Google reads.

  • Syntax

    Rules above the first User-agent, groups that a blank line did not split, unknown or misspelled fields, missing colons, and paths that do not start with a slash.

  • Google support

    Lines Google does not read: noindex and nofollow (dropped in 2019), Crawl-delay, Yandex's Host and Clean-param. Content-Signal and Cloudflare's managed block are labelled, not flagged.

  • Bots

    Search engines kept off the home page, CSS and JavaScript closed to Googlebot, AI search blocked while training is open, and tokens no crawler sends: old names, product names and typos.

  • Sitemap

    Sitemap lines with full URLs, sitemaps that answer, and sitemaps the file blocks from Googlebot itself.

How robots.txt rules work

Six rules explain almost every surprising result. They are the same in Google's parser and in RFC 9309.

  1. 1

    Groups start with User-agent

    One or more User-agent lines, then the rules for them. A crawler obeys the groups that name its token and ignores the * group when any group names it. Tokens are case-insensitive and must match exactly: Claude is not ClaudeBot.

    User-agent: GPTBot
    User-agent: CCBot
    Disallow: /
  2. 2

    Only a rule ends a group

    A blank line does not. Two User-agent lines with a blank line between them and no rule in between share the rules below them: here every crawler is blocked, not only GPTBot. Groups that name the same crawler twice are merged.

    User-agent: *
    
    User-agent: GPTBot
    Disallow: /
  3. 3

    The longest rule wins

    Of the rules that match a path, the one with the longest path wins. On a tie, Allow beats Disallow. The order of lines does not matter.

    Disallow: /docs/
    Allow: /docs/public/
  4. 4

    Two wildcards, case-sensitive paths

    * matches any run of characters and $ anchors the end of the URL. Paths are case-sensitive: /Blog and /blog are different paths. A path always starts with /.

    Disallow: /*.pdf$
    Disallow: /*?sessionid=
  5. 5

    The status code is a rule too

    A 4xx means no restrictions. A 5xx or 429 means stop: Google pauses crawling and then works from its last good copy. More than five redirects counts as a 404. Google caches the file for up to 24 hours.

  6. 6

    500 KiB, UTF-8, one host

    Google reads the first 500 KiB. The file is UTF-8 plain text at the root of each host and protocol: example.com and www.example.com each need their own.

The full rules: Google's robots.txt documentation and RFC 9309.

robots.txt examples

Five files you can copy as they are. Paste any of them into the tester above to see what each crawler gets.

  • Allow everything

    An empty Disallow allows every path. The same as having no file, with a place for the sitemap.

    User-agent: *
    Disallow:
    
    Sitemap: https://yourdomain.com/sitemap.xml
  • Disallow everything

    For a staging site or a private tool. Never ship it to production: it removes the site from search.

    User-agent: *
    Disallow: /
  • Block AI training, stay in AI answers

    Training crawlers out, search and AI search crawlers in. What most publishers who block AI choose.

    User-agent: GPTBot
    User-agent: ClaudeBot
    User-agent: Google-Extended
    User-agent: Applebot-Extended
    User-agent: CCBot
    User-agent: meta-externalagent
    Disallow: /
    
    User-agent: *
    Allow: /
  • Block all known AI crawlers

    Training and AI search. Pages will not be cited by the answers these crawlers feed; Googlebot, and with it AI Overviews, stays in.

    User-agent: GPTBot
    User-agent: OAI-SearchBot
    User-agent: ChatGPT-User
    User-agent: ClaudeBot
    User-agent: Claude-SearchBot
    User-agent: Claude-User
    User-agent: PerplexityBot
    User-agent: Perplexity-User
    User-agent: Google-Extended
    User-agent: CCBot
    User-agent: meta-externalagent
    User-agent: Bytespider
    Disallow: /
  • Point crawlers at your sitemap

    A full URL, anywhere in the file, outside any group. List as many sitemaps as you have.

    Sitemap: https://yourdomain.com/sitemap.xml
    Sitemap: https://yourdomain.com/blog/sitemap.xml

Blocked by robots.txt: how to fix it

Search Console reports robots.txt blocks under two statuses, and they need opposite fixes.

  • Blocked by robots.txt

    Googlebot may not crawl the page, so it is not indexed. If the page should be in Google, a rule is in the way.

  • Indexed, though blocked by robots.txt

    Google indexed the URL from links to it without reading the page, so it can appear with no description. robots.txt cannot remove it: only a noindex Google can see can.

  1. 1

    Paste the page's URL into the tester above and read the Googlebot row: it names the rule and the line that blocks it.

  2. 2

    Decide whether the page belongs in Google. Filters, carts and internal search usually do not; articles, products and landing pages do.

  3. 3

    If it belongs, remove or narrow the rule. An Allow for the exact path beats a shorter Disallow, so you can open one folder without touching the rest.

  4. 4

    If it does not belong but is indexed, allow crawling and add a noindex meta tag or X-Robots-Tag header, then block it again once it has dropped out.

  5. 5

    Publish the file, test the URL again here, then use Validate fix in Search Console. Google caches robots.txt for up to 24 hours.

Common robots.txt mistakes

The tester flags each of these. Most of them are one line.

  • Disallow: / shipped from staging

    The most expensive line in SEO. It takes the site out of search and out of AI Overviews within days.

  • CSS and JavaScript blocked

    Google renders pages. Without the styles and scripts it may see a broken, unusable page.

  • noindex in robots.txt

    Google stopped reading it in 2019. Use a meta tag or header on the page.

  • Blocking a page to deindex it

    A blocked page can stay indexed, because Google can no longer see the noindex on it.

  • Wrong case in paths

    Disallow: /Admin does not match /admin. Paths are case-sensitive.

  • A blank line to split groups

    It does not split them. Two User-agent lines with nothing but a blank line between them share one set of rules.

  • Product names instead of tokens

    User-agent: ChatGPT blocks nothing. The tokens are GPTBot, OAI-SearchBot and ChatGPT-User.

  • Old tokens

    anthropic-ai and Claude-Web are retired; Anthropic's crawlers are ClaudeBot, Claude-SearchBot and Claude-User.

  • Blocking AI search to stop training

    Blocking OAI-SearchBot or PerplexityBot removes you from AI answers. Training has its own tokens.

  • An HTML page at /robots.txt

    A catch-all route that answers every path with the app's page gives crawlers no rules at all.

  • A relative sitemap

    Sitemap takes a full URL with https://. A relative path is ignored.

  • A file past 500 KiB

    Everything after it is ignored. Wildcards replace long lists of URLs.

Frequently asked questions

A plain-text file at the root of a site (yourdomain.com/robots.txt) that tells crawlers which paths they may fetch. It is a list of groups: one or more User-agent lines naming crawlers, then Allow and Disallow rules for them. It is standardised as RFC 9309 since 2022. It controls crawling, not indexing, and it is a request: well-behaved crawlers follow it, nothing enforces it.

Other free tools

  • AI Visibility Checker

    Check if ChatGPT, Gemini, Perplexity and Google AI name your brand, and get your AI visibility score.

    Open tool
  • AI Overview Checker

    See if Google AI Overviews cite your site, which of your pages they pick, and who is cited instead.

    Open tool
  • Perplexity Visibility Tracker

    Track whether Perplexity cites your site and names your brand in its answers.

    Open tool
  • ChatGPT Visibility Tracker

    See if ChatGPT mentions your brand and where you rank in its answers.

    Open tool
  • Gemini Visibility Tracker

    Check if Google Gemini mentions your brand and who it names instead.

    Open tool
  • Google AI Mode Visibility Tracker

    See if Google AI Mode names your brand and which pages it cites.

    Open tool
  • Copilot Visibility Tracker

    See if Microsoft Copilot mentions your brand and cites your site.

    Open tool
  • llms.txt Generator

    Write an llms.txt for any site from its sitemap and page descriptions. No model, no watermark, checked against the spec.

    Open tool
  • llms.txt Checker

    See whether a site's llms.txt exists, follows the spec, has live links and is not blocked for AI crawlers.

    Open tool
  • robots.txt Generator

    Create a robots.txt in a minute: block AI training, stay in AI search, add paths and your sitemap. Tested before you download it.

    Open tool
  • AI Crawlers List

    Every AI crawler with its user agent, IP list and robots.txt token, and whether your site lets each one in.

    Open tool
  • Google Search Console MCP

    Ask Search Console in plain language from Claude, ChatGPT, Cursor or Claude Code. No Google Cloud project.

    Open tool

Next: see whether AI recommends you

robots.txt decides which AI crawlers may read your site. AskWatch shows what ChatGPT, Perplexity, Gemini and Google AI answer when buyers ask about your category, and who they name.

  • Free, no credit card.
  • Report in minutes, link sent to your email.