robots.txt tester and validator
Check robots.txt for any site: which search and AI bots it lets in, the line that decides each one, and the mistakes that block pages you meant to open.
- Free, no account.
- A few seconds.
- Nothing is stored.
Which AI crawlers your robots.txt lets in
The tester checks Googlebot, four other search engines and 21 AI crawlers. Each one below links to its operator's own documentation.
Search engines
- GooglebotGoogleGoogle Search, including AI Overviews and AI Moderobots.txt: Yes
- BingbotMicrosoftBing search, which also feeds Copilotrobots.txt: Yes
- ApplebotAppleSpotlight, Siri and Safari suggestionsrobots.txt: Yes
- DuckDuckBotDuckDuckGoDuckDuckGo searchrobots.txt: Yes
- YandexBotYandexYandex searchrobots.txt: Yes
AI search and answers
- OAI-SearchBotOpenAIPages shown and cited in ChatGPT searchrobots.txt: Yes
- Claude-SearchBotAnthropicSearch results for Claude's answersrobots.txt: Yes
- PerplexityBotPerplexityPerplexity's index for cited answers, not trainingrobots.txt: Yes
- DuckAssistBotDuckDuckGoSources for DuckAssist answers, not trainingrobots.txt: Yes
- Amzn-SearchBotAmazonSearch experiences in Amazon products, such as Alexarobots.txt: Yes
- MistralAI-IndexMistralMistral search, which answers questions in Viberobots.txt: Yes
AI training
- GPTBotOpenAITraining OpenAI's modelsrobots.txt: Yes
- ClaudeBotAnthropicTraining Anthropic's modelsrobots.txt: Yes
- Google-ExtendedGoogleGemini and Vertex AI training; no effect on Search or AI Overviewsrobots.txt: Read by another crawler
- Applebot-ExtendedAppleApple Intelligence training; Applebot reads the rulerobots.txt: Read by another crawler
- CCBotCommon CrawlThe open web archive many models train onrobots.txt: Yes
- meta-externalagentMetaTraining Meta's AI modelsrobots.txt: Yes
- BytespiderByteDanceTraining ByteDance's modelsrobots.txt: Not documented
- AmazonbotAmazonAmazon's services, including model trainingrobots.txt: Yes
- MistralAI-TrainingMistralTraining Mistral's modelsrobots.txt: Yes
Fetches on a user's request
- ChatGPT-UserOpenAIPages a ChatGPT user asks it to openrobots.txt: Not always
- Claude-UserAnthropicPages a Claude user asks it to openrobots.txt: Yes
- Perplexity-UserPerplexityPages a Perplexity user asks it to openrobots.txt: Not always
- meta-externalfetcherMetaLinks a user asks a Meta AI product to fetchrobots.txt: Not always
- Amzn-UserAmazonRequests a user starts in Amazon's productsrobots.txt: Not always
- MistralAI-UserMistralPages a Vibe user's question leads it torobots.txt: Yes
The last column is what each operator says about robots.txt. Tokens checked against their documentation on 29 September 2026.
What blocking each bot costs you
Three kinds of AI crawler, three different prices. Most mistakes come from blocking one when you meant another.
Block AI training
Stays in answersGPTBot, ClaudeBot, Google-Extended, CCBot, meta-externalagent
Your pages stay out of future training sets of the crawlers that obey. AI answers can still find and cite you through the search crawlers. This is the block most publishers choose.
Block AI search
Leaves AI answersOAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot
These crawlers fetch the pages AI answers cite. Block them and those answers cannot link to you: OpenAI says a site that blocks OAI-SearchBot does not appear in ChatGPT search.
Block user fetches
Stays in answersChatGPT-User, Claude-User, Perplexity-User
These open a page because a person asked an assistant to read it. Several say robots.txt may not apply to them, so a Disallow is a request at best. A firewall rule is what actually blocks them.
Google-Extended does not take you out of AI Overviews. It controls Gemini training and grounding. AI Overviews and AI Mode are part of Search and use what Googlebot crawls.
What the robots.txt validator checks
Five groups. Every threshold is Google's or RFC 9309's, and each finding points at its line in the file.
Delivery
The status code and what Google does with it (a 4xx allows everything, a 5xx or 429 stops crawling), redirects past five, the content type, UTF-8, and the 500 KiB Google reads.
Syntax
Rules above the first User-agent, groups that a blank line did not split, unknown or misspelled fields, missing colons, and paths that do not start with a slash.
Google support
Lines Google does not read: noindex and nofollow (dropped in 2019), Crawl-delay, Yandex's Host and Clean-param. Content-Signal and Cloudflare's managed block are labelled, not flagged.
Bots
Search engines kept off the home page, CSS and JavaScript closed to Googlebot, AI search blocked while training is open, and tokens no crawler sends: old names, product names and typos.
Sitemap
Sitemap lines with full URLs, sitemaps that answer, and sitemaps the file blocks from Googlebot itself.
How robots.txt rules work
Six rules explain almost every surprising result. They are the same in Google's parser and in RFC 9309.
- 1
Groups start with User-agent
One or more User-agent lines, then the rules for them. A crawler obeys the groups that name its token and ignores the * group when any group names it. Tokens are case-insensitive and must match exactly: Claude is not ClaudeBot.
User-agent: GPTBot User-agent: CCBot Disallow: /
- 2
Only a rule ends a group
A blank line does not. Two User-agent lines with a blank line between them and no rule in between share the rules below them: here every crawler is blocked, not only GPTBot. Groups that name the same crawler twice are merged.
User-agent: * User-agent: GPTBot Disallow: /
- 3
The longest rule wins
Of the rules that match a path, the one with the longest path wins. On a tie, Allow beats Disallow. The order of lines does not matter.
Disallow: /docs/ Allow: /docs/public/
- 4
Two wildcards, case-sensitive paths
* matches any run of characters and $ anchors the end of the URL. Paths are case-sensitive: /Blog and /blog are different paths. A path always starts with /.
Disallow: /*.pdf$ Disallow: /*?sessionid=
- 5
The status code is a rule too
A 4xx means no restrictions. A 5xx or 429 means stop: Google pauses crawling and then works from its last good copy. More than five redirects counts as a 404. Google caches the file for up to 24 hours.
- 6
500 KiB, UTF-8, one host
Google reads the first 500 KiB. The file is UTF-8 plain text at the root of each host and protocol: example.com and www.example.com each need their own.
The full rules: Google's robots.txt documentation and RFC 9309.
robots.txt examples
Five files you can copy as they are. Paste any of them into the tester above to see what each crawler gets.
Allow everything
An empty Disallow allows every path. The same as having no file, with a place for the sitemap.
User-agent: * Disallow: Sitemap: https://yourdomain.com/sitemap.xml
Disallow everything
For a staging site or a private tool. Never ship it to production: it removes the site from search.
User-agent: * Disallow: /
Block AI training, stay in AI answers
Training crawlers out, search and AI search crawlers in. What most publishers who block AI choose.
User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: CCBot User-agent: meta-externalagent Disallow: / User-agent: * Allow: /
Block all known AI crawlers
Training and AI search. Pages will not be cited by the answers these crawlers feed; Googlebot, and with it AI Overviews, stays in.
User-agent: GPTBot User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Claude-SearchBot User-agent: Claude-User User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Google-Extended User-agent: CCBot User-agent: meta-externalagent User-agent: Bytespider Disallow: /
Point crawlers at your sitemap
A full URL, anywhere in the file, outside any group. List as many sitemaps as you have.
Sitemap: https://yourdomain.com/sitemap.xml Sitemap: https://yourdomain.com/blog/sitemap.xml
Blocked by robots.txt: how to fix it
Search Console reports robots.txt blocks under two statuses, and they need opposite fixes.
- Blocked by robots.txt
Googlebot may not crawl the page, so it is not indexed. If the page should be in Google, a rule is in the way.
- Indexed, though blocked by robots.txt
Google indexed the URL from links to it without reading the page, so it can appear with no description. robots.txt cannot remove it: only a noindex Google can see can.
- 1
Paste the page's URL into the tester above and read the Googlebot row: it names the rule and the line that blocks it.
- 2
Decide whether the page belongs in Google. Filters, carts and internal search usually do not; articles, products and landing pages do.
- 3
If it belongs, remove or narrow the rule. An Allow for the exact path beats a shorter Disallow, so you can open one folder without touching the rest.
- 4
If it does not belong but is indexed, allow crawling and add a noindex meta tag or X-Robots-Tag header, then block it again once it has dropped out.
- 5
Publish the file, test the URL again here, then use Validate fix in Search Console. Google caches robots.txt for up to 24 hours.
Common robots.txt mistakes
The tester flags each of these. Most of them are one line.
Disallow: / shipped from staging
The most expensive line in SEO. It takes the site out of search and out of AI Overviews within days.
CSS and JavaScript blocked
Google renders pages. Without the styles and scripts it may see a broken, unusable page.
noindex in robots.txt
Google stopped reading it in 2019. Use a meta tag or header on the page.
Blocking a page to deindex it
A blocked page can stay indexed, because Google can no longer see the noindex on it.
Wrong case in paths
Disallow: /Admin does not match /admin. Paths are case-sensitive.
A blank line to split groups
It does not split them. Two User-agent lines with nothing but a blank line between them share one set of rules.
Product names instead of tokens
User-agent: ChatGPT blocks nothing. The tokens are GPTBot, OAI-SearchBot and ChatGPT-User.
Old tokens
anthropic-ai and Claude-Web are retired; Anthropic's crawlers are ClaudeBot, Claude-SearchBot and Claude-User.
Blocking AI search to stop training
Blocking OAI-SearchBot or PerplexityBot removes you from AI answers. Training has its own tokens.
An HTML page at /robots.txt
A catch-all route that answers every path with the app's page gives crawlers no rules at all.
A relative sitemap
Sitemap takes a full URL with https://. A relative path is ignored.
A file past 500 KiB
Everything after it is ignored. Wildcards replace long lists of URLs.
Frequently asked questions
A plain-text file at the root of a site (yourdomain.com/robots.txt) that tells crawlers which paths they may fetch. It is a list of groups: one or more User-agent lines naming crawlers, then Allow and Disallow rules for them. It is standardised as RFC 9309 since 2022. It controls crawling, not indexing, and it is a request: well-behaved crawlers follow it, nothing enforces it.
Google retired the robots.txt Tester in Search Console at the end of 2023 and replaced it with the robots.txt report. The report shows which robots.txt files Google found for your properties, when it last fetched them and any lines it could not parse, but it does not test a URL against the rules. Bing Webmaster Tools still has a tester for Bingbot. This tester checks any URL for Googlebot and the AI crawlers at once.
No. GPTBot collects training data. ChatGPT search shows and cites pages that OAI-SearchBot crawls, and OpenAI says a site that blocks OAI-SearchBot will not appear in ChatGPT search answers. Blocking GPTBot and allowing OAI-SearchBot keeps you out of training and in the answers, which is what most sites that block AI want.
No. Google-Extended is a token that controls whether your content is used to train Gemini models and for grounding in Gemini apps and Vertex AI. It does not affect Google Search, and AI Overviews and AI Mode are part of Search: they use what Googlebot crawls. The only way to keep pages out of AI Overviews through robots.txt is to block Googlebot, which removes them from Search too.
Yes. robots.txt is a convention, and some crawlers say in their own documentation that it may not apply: OpenAI for ChatGPT-User, Perplexity for Perplexity-User and Meta for meta-externalfetcher, because those fetch a page a person asked for. Some crawlers publish no documentation at all. To keep a bot out for certain, block it at the server or the firewall, by its verified IP ranges.
No. A Disallow stops Google from crawling the page, not from indexing it: if other sites link to it, Google can index the URL without seeing its content, and Search Console reports it as "Indexed, though blocked by robots.txt". To keep a page out of the index, let Google crawl it and add a noindex meta tag or X-Robots-Tag header.
A line Cloudflare introduced in September 2025, for example Content-Signal: search=yes, ai-train=no. It states how content may be used after it is fetched: for search results, as input to AI answers (ai-input) or for training (ai-train). It is a statement of preference, not a block. Google's parser does not read it, so Search Console reports the line as "Syntax not understood"; that is harmless. The tester shows the signals and labels the line instead of calling it an error.
Because managed robots.txt is on in Cloudflare. It prepends a block between # BEGIN Cloudflare Managed content and # END Cloudflare Managed Content, usually disallowing AI training crawlers and adding a Content-Signal line, and it merges with your own rules for any crawler named in both. You can change it or turn it off in the Cloudflare dashboard. Cloudflare can also block AI crawlers at the edge regardless of robots.txt, so check its bot settings as well.
Not by Google, which ignores Crawl-delay and sets its own rate. Bing honours it, and Anthropic says its crawlers support it. If a crawler is putting load on your server, Crawl-delay is worth setting for the ones that read it.
Google reads the first 500 KiB and ignores the rest, and RFC 9309 sets the same floor for every parser. Rules past that point do nothing. A file that large is usually a list of individual URLs; wildcards and path prefixes say the same thing in a few lines.
robots.txt, plural, lowercase, at the root of the host: yourdomain.com/robots.txt. A file named robot.txt, or one placed in a folder, is not read by any crawler. Each host needs its own file: www.example.com and blog.example.com are read separately.
No. We fetch the file as AskWatchBot and read the rules for each crawler from it, the way the crawler itself would. We do not send another crawler's user agent. Most sites serve one robots.txt to everyone; a few serve a different file depending on who asks, and for those the result is the file served to us. Search Console's robots.txt report shows the copy Google fetched.
No. The result of a domain check is cached for ten minutes so that a link to it does not fetch your site again on every open; after that it is gone. Pasted text is checked in the request and not kept. There is no account and nothing to delete.
Other free tools
- Open tool
AI Visibility Checker
Check if ChatGPT, Gemini, Perplexity and Google AI name your brand, and get your AI visibility score.
- Open tool
AI Overview Checker
See if Google AI Overviews cite your site, which of your pages they pick, and who is cited instead.
- Open tool
Perplexity Visibility Tracker
Track whether Perplexity cites your site and names your brand in its answers.
- Open tool
ChatGPT Visibility Tracker
See if ChatGPT mentions your brand and where you rank in its answers.
- Open tool
Gemini Visibility Tracker
Check if Google Gemini mentions your brand and who it names instead.
- Open tool
Google AI Mode Visibility Tracker
See if Google AI Mode names your brand and which pages it cites.
- Open tool
Copilot Visibility Tracker
See if Microsoft Copilot mentions your brand and cites your site.
- Open tool
llms.txt Generator
Write an llms.txt for any site from its sitemap and page descriptions. No model, no watermark, checked against the spec.
- Open tool
llms.txt Checker
See whether a site's llms.txt exists, follows the spec, has live links and is not blocked for AI crawlers.
- Open tool
robots.txt Generator
Create a robots.txt in a minute: block AI training, stay in AI search, add paths and your sitemap. Tested before you download it.
- Open tool
AI Crawlers List
Every AI crawler with its user agent, IP list and robots.txt token, and whether your site lets each one in.
- Open tool
Google Search Console MCP
Ask Search Console in plain language from Claude, ChatGPT, Cursor or Claude Code. No Google Cloud project.
Next: see whether AI recommends you
robots.txt decides which AI crawlers may read your site. AskWatch shows what ChatGPT, Perplexity, Gemini and Google AI answer when buyers ask about your category, and who they name.
- Free, no credit card.
- Report in minutes, link sent to your email.