Free robots.txt generator
Pick who may crawl your site, from Googlebot to GPTBot, add the paths to close and your sitemap, and download the file. It is checked against the tester's rules as you type.
Start from
A stance on crawlers. Your paths and sitemaps stay when you switch.
Or start from your current robots.txt
Crawlers
Switch a crawler off to keep it out of the whole site. The rest follow the paths below.
Search engines
Googlebot
Google · Google Search, including AI Overviews and AI Mode
AllowedBingbot
Microsoft · Bing search, which also feeds Copilot
AllowedApplebot
Apple · Spotlight, Siri and Safari suggestions
AllowedDuckDuckBot
DuckDuckGo · DuckDuckGo search
AllowedYandexBot
Yandex · Yandex search
Allowed
AI search and answers
OAI-SearchBot
OpenAI · Pages shown and cited in ChatGPT search
AllowedClaude-SearchBot
Anthropic · Search results for Claude's answers
AllowedPerplexityBot
Perplexity · Perplexity's index for cited answers, not training
AllowedDuckAssistBot
DuckDuckGo · Sources for DuckAssist answers, not training
AllowedAmzn-SearchBot
Amazon · Search experiences in Amazon products, such as Alexa
AllowedMistralAI-Index
Mistral · Mistral search, which answers questions in Vibe
Allowed
AI training
GPTBot
OpenAI · Training OpenAI's models
BlockedClaudeBot
Anthropic · Training Anthropic's models
BlockedGoogle-Extended
Google · Gemini and Vertex AI training; no effect on Search or AI Overviews
BlockedApplebot-Extended
Apple · Apple Intelligence training; Applebot reads the rule
BlockedCCBot
Common Crawl · The open web archive many models train on
Blockedmeta-externalagent
Meta · Training Meta's AI models
BlockedBytespider
ByteDance · Training ByteDance's models
BlockedAmazonbot
Amazon · Amazon's services, including model training
BlockedMistralAI-Training
Mistral · Training Mistral's models
Blocked
Fetches on a user's request
ChatGPT-User
OpenAI · Pages a ChatGPT user asks it to open
AllowedClaude-User
Anthropic · Pages a Claude user asks it to open
AllowedPerplexity-User
Perplexity · Pages a Perplexity user asks it to open
Allowedmeta-externalfetcher
Meta · Links a user asks a Meta AI product to fetch
AllowedAmzn-User
Amazon · Requests a user starts in Amazon's products
AllowedMistralAI-User
Mistral · Pages a Vibe user's question leads it to
Allowed
Paths
Closed to every crawler that is not blocked outright. One path per line; * and $ work.
Sitemap
Full URLs, one per line. Every crawler reads these, including the ones you never submit a sitemap to.
Content signals
Optional. Cloudflare's line stating how fetched content may be used. A preference, not a block; Google skips it.
Other rules
Kept exactly as written, after everything above. For crawlers not listed here, or partial blocks.
robots.txt
1User-agent: *2Disallow:3 4# AI training5User-agent: GPTBot6User-agent: ClaudeBot7User-agent: Google-Extended8User-agent: Applebot-Extended9User-agent: CCBot10User-agent: meta-externalagent11User-agent: Bytespider12User-agent: Amazonbot13User-agent: MistralAI-Training14Disallow: /
- Search enginesallowed
- AI searchallowed
- AI trainingall blocked
- User fetchesallowed
Passes every check the robots.txt tester runs.
Save it as robots.txt at the root of the site, then test it for any URL
How to create a robots.txt file
Four steps, and the generator above does the first three.
- 1
Choose who may crawl
Start from a preset. Block AI training suits most sites: search and AI answers stay, training sets do not get your pages.
- 2
Close the paths nobody needs
Admin, cart, account and internal search pages waste crawl budget and never answer a search. One click each.
- 3
Point to your sitemap
A full URL. Every crawler that reads robots.txt finds your page list, not only the ones you submit it to.
- 4
Upload and test
Save it as robots.txt at the root of the site, then test any URL against it for Googlebot and every AI crawler.
Which AI crawlers to block
What each preset lets in. Blocking training does not take you out of AI answers; blocking AI search does.
| Who reads your site | Allow all | Block AI training | Block all AI | Block everything |
|---|---|---|---|---|
| Google Search, AI Overviews, AI Mode | ||||
| Bing and Copilot | ||||
| ChatGPT search, Perplexity, Claude answers | ||||
| Pages a user asks an assistant to open | ||||
| Training of GPT, Claude, Gemini and other models |
Fetches a user asks for (ChatGPT-User, Perplexity-User) may skip robots.txt by their own documentation, so blocking them is a request rather than a guarantee.
Where to put robots.txt
At the root of the site, as plain text. Most platforms manage the file for you; here is where each one keeps it.
- 1.Download the file and upload it to the web root, where your home page lives, so it answers at /robots.txt.
- 2.Open it in a browser to confirm it is served as plain text:
https://yourdomain.com/robots.txt
- 3.Each subdomain (www, blog, shop) serves its own file.
What to disallow, and what to leave open
Close what wastes a crawler's time. Leave open what Google needs to render a page and what you want cited.
Usually disallow
Admin and login
/wp-admin/, /admin/, /login. Nothing there answers a search.
Cart, checkout and account
Personal, empty for a crawler, and endless on some shops.
Internal search results
/search and ?s= pages. Google's guidance is to keep them from being crawled.
Sort, filter and session parameters
The same page under thousands of URLs. Close the parameters, keep the page.
Never disallow
CSS, JavaScript and images
Google renders pages; without them it sees a broken one.
Pages you want out of Google
Use noindex and let Google crawl them, or it never sees the noindex.
Your sitemap
A Sitemap line that points into a closed folder points at nothing.
Pages you want cited by AI
Blog, docs, pricing, comparisons: what AI search crawlers quote and link to.
Frequently asked questions
A tool that writes the robots.txt file for you from a few choices: which crawlers may read the site, which paths are closed, and where the sitemap is. This one also knows every major AI crawler by name and group, so blocking AI training without leaving AI search is one click, and it runs the result through the same checks as our robots.txt tester before you download it.
robots.txt itself is how you talk to AI crawlers: each one has a user-agent token, such as GPTBot, ClaudeBot or PerplexityBot, and obeys the rules for it like any other crawler. What matters is picking the right tokens, because the same company often runs one crawler for training and another for search. There is also llms.txt, a different file that describes your site to AI agents rather than restricting them.
Use the Block AI training or Block all AI crawlers preset. Both leave Googlebot alone, and Googlebot is what Google Search, AI Overviews and AI Mode read. To keep your pages out of Gemini training as well, the presets block Google-Extended, which does not affect Search.
Yes, for the crawlers that choose to obey it, which includes Google, Bing, OpenAI, Anthropic, Perplexity and the other operators listed here. It is a public request, not a lock: fetchers acting on a person's request may skip it, some crawlers publish no policy at all, and anything truly private needs a login or a firewall rule.
It is a technical standard (RFC 9309), not a contract, and following it is voluntary. It still matters as a clear public statement of what you allow, and the major AI operators document that their crawlers follow it. For content you must protect, use authentication rather than robots.txt.
At the root of the host, so it answers at yourdomain.com/robots.txt, served as plain text. A file in a folder is never read. Each subdomain needs its own. Many platforms manage the file for you; the guide above shows where for WordPress, Shopify, Next.js, Webflow, Wix, Blogger and Squarespace.
WordPress serves a virtual robots.txt that closes /wp-admin/ and reopens /wp-admin/admin-ajax.php, and adds the sitemap. That is a good start: add the WordPress admin chip, your sitemap and the AI preset you want. Edit it through Yoast or Rank Math, or upload a physical robots.txt to the site root, which replaces the virtual one.
Yes, but it takes care. Block everyone with User-agent: * and Disallow: /, then give Googlebot its own group with Allow: /, since a named group replaces the * rules for that crawler. Most sites do not want this: Bing also feeds Copilot, and ChatGPT search reads pages through OAI-SearchBot, so you would disappear from both.
Only if a crawler is putting real load on your server. Google ignores it; Bing and Anthropic's crawlers honour it. Add it under Other rules for the crawler in question, for example User-agent: Bingbot followed by Crawl-delay: 5.
Yes. Enter your domain under Start from and the generator reads the file you serve now: blocked crawlers become switches, paths closed to everyone fill the path lists, sitemaps carry over, and anything the form has no control for is kept word for word under Other rules. Cloudflare's managed block is left out, because Cloudflare adds it itself.
No. The file is built in your browser and never sent to us. Starting from your current robots.txt reads that one public file for you and keeps nothing; the result is cached for ten minutes and then gone.
Other free tools
- Open tool
AI Visibility Checker
Check if ChatGPT, Gemini, Perplexity and Google AI name your brand, and get your AI visibility score.
- Open tool
AI Overview Checker
See if Google AI Overviews cite your site, which of your pages they pick, and who is cited instead.
- Open tool
Perplexity Visibility Tracker
Track whether Perplexity cites your site and names your brand in its answers.
- Open tool
ChatGPT Visibility Tracker
See if ChatGPT mentions your brand and where you rank in its answers.
- Open tool
Gemini Visibility Tracker
Check if Google Gemini mentions your brand and who it names instead.
- Open tool
Google AI Mode Visibility Tracker
See if Google AI Mode names your brand and which pages it cites.
- Open tool
Copilot Visibility Tracker
See if Microsoft Copilot mentions your brand and cites your site.
- Open tool
llms.txt Generator
Write an llms.txt for any site from its sitemap and page descriptions. No model, no watermark, checked against the spec.
- Open tool
llms.txt Checker
See whether a site's llms.txt exists, follows the spec, has live links and is not blocked for AI crawlers.
- Open tool
robots.txt Tester
Test any URL against robots.txt for Googlebot and 20+ AI crawlers, and find the line that decides each one.
- Open tool
AI Crawlers List
Every AI crawler with its user agent, IP list and robots.txt token, and whether your site lets each one in.
- Open tool
Google Search Console MCP
Ask Search Console in plain language from Claude, ChatGPT, Cursor or Claude Code. No Google Cloud project.
Next: see whether AI recommends you
robots.txt decides which AI crawlers may read your site. AskWatch shows what ChatGPT, Perplexity, Gemini and Google AI answer when buyers ask about your category, and who they name.
- Free, no credit card.
- Report in minutes, link sent to your email.