Updated29 September 2026

AI crawlers list

Every AI crawler by name: who runs it, what it collects, its user agent and IP list, and what blocking it costs. Enter your domain to see which ones your robots.txt lets in.

  • Free, no account.
  • A few seconds.
  • Nothing is stored.

Search engines

Crawls for a search engine's index. Blocking it removes your pages from that search engine, and from the AI answers built on it (AI Overviews on Googlebot, Copilot on Bingbot).

AI search and answers

Fetches the pages AI answers cite and link to. Blocking it keeps your pages out of those answers: no citation, no link.

AI training

Collects text that may be used to train AI models. Blocking it keeps your pages out of future training data. AI answers can still cite you through the search crawlers.

  • GPTBotOpenAITraining OpenAI's modelsrobots.txt: YesIP list
  • ClaudeBotAnthropicTraining Anthropic's modelsrobots.txt: YesIP list
  • Google-ExtendedGoogleGemini and Vertex AI training; no effect on Search or AI Overviewsrobots.txt: Read by another crawlerIP list: none
  • Applebot-ExtendedAppleApple Intelligence training; Applebot reads the rulerobots.txt: Read by another crawlerIP list: none
  • CCBotCommon CrawlThe open web archive many models train onrobots.txt: YesIP list
  • meta-externalagentMetaTraining Meta's AI modelsrobots.txt: YesIP list: none
  • BytespiderByteDanceTraining ByteDance's modelsrobots.txt: Not documentedIP list: none
  • AmazonbotAmazonAmazon's services, including model trainingrobots.txt: YesIP list
  • MistralAI-TrainingMistralTraining Mistral's modelsrobots.txt: YesIP list: none

Fetches on a user's request

Opens a page because a person asked an assistant to read it. Blocking it stops those visits for the operators that honour robots.txt. Several say they may not, so a firewall rule is the real block.

Download the list:JSONCSVon GitHub, free to reuse under CC BY 4.0.

How to block AI crawlers

Name each token in robots.txt. The usual choice: keep training crawlers out, let search and AI search in.

The file on the right blocks every training crawler in this list and leaves the rest alone, so your pages stay in Google, Bing and the AI answers that cite sources, and out of future training sets of the crawlers that obey.

Blocking AI search crawlers too is one more group, and it takes you out of ChatGPT search, Perplexity and Claude's answers. Fetchers acting on a user's request may not read robots.txt at all; a firewall rule is what stops them.

Build yours in the generator Test the one you have

User-agent: *
Disallow:

# AI training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
User-agent: Bytespider
User-agent: Amazonbot
User-agent: MistralAI-Training
Disallow: /

How to tell a real AI crawler from a fake one

User agents are free to copy, and scrapers often borrow a known crawler's name. Where a request comes from is what proves it.

Check the IP

17 of the 26 crawlers here come with a published IP list. A request outside the list is not that crawler.

Or reverse DNS

Google, Bing, Apple and Yandex document host names instead: the IP resolves to their domain, and the host resolves back to the same IP.

Block the rest at the edge

What fails the check is a scraper. Deny it at your firewall or CDN; robots.txt only reaches crawlers that choose to read it.

Frequently asked questions

A bot run by an AI company that fetches web pages. Some collect text to train models (GPTBot, ClaudeBot, CCBot), some build the index AI answers cite (OAI-SearchBot, PerplexityBot, Claude-SearchBot), and some fetch a page because a person asked an assistant to read it (ChatGPT-User, Claude-User). Each one has a user-agent token you can name in robots.txt.

Other free tools

  • AI Visibility Checker

    Check if ChatGPT, Gemini, Perplexity and Google AI name your brand, and get your AI visibility score.

    Open tool
  • AI Overview Checker

    See if Google AI Overviews cite your site, which of your pages they pick, and who is cited instead.

    Open tool
  • Perplexity Visibility Tracker

    Track whether Perplexity cites your site and names your brand in its answers.

    Open tool
  • ChatGPT Visibility Tracker

    See if ChatGPT mentions your brand and where you rank in its answers.

    Open tool
  • Gemini Visibility Tracker

    Check if Google Gemini mentions your brand and who it names instead.

    Open tool
  • Google AI Mode Visibility Tracker

    See if Google AI Mode names your brand and which pages it cites.

    Open tool
  • Copilot Visibility Tracker

    See if Microsoft Copilot mentions your brand and cites your site.

    Open tool
  • llms.txt Generator

    Write an llms.txt for any site from its sitemap and page descriptions. No model, no watermark, checked against the spec.

    Open tool
  • llms.txt Checker

    See whether a site's llms.txt exists, follows the spec, has live links and is not blocked for AI crawlers.

    Open tool
  • robots.txt Generator

    Create a robots.txt in a minute: block AI training, stay in AI search, add paths and your sitemap. Tested before you download it.

    Open tool
  • robots.txt Tester

    Test any URL against robots.txt for Googlebot and 20+ AI crawlers, and find the line that decides each one.

    Open tool
  • Google Search Console MCP

    Ask Search Console in plain language from Claude, ChatGPT, Cursor or Claude Code. No Google Cloud project.

    Open tool

Next: see whether AI recommends you

Letting the right crawlers in is the first step. AskWatch shows what ChatGPT, Perplexity, Gemini and Google AI answer when buyers ask about your category, and who they name.

  • Free, no credit card.
  • Report in minutes, link sent to your email.