meta-externalagent: Meta's AI training crawler
Meta's crawler for use cases such as training foundation AI models or improving products by indexing content directly.
- Free, no account.
- A few seconds.
- Nothing is stored.
meta-externalagent at a glance
Everything Meta documents about it, with links to the source.
- Operator
- Meta
- Type
- AI training
- Used for
- Training Meta's AI models
- Obeys robots.txt
- Yes. Meta documents that meta-externalagent follows robots.txt, so a rule for its token is honoured.
- robots.txt token
meta-externalagent- User-agent string
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)
- IP addresses
- Not published.
How to block meta-externalagent
Collects text that may be used to train AI models. Blocking it keeps your pages out of future training data. AI answers can still cite you through the search crawlers.
Block the whole site
Add this group to robots.txt at the root of your site.
User-agent: meta-externalagent Disallow: /
Block only some paths
List the paths; everything else stays open to meta-externalagent.
User-agent: meta-externalagent Disallow: /private/ Disallow: /drafts/
How to verify meta-externalagent
Anyone can send a crawler's user-agent string. Where the request comes from is what proves it.
- 1
Find requests whose user agent contains "meta-externalagent" in your server or CDN logs.
- 2
Meta publishes no IP list for meta-externalagent, so it cannot be verified from the request alone. Treat unusual volume as a scraper using its name, and rate-limit or block it at the firewall.
- 3
Block what fails the check at the firewall. A robots.txt rule only reaches crawlers that choose to read it.
Other crawlers from Meta
Each has its own token, and blocking one leaves the others untouched.
Frequently asked questions
Meta's crawler for use cases such as training foundation AI models or improving products by indexing content directly. Meta asks sites to allow up to 24 hours for robots.txt changes to take effect.
Yes. Meta documents that meta-externalagent follows robots.txt, so a rule for its token is honoured.
Add a group for its token to robots.txt at the root of your site: "User-agent: meta-externalagent" followed by "Disallow: /". To close only part of the site, list those paths instead of "/".
No. meta-externalagent is a training crawler: blocking it keeps your pages out of data that may train Meta's models. AI answers that cite pages come from search crawlers, which have their own tokens, so blocking meta-externalagent alone does not remove you from them.
Checked against Meta's documentation on 29 September 2026.
Next: see whether AI recommends you
Letting the right crawlers in is the first step. AskWatch shows what ChatGPT, Perplexity, Gemini and Google AI answer when buyers ask about your category, and who they name.
- Free, no credit card.
- Report in minutes, link sent to your email.