robots.txtNot documented

Bytespider: ByteDance's AI training crawler

ByteDance's crawler, widely reported as collecting data for its AI models.

  • Free, no account.
  • A few seconds.
  • Nothing is stored.

Bytespider at a glance

Everything ByteDance documents about it, with links to the source.

Operator
ByteDance
Type
AI training
Used for
Training ByteDance's models
Obeys robots.txt
ByteDance publishes no documentation for Bytespider, so there is no statement to rely on. A robots.txt rule is a request; a firewall rule is the only enforcement.
robots.txt token
Bytespider
User-agent string
ByteDance does not publish the full string. Match requests on the token Bytespider; there is no published IP list to verify them against.
IP addresses
Not published.
Documentation
None published.

How to block Bytespider

Collects text that may be used to train AI models. Blocking it keeps your pages out of future training data. AI answers can still cite you through the search crawlers.

Block the whole site

Add this group to robots.txt at the root of your site.

User-agent: Bytespider
Disallow: /

Block only some paths

List the paths; everything else stays open to Bytespider.

User-agent: Bytespider
Disallow: /private/
Disallow: /drafts/

robots.txt is a request Bytespider may not follow. To enforce it, block its requests at your server or firewall.

Generate a full robots.txt Test it for any URL

How to verify Bytespider

Anyone can send a crawler's user-agent string. Where the request comes from is what proves it.

  1. 1

    Find requests whose user agent contains "Bytespider" in your server or CDN logs.

  2. 2

    ByteDance publishes no IP list for Bytespider, so it cannot be verified from the request alone. Treat unusual volume as a scraper using its name, and rate-limit or block it at the firewall.

  3. 3

    Block what fails the check at the firewall. A robots.txt rule only reaches crawlers that choose to read it.

Frequently asked questions

ByteDance's crawler, widely reported as collecting data for its AI models. ByteDance publishes no documentation for it: no user-agent string, no IP list and no statement about robots.txt, so a robots.txt rule is the only request you can make, and a firewall rule the only enforcement.

Checked against ByteDance's documentation on 29 September 2026.

Next: see whether AI recommends you

Letting the right crawlers in is the first step. AskWatch shows what ChatGPT, Perplexity, Gemini and Google AI answer when buyers ask about your category, and who they name.

  • Free, no credit card.
  • Report in minutes, link sent to your email.