llms.txt checker and validator
Enter a domain and see whether the file is there, whether it follows the spec, whether its links answer, and whether robots.txt lets AI crawlers read it. Or paste a file to check it before you publish.
- Free, no account.
- A few seconds.
- Nothing is stored.
Who reads llms.txt, and what for
Agents and tools that read a site on someone's behalf, and the audits that check you are ready for them.
Coding agents
Claude Code, Cursor, Windsurf
A developer working with a library or an API points the agent at its llms.txt, and the agent finds the right docs page instead of guessing. In Ahrefs' 2026 study, Claude Code requested llms.txt files more often than any AI search crawler.
Documentation tools
MCP doc servers, IDE doc indexes
Tools that load a product's documentation into an agent use the file as the index: the list of pages to fetch, with a line on each. LangChain's mcpdoc server, for one, serves llms.txt files to coding agents.
People working with AI assistants
ChatGPT, Claude, Perplexity
Pasting a site's llms.txt into a chat is the fastest way to give an assistant the whole site in one go: what it is, and which page says what.
Audits
Lighthouse, SEO audit tools
Lighthouse checks for the file in its Agentic Browsing category, and SEO audit tools now flag a missing or broken one. In Ahrefs' study they were the largest group of readers.
Published on their sites or developer docs by
- Stripe
- Vercel
- Mintlify
- Anthropic
- OpenAI
- Cloudflare
Yoast SEO, Rank Math, Mintlify, GitBook and Docusaurus generate it for their sites, and an Ahrefs study of 137,000 domains in 2026 found a valid file on 28% of them.
What the llms.txt validator checks
Four groups, sixteen checks. The spec is short and most of the thresholds are not in it, so each card says where its number comes from.
Delivery
Answers 200 at /llms.txt, served as text/plain or text/markdown, under 50 KB (a warning past that, a failure past 256 KB, the thresholds validators settled on since the spec sets none).
Structure
Opens with one H1, carries a blockquote summary, lists pages in H2 sections as markdown link items, and has no HTML, no duplicate URLs and no headings deeper than H2.
Links
The first forty links are probed. A 404 or a host that does not resolve is dead; a 403 from a bot wall or a timeout is reported, not failed.
Discovery
robots.txt does not block GPTBot, ClaudeBot, PerplexityBot and the rest from the file; whether llms-full.txt exists; whether the home page points at the file with a markdown alternate link.
What a good llms.txt looks like
Six habits of the files that are worth an agent's time.
- 1
Say what the site is in the first two lines
The H1 names it, the blockquote says what it does and for whom. An agent that reads nothing else should still know.
- 2
Lead with the pages people ask about
Pricing, the product, the docs, comparisons: the pages that answer the questions buyers and developers bring. Legal pages and tag archives go last or in Optional.
- 3
One plain line per page
What the reader finds there, in words a person would use. A page title repeated as its own description tells an agent nothing.
- 4
Absolute links that answer 200
https:// addresses, no redirects, no pages behind a login. A dead link in the index is a dead end for the agent.
- 5
An index, not the content
Keep it under 50 KB. The full text belongs on the pages, or in an llms-full.txt beside it.
- 6
Built from the site, so it stays true
A file written by hand once goes stale the day a page is renamed. Generate it from the same data the pages come from, or regenerate it when they change.
llms.txt, robots.txt and sitemap.xml
Three files at the root of a site, three different readers. You want all three, and none replaces another.
llms.txt
- Written for
- Language models and AI agents
- What it says
- What the site is and which pages matter, in words
- Format
- Markdown
- Where it lives
- /llms.txt
- Standard
- llmstxt.org, 2024
- What it changes
- How easily an agent reads the site
robots.txt
- Written for
- Every crawler
- What it says
- What may and may not be fetched
- Format
- Plain-text directives
- Where it lives
- /robots.txt
- Standard
- RFC 9309
- What it changes
- Who may crawl what
sitemap.xml
- Written for
- Search engine crawlers
- What it says
- Which URLs exist and when they changed
- Format
- XML
- Where it lives
- /sitemap.xml, named in robots.txt
- Standard
- sitemaps.org
- What it changes
- How fast new pages are found
Frequently asked questions
A plain-text markdown file at the root of a site (/llms.txt) that tells a language model what the site is and which pages matter. It opens with an H1 naming the site, an optional one-line summary in a blockquote, and then H2 sections listing pages as markdown links with a short note each. The format was proposed by Jeremy Howard in September 2024 and is documented at llmstxt.org.
llms.txt, with an s, at the root of the site: yourdomain.com/llms.txt. llm.txt is a common misspelling, and a file published under that name is one that tools and agents looking for llms.txt will not find. If you have one at /llm.txt, rename it or redirect it.
Four groups. Delivery: the file answers 200 at /llms.txt, is served as text/plain or text/markdown, and is a sensible size. Structure: it opens with a single H1, has a blockquote summary, uses H2 sections of link items, and has no HTML, duplicate URLs or headings deeper than H2. Links: the first forty links answer. Discovery: robots.txt does not block AI crawlers from the file, and whether llms-full.txt and a markdown alternate link exist.
Publishers include Stripe, Vercel and Mintlify on their main sites and Anthropic, OpenAI and Cloudflare on their developer docs, and an Ahrefs study of 137,000 domains in 2026 found a valid file on 28% of them. Readers are coding agents such as Claude Code and Cursor, documentation tools that load a product's docs into an agent, Lighthouse's Agentic Browsing audit, and SEO audit tools.
Google Search does not use it as a ranking signal, and Google has said so. The file is for a different reader: AI agents and assistants that read a site on someone's behalf. Google's own Lighthouse does check it, in the Agentic Browsing category added in version 13.3.
Since version 13.3, Lighthouse has an Agentic Browsing category with an llms.txt audit. A 404 is reported as not applicable, so a missing file does not fail the audit. A server error, or a file that cannot be parsed, does. This checker separates those two cases the same way.
Because robots.txt has a rule that keeps one or more AI crawlers away from /llms.txt, often a blanket Disallow: / for GPTBot or ClaudeBot. A crawler that may not fetch the file does not have it, however well written it is. The check names the crawlers and the line in robots.txt.
A companion file, by convention, that holds the full text of every page rather than a list of links. Documentation platforms such as Mintlify publish it automatically. It is not part of the spec, so its absence is noted, not failed.
Not in the spec. Validators settled on warning past 50 KB and failing past 256 KB, because the file is meant to be an index an agent reads before it decides what to fetch, and an index that fills the agent's context defeats that. This checker uses the same thresholds and says so.
Yes. Paste mode runs the structure checks on the text you give it, which is useful before the file is live. It cannot probe links, robots.txt or the home page, because there is no site to ask, and it says which checks it skipped.
Cloudflare with default settings serves /llms.txt and robots.txt without a challenge; Bot Fight Mode, Under Attack mode and Imperva usually answer with a challenge page instead, and the checker says so rather than judging a file it never saw. That does not mean AI crawlers are blocked: these firewalls let verified bots such as GPTBot and ClaudeBot through by IP, which is where to look. To check the file itself, paste it. We do not ask you to allow our crawler; a user-agent allow rule is one anyone can claim.
No. The result of a domain check is cached for ten minutes so that a link to it does not fetch your site again on every open; after that it is gone. Pasted text is checked in the request and not kept. There is no account and nothing to delete.
The checker reads the file you have. The generator writes one from your sitemap and page descriptions. If the checker finds no file, it offers to generate one for the same domain; the generator runs its output through the same structure checks before handing it to you.
Other free tools
- Open tool
AI Visibility Checker
Check if ChatGPT, Gemini, Perplexity and Google AI name your brand, and get your AI visibility score.
- Open tool
AI Overview Checker
See if Google AI Overviews cite your site, which of your pages they pick, and who is cited instead.
- Open tool
Perplexity Visibility Tracker
Track whether Perplexity cites your site and names your brand in its answers.
- Open tool
ChatGPT Visibility Tracker
See if ChatGPT mentions your brand and where you rank in its answers.
- Open tool
Gemini Visibility Tracker
Check if Google Gemini mentions your brand and who it names instead.
- Open tool
Google AI Mode Visibility Tracker
See if Google AI Mode names your brand and which pages it cites.
- Open tool
Copilot Visibility Tracker
See if Microsoft Copilot mentions your brand and cites your site.
- Open tool
llms.txt Generator
Write an llms.txt for any site from its sitemap and page descriptions. No model, no watermark, checked against the spec.
- Open tool
robots.txt Generator
Create a robots.txt in a minute: block AI training, stay in AI search, add paths and your sitemap. Tested before you download it.
- Open tool
robots.txt Tester
Test any URL against robots.txt for Googlebot and 20+ AI crawlers, and find the line that decides each one.
- Open tool
AI Crawlers List
Every AI crawler with its user agent, IP list and robots.txt token, and whether your site lets each one in.
- Open tool
Google Search Console MCP
Ask Search Console in plain language from Claude, ChatGPT, Cursor or Claude Code. No Google Cloud project.
Next: see whether AI recommends you
llms.txt helps AI read your site. AskWatch shows what ChatGPT, Perplexity, Gemini and Google AI answer when buyers ask about your category, and who they name.
- Free, no credit card.
- Report in minutes, link sent to your email.