In short
Most AI crawlers read your site like a basic HTTP client: they check robots.txt for their own token, request the HTML and, according to third-party analysis, mostly do not run JavaScript. Whether they get your content depends on four layers you control: per-crawler robots.txt rules, your CDN or firewall, page directives such as noindex and nosnippet, and whether the text is in the server-rendered HTML. Search indexers and user-triggered fetchers decide whether you can be cited; training crawlers such as GPTBot and ClaudeBot can be opted out separately.
Which AI crawlers visit your site, and what does each one do?
Vendors now split their bots by job: training crawlers collect content for models, search indexers build the index an assistant searches, and user-triggered fetchers load a page because a user asked about it. Google-Extended and Applebot-Extended are only robots.txt control tokens and never fetch pages.
| Crawler | Operator | Purpose (vendor wording, condensed) | robots.txt token | Source |
|---|---|---|---|---|
| GPTBot | OpenAI | Content that may be used to train OpenAI's foundation models | GPTBot | OpenAI (2026) |
| OAI-SearchBot | OpenAI | Surfaces websites in ChatGPT search | OAI-SearchBot | OpenAI (2026) |
| ChatGPT-User | OpenAI | User actions in ChatGPT and Custom GPTs | None; rules “may not apply” | OpenAI (2026) |
| ClaudeBot | Anthropic | Web content for model training | ClaudeBot | Anthropic (2026) |
| Claude-SearchBot | Anthropic | Indexing to improve Claude's search results | Claude-SearchBot | Anthropic (2026) |
| Claude-User | Anthropic | Fetches pages when a user asks Claude | Claude-User | Anthropic (2026) |
| PerplexityBot | Perplexity | Perplexity search results; not used for training | PerplexityBot | Perplexity (2026) |
| Perplexity-User | Perplexity | Visits pages to answer a user's question | Perplexity-User (“generally ignores” robots.txt) | Perplexity (2026) |
| Googlebot | Google Search, including AI Overviews and AI Mode | Googlebot | Google (2025) | |
| Google-Extended | Control token: Gemini training and grounding in Gemini Apps and Vertex AI | Google-Extended | Google (2026) | |
| Bingbot | Microsoft | Bing index, which also powers Copilot | bingbot | Bing (2025) |
| Applebot | Apple | Spotlight, Siri and Safari search; may also train Apple models | Applebot | Apple (2026) |
| Applebot-Extended | Apple | Control token: opt out of Apple model training | Applebot-Extended | Apple (2026) |
| Meta-WebIndexer | Meta | Meta AI search results | meta-webindexer | Meta (2026) |
| Meta-ExternalAgent | Meta | Training foundation models or improving products | meta-externalagent | Meta (2026) |
| Meta-ExternalFetcher | Meta | Links fetched at a user's request; “may bypass robots.txt rules” | meta-externalfetcher | Meta (2026) |
| CCBot | Common Crawl | Common Crawl's open web archive | CCBot | Common Crawl (2026) |
Do AI crawlers obey robots.txt?
The automated crawlers do, by their vendors' own statements. For user-triggered fetchers the wording differs:
- OpenAI: GPTBot and OAI-SearchBot settings are “independent of the others”. For ChatGPT-User, “robots.txt rules may not apply” because a user initiated the action.
- Anthropic: its bots “respect ‘do not crawl’ signals by honoring industry standard directives in robots.txt”, Claude-User included. It also supports the non-standard
Crawl-delay. - Perplexity: PerplexityBot follows robots.txt; Perplexity-User “generally ignores robots.txt rules” since a user requested the fetch.
- Google: common crawlers “always obey robots.txt rules when crawling automatically”.
- Meta: Meta-ExternalFetcher “may bypass robots.txt rules”.
- Apple: Applebot respects robots.txt directives aimed at it in general search crawls.
Changes take time: OpenAI cites about 24 hours, Perplexity up to 24 hours. And robots.txt is a request, not a lock; RFC 9309 says its rules “are not a form of access authorization”. OpenAI, Anthropic, Perplexity, Apple and Common Crawl publish IP ranges for verification, and Google documents reverse-DNS patterns. Anthropic warns that blocking its IPs can stop it from reading your robots.txt, so robots.txt remains its documented opt-out.
How do you allow AI search but opt out of training?
Put search indexers and training crawlers in separate groups. This starting point keeps a site citable in AI answers while opting out of model training; replace the example paths with your own.
# AI search indexers and user-triggered fetchers: allowedUser-agent: OAI-SearchBotUser-agent: Claude-SearchBotUser-agent: Claude-UserUser-agent: PerplexityBotUser-agent: meta-webindexerDisallow: /cart/Disallow: /account/# Model training (crawlers and control tokens): opted outUser-agent: GPTBotUser-agent: ClaudeBotUser-agent: Google-ExtendedUser-agent: Applebot-ExtendedUser-agent: meta-externalagentUser-agent: CCBotDisallow: /# Everyone else, including Googlebot, Bingbot and ApplebotUser-agent: *Disallow: /cart/Disallow: /account/Sitemap: https://www.example.com/sitemap.xmlWhat are the caveats?
- Groups don't inherit. Under RFC 9309 a crawler follows the group naming it and uses the wildcard group only if none does, so repeat your exclusions in each group.
- Google-Extended is broader than training. It also covers grounding in Gemini Apps and on Vertex AI, but it does not affect Google Search, so it won't remove you from AI Overviews or AI Mode.
- Some crawlers are mixed-purpose. Applebot data may also train Apple's models (Applebot-Extended opts out), and Meta-ExternalAgent serves training or product improvement.
- Blocking Googlebot or Bingbot is not a training opt-out. It removes you from Google Search, including its AI features, or from the Bing index behind Copilot.
- User-triggered fetchers such as ChatGPT-User, Perplexity-User and Meta-ExternalFetcher may ignore these rules, so the file omits them.
- Keep the file reachable. RFC 9309 tells crawlers to assume complete disallow when robots.txt returns a server error (5xx).
- CCBot is a judgment call: it feeds a public archive whose reuse you don't control.
Can AI crawlers read JavaScript-rendered content?
Some can, and most vendors don't say. Google documents a crawl, render and index pipeline in which Googlebot runs JavaScript in “an evergreen version of Chromium”, yet still calls server-side or pre-rendering “a great idea” because “not all bots can run JavaScript”. Apple says Applebot “may render the content of your website within a browser”. The OpenAI, Anthropic and Perplexity crawler pages don't address rendering.
The best public evidence is third-party. In The rise of the AI crawler (December 2024), Vercel and MERJ analyzed traffic on Vercel's network and found that “none of the major AI crawlers currently render JavaScript”. GPTBot and Claude downloaded JavaScript files (11.50% and 23.84% of their requests) without executing them; Google's Gemini and AppleBot rendered pages.
That is one platform's snapshot from almost two years ago. The conclusion holds either way: put what you want quoted (product facts, prices, specs, answers, internal links) into the initial HTML via server-side rendering or static generation.
Which page-level directives keep content out of AI answers?
For Google, AI features inherit Search rules. A supporting link in AI Overviews or AI Mode “must be indexed and eligible to be shown in Google Search with a snippet”, with no additional requirements. The controls are noindex, nosnippet, max-snippet and data-nosnippet, set as a meta robots tag or an X-Robots-Tag header, which can target one crawler (X-Robots-Tag: googlebot: nofollow).
| Claim | What the source says | Source (year) |
|---|---|---|
nosnippet affects Google's AI features | It “will also prevent the content from being used as a direct input for AI Overviews and AI Mode” | Google Search Central (2026) |
max-snippet caps AI input | It “will also limit how much of the content may be used as a direct input” for both | Google Search Central (2026) |
Bing applies data-nosnippet to AI answers | Marked content is “excluded from snippets and AI summaries” | Bing Webmaster Blog (2025) |
| Google needs no special AI files | “You don't need to create new machine readable files, AI text files, or markup to appear in these features.” | Google Search Central (2025) |
The trade-off: these directives also cut your regular snippets, and Google documents no switch for its AI features alone. Check templates for leftovers such as a staging max-snippet:0. For how the two Google surfaces differ, see AI Mode vs AI Overviews.
Can a CDN or firewall block AI crawlers without you noticing?
Yes. robots.txt can allow a crawler that your edge then blocks or challenges. Cloudflare is the clearest case because its defaults just changed. As of September 2026, its documentation sorts AI bots by behavior: Search (indexing to answer questions later), Agent (real-time activity on a person's behalf, “such as chat fetch bots and browser-use agents”) and Training. Each can be blocked on all pages, blocked on pages with ads, or allowed.
Since September 15, 2026, new domains get Training and Agent blocked on pages that display ads, with Search allowed. Mixed-purpose crawlers that combine Search and Training are blocked by every configuration that blocks training, including the legacy Block AI bots setting. So on a new domain, a user-triggered fetch of an ad-supported page can fail by default. Review Security Settings → Configure AI bot policies (labels as of September 2026) and any custom firewall rules on whichever CDN you use, then test with real requests.
Does llms.txt help AI crawlers read your site?
There's no official evidence that it does, and it controls nothing. llms.txt is a proposal by Jeremy Howard, first published on September 3, 2024 and revised as v2 on August 10, 2026: a Markdown file at /llms.txt with the site name as H1, a short summary, and H2 sections listing key links. It is not a formal standard and neither grants nor denies access.
The proposal notes that OpenAI, Anthropic and Google's Gemini team publish llms.txt files for their developer docs, but publishing is not consuming. Google says no AI text files are needed for AI Overviews or AI Mode. None of the crawler documentation cited here mentions reading llms.txt, and as of September 2026 we found no official statement from OpenAI, Anthropic, Perplexity, Microsoft or Google committing to use it.
If you publish one anyway, keep it accurate and in sync with your sitemap, and never make it the only place important content lives.
What does this mean for you?
Access is the precondition for citations, not a guarantee (for what earns them, see which GEO techniques work). For access itself:
- Decide per purpose, not per vendor. Allow search indexers and user-triggered fetchers; decide on training crawlers separately.
- Test the path, not just the file. Request key pages with each crawler's user agent and re-check CDN bot settings after every security change.
- Server-render what should be quoted: pricing, product facts, docs and FAQs.
- Audit templates for
noindex,nosnippetandmax-snippet:0on pages you want cited. - Keep robots.txt fast and returning 200, list your sitemap, and allow a day for changes.
- Verify in logs by IP, not user agent. Anyone can send a GPTBot user-agent string.
- Treat llms.txt as optional housekeeping.
Example: fixing crawler access for a pricing page
Example (hypothetical; all numbers are illustrative): a B2B software company wants to be cited by ChatGPT, Claude and Perplexity but opt out of model training.
- Its robots.txt blocks GPTBot and ClaudeBot on purpose, and a crawlability check confirms OAI-SearchBot, Claude-SearchBot and PerplexityBot are allowed.
- The live fetch disagrees: with those three user agents,
/pricingreturns HTTP 403 while a browser gets 200. A firewall rule from a past scraping incident blocks most non-browser user agents. - The raw HTML of
/pricinghas 30 words and an empty#rootcontainer, so crawlers that skip JavaScript see no prices. - The docs template sets
max-snippet:0site-wide, left over from staging. - The team exempts the search crawlers' published IP ranges from the firewall rule, pre-renders
/pricing, removesmax-snippet:0and re-runs the check. - Two weeks of logs show OAI-SearchBot and PerplexityBot fetching
/pricingwith status 200 from IPs inside the vendors' ranges: the fix reached the real crawlers, not just a test.
How can you check crawler access with AutoSEO?
The AI crawlability check evaluates robots.txt for 25 AI and search crawler tokens (plus two SEO tool crawlers, shown but not scored), including the share of up to 300 sitemap URLs each may not fetch, and treats a 5xx robots.txt as a full block. It requests up to three key pages with each crawler's user agent and compares status, redirects and word count with a browser, which exposes firewall challenges. It also reads meta robots and X-Robots-Tag (including bot-specific tags such as gptbot), judges from the raw HTML whether pages depend on JavaScript, validates llms.txt, and checks sitemaps, canonicals and JSON-LD.
Findings include fixes, such as a robots.txt snippet that re-allows blocked search crawlers while keeping your exclusions, or an llms.txt draft built from your latest site audit. Checks can run weekly or monthly with score-drop alerts, or via run_crawlability_check on the MCP server. The score weighs search and user-triggered crawlers three times as much as training crawlers; llms.txt carries 10 of 100 points, so given the evidence above, weigh a missing file below a blocked search crawler.
The user-agent test runs from AutoSEO's servers, so a firewall that admits only verified crawler IPs may block the test but not the real crawler. Logs settle it: AI bot traffic imports access logs (nginx, Apache, Cloudflare Logpush, Akamai DataStream 2, NDJSON; up to 1 GB) or streams them via a Cloudflare Worker (bot traffic integrations, Cloudflare). Each hit is matched to a crawler and checked against the published IP ranges of OpenAI, Anthropic, Perplexity, Google, Microsoft and Apple, separating verified visits from likely spoofed ones, with status codes per path.
Frequently asked questions
Does blocking GPTBot remove my site from ChatGPT search?
No. OpenAI says GPTBot (training) and OAI-SearchBot (ChatGPT search) are controlled independently. Sites opted out of OAI-SearchBot “will not be shown in ChatGPT search answers, though can still appear as navigational links”.
Will blocking Google-Extended keep my content out of AI Overviews?
No. Google says Google-Extended does not affect inclusion in Google Search, and AI Overviews and AI Mode link to indexed, snippet-eligible pages. To limit what appears there, use nosnippet, data-nosnippet, max-snippet or noindex, which also affect regular snippets.
How can I tell whether a request claiming to be GPTBot or ClaudeBot is real?
Check the source IP against the vendor's published list, such as openai.com/gptbot.json or claude.com/crawling/bots.json; Google and Apple also document reverse-DNS checks. A matching user-agent string alone proves nothing.
How long do robots.txt changes take to reach AI crawlers?
OpenAI cites about 24 hours for its search systems, Perplexity up to 24 hours. RFC 9309 says crawlers should not use a cached robots.txt for more than 24 hours unless the file is unreachable. Plan for a day or more and confirm the change in your logs.
Should I still publish an llms.txt file?
It is optional. No major AI vendor has officially said its crawlers use it, and Google says no AI text files are needed for AI Overviews or AI Mode. If you publish one, keep it accurate and treat it as a convenience for agents, not as access control.
Sources
Primary sources used for this article, as available on the publication date.
- OpenAI: Overview of OpenAI Crawlers (2026)developers.openai.com
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler? (2026)support.claude.com
- Perplexity: Perplexity Crawlers (2026)docs.perplexity.ai
- Google Search Central: Google's common crawlers (2026)developers.google.com
- Google Search Central: AI features and your website (2025)developers.google.com
- Google Search Central: Robots meta tag, data-nosnippet, and X-Robots-Tag specifications (2026)developers.google.com
- Google Search Central: Understand JavaScript SEO basics (2026)developers.google.com
- Apple: About Applebot (2026)support.apple.com
- Bing Webmaster Blog: Bing Introduces Support for the data-nosnippet HTML Attribute (2025)blogs.bing.com
- Meta for Developers: Meta Web Crawlers (2026)developers.facebook.com
- Common Crawl: CCBot (2026)commoncrawl.org
- Vercel and MERJ: The rise of the AI crawler (2024)vercel.com
- Jeremy Howard: The /llms.txt file, llmstxt.org (2024, v2 2026)llmstxt.org
- Cloudflare Docs: Block AI Bots and AI bot policies (2026)developers.cloudflare.com
- IETF: RFC 9309, Robots Exclusion Protocol (2022)rfc-editor.org