Playbook · methodology
Fix AI crawler access
If AI crawlers can't fetch or read your pages, no content strategy will help. This playbook runs the technical checks first, fixes what blocks AI engines and verifies the result with real crawler visits.
- Goal
- AI crawlers can fetch and read your key pages
- For
- SEO, web and platform teams
- Baseline
- One crawlability check plus 7 days of tracking
- Re-measure
- After every fix and every release
Goal
Make sure the crawlers behind AI search and assistants — such as OAI-SearchBot, ChatGPT-User, Claude-SearchBot, PerplexityBot and Googlebot — are allowed where you want them, receive the same content as browsers and find your important pages. This is the precondition for every other playbook.
Crawlers serve different purposes: search crawlers index pages for AI search, user agents fetch a page when someone asks about it, and training crawlers such as GPTBot and ClaudeBot collect data for future models. You can treat each group differently.
Who it's for
- SEO teams who suspect technical reasons for low AI visibility
- Web and platform teams running CDNs, firewalls or bot protection
- Sites built as JavaScript single-page apps
- Anyone who changed robots.txt or bot settings and wants to verify the effect
Prompt set
Ten retrieval prompts: factual questions that engines can only answer correctly if they can read your pages. They are the practical test of crawler access.
Import it in AutoSEO under AI Visibility → Tracker → Import CSV (paste the text or upload the file).
Replace these placeholders before you import — the importer takes every line literally:
[brand]— your brand name[product]— a product or plan[feature]— a feature documented on your site[domain]— your domain, e.g. “example.com”
prompt,tags"What does [brand] cost?","Retrieval|Pricing""Does [brand] offer [feature]?","Retrieval|Product""What is [product] by [brand]?","Retrieval|Product""What integrations does [brand] support?","Retrieval|Product""Who is [brand] for?","Retrieval|Positioning""What is [brand]'s refund or cancellation policy?","Retrieval|Policy""How do I contact [brand] support?","Retrieval|Support""What is on the pricing page of [domain]?","Retrieval|Pricing""What is new at [brand]?","Retrieval|Freshness""Where is [brand] based and who runs it?","Retrieval|Company"Engines to track
Track the engines whose crawlers you are fixing, so the effect can show up in answers.
| Engine | Crawlers involved |
|---|---|
| ChatGPT | OAI-SearchBot indexes for search, ChatGPT-User fetches pages on request, GPTBot collects training data. |
| Claude | Claude-SearchBot, Claude-User and ClaudeBot (training) — each can be allowed or blocked separately. |
| Perplexity | PerplexityBot indexes pages; Perplexity-User fetches them when someone asks. |
| Google AI Overviews, AI Mode and Gemini | Rely on Googlebot's index. Google-Extended is a robots.txt token without its own crawler that governs use for Gemini models; it does not remove you from AI Overviews. |
Baseline measurement in AutoSEO
- 1
Run the crawlability check
Optimizations → Crawlability checks robots.txt rules per AI crawler, crawl delay, whether crawler user agents get the same pages as browsers, server-side rendering, meta robots and X-Robots-Tag, sitemaps, llms.txt, structured data and canonicals.
- 2
Decide your crawler policy
Decide per purpose — search, user requests, training — which crawlers you allow. Write the decision down before you change anything.
- 3
Connect bot traffic
Under Analytics → Bot Traffic, connect Cloudflare or Akamai, upload server logs or send them via webhook. Visits are checked against the IP ranges crawler operators publish.
- 4
Track the retrieval prompts
Import the prompt set and track it daily for a week. Note which questions each engine answers correctly from your pages.
Actions
- 01
Unblock the crawlers you want
Remove disallow rules for the AI search and user crawlers you decided to allow. Look out for broad rules like
User-agent: *withDisallow: /and for rules left over from staging. - 02
Check CDN and firewall rules
Bot protection and WAF rules often block AI crawlers before robots.txt even matters. Allow verified crawlers explicitly, then re-run the check.
- 03
Serve content in the HTML
Many crawlers don't execute JavaScript and only see the raw HTML. Render key text, headings and links on the server.
- 04
Remove accidental noindex and noai
Check meta robots and X-Robots-Tag on key pages. noindex, noai or nosnippet can keep content out of answers.
- 05
Publish sitemaps and an llms.txt
A complete XML sitemap helps crawlers discover pages. llms.txt is optional and not used by every engine, but it is cheap to provide.
- 06
Add structured data and clean canonicals
JSON-LD on the homepage and key pages and unambiguous canonical URLs make it clear which page states which fact.
How to measure change
- Crawlability findings resolved, re-checked after each fix
- Verified visits per AI crawler under Bot Traffic, compared week over week
- Share of retrieval prompts answered correctly, per engine
- Citations of your own pages in the retrieval prompts' answers
- A new check after every release, because deployments often reintroduce blocks
Expect a delay
Crawlers revisit on their own schedule. It can take days or weeks until an unblocked page is fetched, and longer until it shows up in answers.
Expected variance
The technical checks are deterministic: a robots.txt rule either blocks a crawler or it doesn't. The effect on answers is not. Whether an engine fetches your page for a question depends on its retrieval, and training crawlers only influence future model versions.
Crawler visits also fluctuate with crawl budgets and your own publishing activity. Compare weekly totals, not individual days.
Limitations
- Allowing crawlers makes you eligible, not visible. Content and authority still decide whether engines use your pages.
- robots.txt is a request, not access control. Crawlers that ignore it have to be stopped at the CDN or firewall.
- Visits from crawlers whose operators publish no IP ranges can't be verified.
- Blocking training crawlers today does not remove your content from models that already exist.
Questions about this playbook
Should I block AI training crawlers?
That is a business decision, not a technical one. Blocking training crawlers such as GPTBot or ClaudeBot keeps your content out of future training data but does not affect content already used. Search and user crawlers such as OAI-SearchBot, ChatGPT-User, Claude-SearchBot and PerplexityBot let engines fetch your pages to answer questions — blocking them usually costs visibility.
Does llms.txt improve AI visibility?
No major engine has confirmed that it uses llms.txt for ranking or citations. The file is cheap to provide and helps some AI agents navigate your site, so AutoSEO checks it — but fix robots.txt, rendering and meta directives first.
Why would crawlers get different pages than browsers?
Usually because of bot protection: CDNs and firewalls serve challenges, blocks or stripped pages to unknown bots. AutoSEO fetches your pages with real crawler user agents and compares the result with a browser request, so you can see the difference.
How do I know a visit really came from OpenAI or Anthropic?
User agents can be faked. AutoSEO checks crawler visits against the IP ranges that OpenAI, Anthropic, Perplexity, Google and other operators publish, and separates verified from unverified visits.
Do I need a developer for these fixes?
For robots.txt and CDN settings, often not. Rendering problems, meta directives and structured data usually need your web team. Turn the findings into tasks and push them to Jira, Linear or your project tool from Optimizations → Tasks.
Run this playbook in AutoSEO
Every feature is included in the free self-hosted edition and in AutoSEO Cloud.