AI Crawlers, robots.txt and llms.txt for SaaS
AI companies run different crawlers for different jobs. Blocking the wrong one can remove you from AI search while doing nothing about training. Bot names below were checked against each platform’s documentation.
Which AI crawlers should a SaaS site allow?
A SaaS site that wants to appear in AI answers should allow the search crawlers: OAI-SearchBot, Claude-SearchBot, PerplexityBot, and Googlebot and Bingbot as usual. Training crawlers like GPTBot, ClaudeBot and Google-Extended are a separate choice. Blocking them doesn’t remove you from AI search on those platforms, according to their documentation.
What’s the difference between search, training and user-triggered crawlers?
Search crawlers index pages so AI search can find and link them. Training crawlers collect content that may be used to train models. User-triggered fetchers visit a page because a person asked the assistant to, and some don’t follow robots.txt. Each platform documents its own names and rules.
| Company | Search retrieval | Model training | User-triggered fetching |
|---|---|---|---|
| OpenAI | OAI-SearchBot | GPTBot | ChatGPT-User: robots.txt rules may not apply |
| Anthropic | Claude-SearchBot | ClaudeBot | Claude-User: follows robots.txt, per Anthropic |
| Perplexity | PerplexityBot (not used for model training) | No separate bot documented | Perplexity-User: generally ignores robots.txt |
| Googlebot (Search, including AI Overviews and AI Mode) | Google-Extended token (Gemini training and some grounding) | Not covered here | |
| Microsoft | Bingbot (Bing, which reports Copilot citations) | Not covered here | Not covered here |
Checked September 26, 2026 against OpenAI’s crawler overview, Anthropic’s crawler help article and Google’s common crawlers, plus Perplexity’s crawler documentation. Names change; re-check before editing robots.txt.
What should robots.txt look like?
For most B2B SaaS sites, robots.txt should allow the search crawlers and make a deliberate choice about training crawlers. The example below allows AI search retrieval and opts out of training. OpenAI says each of its settings is independent, so blocking GPTBot doesn’t block OAI-SearchBot.
# Search retrieval: let AI search tools find and link your pages User-agent: OAI-SearchBot Allow: / User-agent: Claude-SearchBot Allow: / User-agent: PerplexityBot Allow: / # Model training: optional opt-out, a business decision User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: /
Google-Extended doesn’t affect inclusion or ranking in Google Search, per Google, but it does control use in Gemini and some grounding. Leave it allowed if appearing in Gemini matters more to you than opting out of training.
Should you add an llms.txt file?
An llms.txt file is optional. It is an emerging convention, proposed in September 2024, for a Markdown file at your site root that points AI tools to your most useful pages. It isn’t a requirement or a ranking factor: Google says you don’t need AI text files to appear in its AI features.
# Fabrikam Desk > Help desk software for B2B SaaS support teams of 2 to 20 agents. ## Product - [Pricing](https://fabrikam.example/pricing): plans, price per agent, what each includes - [HubSpot integration](https://fabrikam.example/integrations/hubspot): two-way sync details ## Comparisons - [Fabrikam Desk vs Litware Help](https://fabrikam.example/compare/litware)
The format is described in the proposal at llmstxt.org. If you add one, keep it short and current, and don’t expect it to change citations on its own.
Know which door you’re closing.
Search, training and user fetches are separate bots. Decide each one on purpose, and write the decision down.
How do you check what crawlers can reach?
Check three things: robots.txt rules for each bot name above, your CDN or firewall bot settings, which often block AI crawlers by default, and your server logs for the user agents. A CDN rule can block OAI-SearchBot even when robots.txt allows it.
If you are allowed but still missing from answers, work through why your SaaS isn’t showing up in AI answers. General crawl and indexing checks belong in the SEO for SaaS guide.
Frequently asked questions
Should I block GPTBot?
It’s a business choice. Blocking GPTBot opts out of training; OpenAI says it doesn’t stop OAI-SearchBot from showing your site in ChatGPT search.
Which bot lets my site appear in ChatGPT search?
OAI-SearchBot. OpenAI says sites that block it won’t be shown in ChatGPT search answers, though they can still appear as navigational links.
Does blocking Google-Extended remove me from AI Overviews?
Google says Google-Extended doesn’t affect inclusion or ranking in Google Search. AI Overviews and AI Mode are part of Search.
Do AI user agents respect robots.txt?
Search and training crawlers do, per their docs. User-triggered fetchers vary: OpenAI and Perplexity say theirs may not.
Is llms.txt required?
No. It is an optional, emerging convention. Google says you don’t need AI text files to appear in its AI features.
How long do robots.txt changes take?
OpenAI says about 24 hours for its systems to adjust. Other platforms don’t all publish a time.
Sources & further reading
Platform behaviour and crawler rules come from each platform’s documentation. Survey and study figures are named with their date on each page.