robots.txt for AI crawlers: the practical guide
Every AI assistant that talks about your business first sent a bot to your website. robots.txt is where you decide what those bots may read. Most websites never made that decision on purpose.
Why this file suddenly matters
robots.txt has existed since 1994, but until recently only search engines read it. Now AI bots read it too. OpenAI, Anthropic, Google, Perplexity and Microsoft all check it before reading your pages, and what they find there decides whether AI assistants can learn about your business at all.
One thing to know before the details: the bots have different jobs. Blocking one doesn't make you "invisible to AI", and allowing one doesn't make you visible. There are three roles.
The AI bots and what each one does
| Bot | Operator | Role | If you block it |
|---|---|---|---|
| GPTBot | OpenAI | Training: builds the model's knowledge | Future models know less about you |
| OAI-SearchBot | OpenAI | Search: powers ChatGPT Search citations | ChatGPT stops citing your pages live |
| ChatGPT-User | OpenAI | User-fetch: opens your page when a user asks | Little effect, mostly ignores robots.txt |
| ClaudeBot | Anthropic | Training | Claude models know less about you |
| Claude-SearchBot | Anthropic | Search: citations in Claude's web search | Claude stops citing your pages |
| Claude-User | Anthropic | User-fetch | Little effect, mostly ignores robots.txt |
| PerplexityBot | Perplexity | Search index (does not train models) | You drop out of Perplexity answers |
| Google-Extended | Gemini training and Search grounding | Excluded from grounded Gemini answers | |
| Bingbot | Microsoft | Bing index, which also feeds ChatGPT's search | Weaker presence in Bing and ChatGPT Search |
Sources: the official bot documentation of OpenAI (developers.openai.com), Anthropic (support.claude.com), Perplexity (docs.perplexity.ai), Google (developers.google.com) and Bing, as of July 2026. Google's AI Overviews use the regular Googlebot, so they are governed by your normal search rules. User-fetch bots act on a person's direct request and largely operate outside robots.txt.
The setup most businesses want
If you want AI assistants to know and recommend your business, make sure nothing blocks the AI bots. Here is a minimal robots.txt that lets everyone in and points to your sitemap:
User-agent: *
Allow: /
Sitemap: https://www.example.com/sitemap.xml
If you don't want your content used for model training but still want to be cited in AI search, block the training bots and leave the search bots alone:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
# …everyone else (incl. search bots) stays allowed
User-agent: *
Allow: /
That trade-off is real, though: training is how models learn who you are. For most small businesses trying to win customers, blocking training bots costs more than it protects.
The four mistakes we keep finding
We benchmarked 419 Norwegian trade businesses with our audit methodology. The robots.txt failures fall into four patterns:
1. No robots.txt at all (7% of businesses). Not fatal, since no file means "everything allowed". But you lose the sitemap pointer, and nobody is making a choice.
2. A global Disallow: /. Usually a leftover from a staging environment. It tells every crawler, AI included, to read nothing. We see it rarely, but when we do, the business is invisible in every AI system at once.
3. Copy-pasted "block AI bots" lists. Templates that went around in 2023–24 block GPTBot "to protect content". They were written before AI search existed. In 2026 the result is that your competitors get cited and you don't. 4% of the benchmark still blocks GPTBot, mostly through templates like these.
4. The firewall blocks what robots.txt allows (1 in 5 businesses). This one is easy to miss. Your robots.txt says "welcome", but the firewall (WAF) or hosting provider returns 403 to requests with AI bot user-agents. In our third wave (28 September 2026), 18% of 419 websites blocked at least one of nine AI crawlers this way. In July, when we tested GPTBot only, it was 9% of 447. None of them knew. In one audit we found a firewall blocking OpenAI's training crawler while letting the search bots through. The site owner had never asked for that.
Test it in two minutes
Don't trust the config. Test what the server does. From any terminal, fetch your homepage as an AI bot and compare with a normal request:
curl -sI -A "OAI-SearchBot" https://www.example.com/ | head -1
curl -sI -A "ClaudeBot" https://www.example.com/ | head -1
curl -sI -A "PerplexityBot" https://www.example.com/ | head -1
All four should return 200, the same as a normal browser request. A 403 on any of them means your firewall is making AI policy for you. Our audit runs this exact test as one of its 35 checks.
Check the rest of your setup too
robots.txt is 1 of the 35 things we test. The free mini-audit checks your domain in under a minute, including whether AI bots can reach your pages at all.