AI crawlers and robots.txt: should your store block them?
Last updated:
If your robots.txt blocks AI crawlers, assistants such as ChatGPT, Claude and Perplexity cannot read your product pages, so they have little reason to recommend your store. Many stores block them by accident, through a copied template or a security plugin. Checking takes two minutes.
How AI crawlers read robots.txt
Well-behaved crawlers fetch https://your-store.com/robots.txt before reading anything else and follow the group that matches their name. If there is no group for them, they fall back to the User-agent: * group. That fallback is the usual cause of accidental blocking.
The crawlers worth knowing
Names change over time, so check each provider's documentation. Commonly seen ones include:
- GPTBot, OAI-SearchBot, ChatGPT-User (OpenAI)
- ClaudeBot, anthropic-ai (Anthropic)
- PerplexityBot (Perplexity)
- Google-Extended and Applebot-Extended (control whether Google and Apple use your content for their AI products)
- CCBot (Common Crawl), Bytespider, Amazonbot
Some are for training data, others fetch pages live to answer a question. You may want different rules for each type.
Allowing AI crawlers
To let them read your public pages while keeping private areas off limits:
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: PerplexityBot
Allow: /
Disallow: /cart
Disallow: /checkout
Disallow: /account
User-agent: *
Disallow: /cart
Disallow: /checkout
Disallow: /account
Sitemap: https://your-store.com/sitemap.xml
Allowing one bot does not unblock the others: each needs its own group or a permissive * group.
When blocking makes sense
Some merchants deliberately opt out of AI training, for example to protect original photography or editorial content. That is a legitimate business decision. The mistake is blocking by accident, or blocking everything when you only meant to block training bots. If you want to be cited, allow the search and browsing agents at least.
How to check
- Open
/robots.txtand look forDisallow: /underUser-agent: *or under any AI bot name. - Check that your platform or a plugin has not added its own AI-blocking rules.
- Recheck after theme, plugin or hosting changes, and after launching from a staging site.
Pair this with a clear llms.txt file: blocking crawlers while publishing llms.txt sends contradictory signals.
How Averifly helps
Averifly's AI_CRAWLERS_BLOCKED rule reads your robots.txt and flags when major AI crawlers are disallowed, alongside the llms.txt and structured-data checks in our GEO rules. Run a free scan to see whether AI assistants can read your store.
Frequently asked questions
Does blocking GPTBot remove my store from ChatGPT?
Blocking GPTBot tells OpenAI not to use your pages for model training. Other agents, such as OAI-SearchBot and ChatGPT-User, have separate roles in search and browsing, so you can treat them differently.
Can a blanket Disallow rule block AI crawlers by accident?
Yes. A `User-agent: *` group with `Disallow: /` blocks every bot that has no group of its own, including AI crawlers. This often happens when a staging robots.txt is copied to production.
Is allowing AI crawlers safe for my store?
Allowing them means they can read your public pages, the same pages any visitor sees. Keep private areas such as the cart, checkout and account pages disallowed for every bot.
