Technical SEO
Robots.txt for SEO: Rules, AI Crawlers and Common Mistakes
Robots.txt tells crawlers which URLs they may fetch. It controls crawling, not indexing. To keep a page out of Google, use noindex instead, and never block the CSS, JS or pages you want to rank.
By SEORecheck Research Team, SEO auditorsPublished 9 min read
Robots.txt for SEO is a plain-text file at the root of your domain that tells search engine and AI crawlers which URLs they may request. It controls crawling, not indexing. Used well, it keeps bots out of low-value URLs. One wrong line, though, can hide your whole site from Google or from AI search tools like ChatGPT search.
This guide covers what Google actually supports, how to handle the growing list of AI crawlers, and the robots.txt mistakes that come up most often in real audits.
What does robots.txt actually do?
When a well-behaved crawler arrives at https://example.com, it first requests https://example.com/robots.txt and reads the rules for its user agent. If a URL is disallowed, the crawler skips it. That's all the file does. It doesn't remove pages from search results, it doesn't secure anything, and it doesn't pass ranking signals.
Google's documentation says this plainly: robots.txt "is not a mechanism for keeping a web page out of Google." A disallowed page "can still be indexed if linked to from other sites." It then shows up in results with no description, because Google was never allowed to read it. To keep a page out of the index, Google recommends noindex or password protection.
Some practical points follow from this:
- Each host needs its own file.
blog.example.comandexample.comare separate, and so arehttpandhttps. - The file must sit at the root.
/folder/robots.txtis ignored. - Path values are case-sensitive.
Disallow: /Admin/doesn't block/admin/.
Robots.txt syntax: what Google supports
According to Google Search Central, Google recognizes only four fields: user-agent, allow, disallow and sitemap. Other fields, including crawl-delay, aren't supported by Google, although some other crawlers do respect it.
A typical, healthy file for a small business site looks like this:
User-agent: *
Disallow: /cart/
Disallow: /checkout/
Disallow: /search
Disallow: /*?sort=
Allow: /wp-admin/admin-ajax.php
Disallow: /wp-admin/
Sitemap: https://example.com/sitemap.xml
How wildcards and conflicts work
*matches any sequence of characters.$marks the end of a URL. For example,Disallow: /*.pdf$blocks URLs that end in.pdf.- When an
allowrule and adisallowrule both match, Google uses the most specific rule, meaning the longest matching path. If they're equally specific, it uses the least restrictive one. - A crawler follows only the most specific
User-agentgroup that matches it. If you create aUser-agent: Googlebotgroup, Googlebot ignores everything underUser-agent: *. You have to repeat any shared rules inside the Googlebot group.
Limits, caching and server errors
| Situation | How Google handles it (per Search Central) |
|---|---|
| File larger than 500 KiB | Content beyond 500 KiB is ignored |
| You edit the file | Google generally caches it for up to 24 hours |
| robots.txt returns 4xx (except 429) | Treated as if no robots.txt exists, so everything may be crawled |
| robots.txt returns 5xx | Crawling stops for 12 hours, then Google uses the cached copy for up to 30 days |
noindex line inside robots.txt |
Not supported, so it's ignored |
The 5xx row is the one that catches people out. A server or CDN that errors on /robots.txt can pause crawling of your whole site, even when every page itself loads fine.
Should you block AI crawlers in robots.txt?
This question now comes up in almost every audit. Most AI companies run several separate bots: one collects training data, one indexes content for AI search, and one fetches pages live when a user asks something. Blocking the wrong one can quietly remove you from AI answers.
| Company | Bot | Purpose | Respects robots.txt? |
|---|---|---|---|
| Googlebot | Google Search, including AI features in Search | Yes | |
| Google-Extended | Token controlling Gemini training and grounding in Gemini Apps | Yes | |
| OpenAI | OAI-SearchBot | Surfacing sites in ChatGPT search | Yes |
| OpenAI | GPTBot | Training foundation models | Yes |
| OpenAI | ChatGPT-User | User-initiated fetches | Robots.txt rules "may not apply" |
| Anthropic | Claude-SearchBot | Indexing for search results in Claude | Yes |
| Anthropic | ClaudeBot | Model training | Yes |
| Anthropic | Claude-User | Fetching pages for user queries | Yes, per Anthropic |
| Perplexity | PerplexityBot | Surfacing and linking sites in Perplexity | Yes |
| Perplexity | Perplexity-User | User-initiated fetches | Generally ignores robots.txt |
Sources: Google's crawler documentation, OpenAI's bots page, Anthropic's support center and Perplexity's crawler docs, checked October 2026.
What blocking each bot actually changes
The vendors' own documentation explains the trade-offs:
- Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal," according to Google. Blocking it opts you out of Gemini training and grounding. It does not remove you from AI Overviews, which rely on regular Googlebot crawling. The robots.txt lever for those is blocking Googlebot, which would also remove you from Search. Snippet controls such as
nosnippetare the more precise tool. - OpenAI says each setting is independent. You can allow OAI-SearchBot so you appear in ChatGPT search answers and disallow GPTBot so your content isn't used for training.
- Anthropic warns that blocking Claude-SearchBot or Claude-User "may reduce your site's visibility" in Claude's search results.
For most businesses that want organic and AI visibility, a sensible default is to allow the search and user bots and make a deliberate business decision about the training bots:
# Opt out of model training, stay visible in AI search
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: *
Disallow: /cart/
Disallow: /checkout/
Sitemap: https://example.com/sitemap.xml
OAI-SearchBot, Claude-SearchBot and PerplexityBot have no group of their own here, so they fall back to User-agent: * and can crawl your content. Keep in mind that llms.txt is a separate file and has no effect on Google Search. Robots.txt is the file the crawlers actually check. For the wider picture of AI visibility, see our AI search optimization checklist.
Robots.txt vs noindex vs canonical: which should you use?
Each of these tools does a different job, and mixing them up is one of the most common technical SEO errors.
| Goal | Right tool | Why |
|---|---|---|
| Save crawl effort on infinite filters or internal search | robots.txt Disallow |
Stops the crawl requests themselves |
| Keep a page out of search results | noindex meta tag or X-Robots-Tag header |
Removes the page from the index, but only if Google can crawl it to see the tag |
| Consolidate duplicate URLs | rel="canonical" |
Passes signals to the preferred version (see our canonical tag guide) |
| Hide private content | Password or authentication | Robots.txt is public and anyone can read it |
Why robots.txt and noindex on the same page fails
When a page is disallowed in robots.txt, Google never fetches it, so it never sees the noindex tag. If other sites link to that URL, it can stay indexed indefinitely as a bare listing with no snippet. The fix may feel backwards, but it works: remove the disallow first, let Google crawl the page and process the noindex, then add the block back if you still need it. The same rule applies to canonicals. Google can't read a canonical tag on a page it isn't allowed to crawl.
The most common robots.txt mistakes
These are the problems that show up again and again in technical audits, roughly in order of how much damage they do.
Disallow: /left over from staging. A developer blocks the staging site and the file goes live with the launch. Traffic falls over the following days as Google stops recrawling. It's one of the first things to check in any website traffic drop diagnosis.- Blocking CSS, JavaScript or image folders. Old WordPress-era rules such as
Disallow: /wp-includes/orDisallow: /assets/stop Google from rendering the page properly. Google then sees a broken layout and may misjudge mobile-friendliness or miss content that JavaScript loads. - Blocking pages you want to rank. A broad rule like
Disallow: /productalso matches/products/and/product-guides/, because rules match URL prefixes. - A robots.txt that returns 5xx or times out. As the table above shows, this can halt crawling of the whole site.
- Blocking parameters that carry real content. If your category pages paginate with
?page=2, blocking/*?hides those products from discovery. - Accidentally blocking AI search bots. A blanket "block all AI" snippet copied from a forum often includes OAI-SearchBot, Claude-SearchBot or PerplexityBot, which removes you from those answer engines.
- Relying on unsupported directives.
Noindex:,Crawl-delay:(for Google) and comments in non-standard formats do nothing for Googlebot. - A relative sitemap URL. The
Sitemap:line must use a full absolute URL, such ashttps://example.com/sitemap.xml.
How do you test your robots.txt file?
Testing takes about 15 minutes and is worth repeating after every site release.
- Open the live file at
yourdomain.com/robots.txt. Check that it returns HTTP 200 and that the content is what you expect, and do this for every subdomain and protocol variant. - Check the robots.txt report in Google Search Console (Settings → robots.txt). It shows which robots.txt files Google found, when it last fetched them, and any fetch errors or parsing warnings. It replaced the old standalone robots.txt Tester.
- Run URL Inspection on your most important pages. If a page shows "Blocked by robots.txt," find the rule responsible.
- Cross-check against your sitemap. Any URL listed in your XML sitemap but disallowed in robots.txt sends Google contradictory signals and should be fixed.
- Review server logs or your CDN's bot settings. Some CDNs and security plugins block AI crawlers at the firewall, whatever robots.txt says. Confirm the bots you want are getting 200 responses.
If key pages appear under "Discovered – currently not indexed" alongside crawl restrictions, our guide to Discovered – Currently Not Indexed explains how crawl budget and blocked resources interact.
How SEORecheck checks robots.txt in an audit
Robots.txt problems are easy to miss by eye, because the file looks fine until you compare it with what's actually on your site. An automated audit fetches the file, then checks every crawled URL against it. It flags important pages that are blocked, blocked rendering resources, sitemap URLs that contradict the rules, and AI search crawlers that are shut out. Each finding lists the affected URLs and the exact rule causing it. You can see the format in the sample report.
Once you've fixed things, a recheck 30–60 days later shows which issues are fixed, improved, still unresolved or new. That helps when a later deploy quietly brings an old rule back. Here's an example comparison.
Key takeaways
- Robots.txt controls crawling, not indexing. Use
noindexor authentication to keep pages out of Google. - Google supports only
user-agent,allow,disallowandsitemap. It ignorescrawl-delayandnoindexin robots.txt. - A robots.txt that returns 5xx can pause crawling of your whole site, so monitor it like a critical page.
- Never disallow a page and expect Google to see its
noindexor canonical tag. - AI companies run separate training and search bots. Blocking GPTBot, ClaudeBot or Google-Extended doesn't remove you from ChatGPT search, Claude search or Google Search, but blocking OAI-SearchBot, Claude-SearchBot or PerplexityBot does remove you from those tools.
- Test the file after every release: check the live file, the Search Console robots.txt report and URL Inspection on your key pages.
Check your robots.txt as part of a full audit
If you aren't sure whether your robots.txt is blocking pages, resources or AI search crawlers that matter, a quick audit will tell you. Request a free SEO audit preview to get your SEO score and several real issues found on your own pages, with no account needed. For the full list of what an audit covers, see the SEO audit checklist.