Technical SEO

Robots.txt for SEO: Rules, AI Crawlers and Common Mistakes

Robots.txt tells crawlers which URLs they may fetch. It controls crawling, not indexing. To keep a page out of Google, use noindex instead, and never block the CSS, JS or pages you want to rank.

By SEORecheck Research Team, SEO auditorsPublished 9 min read

Robots.txt for SEO is a plain-text file at the root of your domain that tells search engine and AI crawlers which URLs they may request. It controls crawling, not indexing. Used well, it keeps bots out of low-value URLs. One wrong line, though, can hide your whole site from Google or from AI search tools like ChatGPT search.

This guide covers what Google actually supports, how to handle the growing list of AI crawlers, and the robots.txt mistakes that come up most often in real audits.

What does robots.txt actually do?

When a well-behaved crawler arrives at https://example.com, it first requests https://example.com/robots.txt and reads the rules for its user agent. If a URL is disallowed, the crawler skips it. That's all the file does. It doesn't remove pages from search results, it doesn't secure anything, and it doesn't pass ranking signals.

Google's documentation says this plainly: robots.txt "is not a mechanism for keeping a web page out of Google." A disallowed page "can still be indexed if linked to from other sites." It then shows up in results with no description, because Google was never allowed to read it. To keep a page out of the index, Google recommends noindex or password protection.

Some practical points follow from this:

  • Each host needs its own file. blog.example.com and example.com are separate, and so are http and https.
  • The file must sit at the root. /folder/robots.txt is ignored.
  • Path values are case-sensitive. Disallow: /Admin/ doesn't block /admin/.

Robots.txt syntax: what Google supports

According to Google Search Central, Google recognizes only four fields: user-agent, allow, disallow and sitemap. Other fields, including crawl-delay, aren't supported by Google, although some other crawlers do respect it.

A typical, healthy file for a small business site looks like this:

User-agent: *
Disallow: /cart/
Disallow: /checkout/
Disallow: /search
Disallow: /*?sort=
Allow: /wp-admin/admin-ajax.php
Disallow: /wp-admin/

Sitemap: https://example.com/sitemap.xml

How wildcards and conflicts work

  • * matches any sequence of characters. $ marks the end of a URL. For example, Disallow: /*.pdf$ blocks URLs that end in .pdf.
  • When an allow rule and a disallow rule both match, Google uses the most specific rule, meaning the longest matching path. If they're equally specific, it uses the least restrictive one.
  • A crawler follows only the most specific User-agent group that matches it. If you create a User-agent: Googlebot group, Googlebot ignores everything under User-agent: *. You have to repeat any shared rules inside the Googlebot group.

Limits, caching and server errors

Situation How Google handles it (per Search Central)
File larger than 500 KiB Content beyond 500 KiB is ignored
You edit the file Google generally caches it for up to 24 hours
robots.txt returns 4xx (except 429) Treated as if no robots.txt exists, so everything may be crawled
robots.txt returns 5xx Crawling stops for 12 hours, then Google uses the cached copy for up to 30 days
noindex line inside robots.txt Not supported, so it's ignored

The 5xx row is the one that catches people out. A server or CDN that errors on /robots.txt can pause crawling of your whole site, even when every page itself loads fine.

Should you block AI crawlers in robots.txt?

This question now comes up in almost every audit. Most AI companies run several separate bots: one collects training data, one indexes content for AI search, and one fetches pages live when a user asks something. Blocking the wrong one can quietly remove you from AI answers.

Company Bot Purpose Respects robots.txt?
Google Googlebot Google Search, including AI features in Search Yes
Google Google-Extended Token controlling Gemini training and grounding in Gemini Apps Yes
OpenAI OAI-SearchBot Surfacing sites in ChatGPT search Yes
OpenAI GPTBot Training foundation models Yes
OpenAI ChatGPT-User User-initiated fetches Robots.txt rules "may not apply"
Anthropic Claude-SearchBot Indexing for search results in Claude Yes
Anthropic ClaudeBot Model training Yes
Anthropic Claude-User Fetching pages for user queries Yes, per Anthropic
Perplexity PerplexityBot Surfacing and linking sites in Perplexity Yes
Perplexity Perplexity-User User-initiated fetches Generally ignores robots.txt

Sources: Google's crawler documentation, OpenAI's bots page, Anthropic's support center and Perplexity's crawler docs, checked October 2026.

What blocking each bot actually changes

The vendors' own documentation explains the trade-offs:

  • Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal," according to Google. Blocking it opts you out of Gemini training and grounding. It does not remove you from AI Overviews, which rely on regular Googlebot crawling. The robots.txt lever for those is blocking Googlebot, which would also remove you from Search. Snippet controls such as nosnippet are the more precise tool.
  • OpenAI says each setting is independent. You can allow OAI-SearchBot so you appear in ChatGPT search answers and disallow GPTBot so your content isn't used for training.
  • Anthropic warns that blocking Claude-SearchBot or Claude-User "may reduce your site's visibility" in Claude's search results.

For most businesses that want organic and AI visibility, a sensible default is to allow the search and user bots and make a deliberate business decision about the training bots:

# Opt out of model training, stay visible in AI search
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: *
Disallow: /cart/
Disallow: /checkout/

Sitemap: https://example.com/sitemap.xml

OAI-SearchBot, Claude-SearchBot and PerplexityBot have no group of their own here, so they fall back to User-agent: * and can crawl your content. Keep in mind that llms.txt is a separate file and has no effect on Google Search. Robots.txt is the file the crawlers actually check. For the wider picture of AI visibility, see our AI search optimization checklist.

Robots.txt vs noindex vs canonical: which should you use?

Each of these tools does a different job, and mixing them up is one of the most common technical SEO errors.

Goal Right tool Why
Save crawl effort on infinite filters or internal search robots.txt Disallow Stops the crawl requests themselves
Keep a page out of search results noindex meta tag or X-Robots-Tag header Removes the page from the index, but only if Google can crawl it to see the tag
Consolidate duplicate URLs rel="canonical" Passes signals to the preferred version (see our canonical tag guide)
Hide private content Password or authentication Robots.txt is public and anyone can read it

Why robots.txt and noindex on the same page fails

When a page is disallowed in robots.txt, Google never fetches it, so it never sees the noindex tag. If other sites link to that URL, it can stay indexed indefinitely as a bare listing with no snippet. The fix may feel backwards, but it works: remove the disallow first, let Google crawl the page and process the noindex, then add the block back if you still need it. The same rule applies to canonicals. Google can't read a canonical tag on a page it isn't allowed to crawl.

The most common robots.txt mistakes

These are the problems that show up again and again in technical audits, roughly in order of how much damage they do.

  1. Disallow: / left over from staging. A developer blocks the staging site and the file goes live with the launch. Traffic falls over the following days as Google stops recrawling. It's one of the first things to check in any website traffic drop diagnosis.
  2. Blocking CSS, JavaScript or image folders. Old WordPress-era rules such as Disallow: /wp-includes/ or Disallow: /assets/ stop Google from rendering the page properly. Google then sees a broken layout and may misjudge mobile-friendliness or miss content that JavaScript loads.
  3. Blocking pages you want to rank. A broad rule like Disallow: /product also matches /products/ and /product-guides/, because rules match URL prefixes.
  4. A robots.txt that returns 5xx or times out. As the table above shows, this can halt crawling of the whole site.
  5. Blocking parameters that carry real content. If your category pages paginate with ?page=2, blocking /*? hides those products from discovery.
  6. Accidentally blocking AI search bots. A blanket "block all AI" snippet copied from a forum often includes OAI-SearchBot, Claude-SearchBot or PerplexityBot, which removes you from those answer engines.
  7. Relying on unsupported directives. Noindex:, Crawl-delay: (for Google) and comments in non-standard formats do nothing for Googlebot.
  8. A relative sitemap URL. The Sitemap: line must use a full absolute URL, such as https://example.com/sitemap.xml.

How do you test your robots.txt file?

Testing takes about 15 minutes and is worth repeating after every site release.

  1. Open the live file at yourdomain.com/robots.txt. Check that it returns HTTP 200 and that the content is what you expect, and do this for every subdomain and protocol variant.
  2. Check the robots.txt report in Google Search Console (Settings → robots.txt). It shows which robots.txt files Google found, when it last fetched them, and any fetch errors or parsing warnings. It replaced the old standalone robots.txt Tester.
  3. Run URL Inspection on your most important pages. If a page shows "Blocked by robots.txt," find the rule responsible.
  4. Cross-check against your sitemap. Any URL listed in your XML sitemap but disallowed in robots.txt sends Google contradictory signals and should be fixed.
  5. Review server logs or your CDN's bot settings. Some CDNs and security plugins block AI crawlers at the firewall, whatever robots.txt says. Confirm the bots you want are getting 200 responses.

If key pages appear under "Discovered – currently not indexed" alongside crawl restrictions, our guide to Discovered – Currently Not Indexed explains how crawl budget and blocked resources interact.

How SEORecheck checks robots.txt in an audit

Robots.txt problems are easy to miss by eye, because the file looks fine until you compare it with what's actually on your site. An automated audit fetches the file, then checks every crawled URL against it. It flags important pages that are blocked, blocked rendering resources, sitemap URLs that contradict the rules, and AI search crawlers that are shut out. Each finding lists the affected URLs and the exact rule causing it. You can see the format in the sample report.

Once you've fixed things, a recheck 30–60 days later shows which issues are fixed, improved, still unresolved or new. That helps when a later deploy quietly brings an old rule back. Here's an example comparison.

Key takeaways

  • Robots.txt controls crawling, not indexing. Use noindex or authentication to keep pages out of Google.
  • Google supports only user-agent, allow, disallow and sitemap. It ignores crawl-delay and noindex in robots.txt.
  • A robots.txt that returns 5xx can pause crawling of your whole site, so monitor it like a critical page.
  • Never disallow a page and expect Google to see its noindex or canonical tag.
  • AI companies run separate training and search bots. Blocking GPTBot, ClaudeBot or Google-Extended doesn't remove you from ChatGPT search, Claude search or Google Search, but blocking OAI-SearchBot, Claude-SearchBot or PerplexityBot does remove you from those tools.
  • Test the file after every release: check the live file, the Search Console robots.txt report and URL Inspection on your key pages.

Check your robots.txt as part of a full audit

If you aren't sure whether your robots.txt is blocking pages, resources or AI search crawlers that matter, a quick audit will tell you. Request a free SEO audit preview to get your SEO score and several real issues found on your own pages, with no account needed. For the full list of what an audit covers, see the SEO audit checklist.

All articles