robots.txt controls crawling, not indexing#
This is the single most common misunderstanding about the file. Disallowing a path tells well-behaved crawlers not to *fetch* it — it does not tell Google not to *index* it. A disallowed URL that is linked to from somewhere else can still show up in search results, just without a snippet, because Google never crawled the page to generate one.
To actually keep a page out of search results, use a noindex meta tag or HTTP header on that page instead — and critically, do not also disallow it in robots.txt, because Google needs to crawl a page to see its noindex tag in the first place. Blocking crawling of a page you also want deindexed is a common way to accidentally keep an unwanted page indexed indefinitely.
The mistake that deindexes an entire site#
Disallow: / under User-agent: * blocks every well-behaved crawler from every page on the site. It is a two-character file that has taken down real production sites' search visibility — often left over from a staging environment's robots.txt that shipped to production by accident.
This tool checks for exactly that combination and warns before you copy the output, but the check only catches what is visible in the config — always view the live file at /robots.txt after deploying, not just the version generated here.
What crawlers actually respect Crawl-delay#
Crawl-delay asks a crawler to wait a number of seconds between requests. Bing and Yandex honor it. Google ignores it entirely and has for years — to control Googlebot's crawl rate, the setting lives in Google Search Console instead, under crawl rate settings for the property.
Setting Crawl-delay is harmless to include for the crawlers that respect it, but it will not do anything to Googlebot's behaviour, which surprises people who assume robots.txt is the one place all crawler behaviour is configured.
Blocking AI crawlers specifically#
Naming a specific user-agent, like GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (Google's AI training crawler, separate from Googlebot) or CCBot (Common Crawl), lets you block AI training crawlers without touching how search engines index the site — since Googlebot and GPTBot are different user-agents, a rule for one does not affect the other.
This is a policy decision, not a technical guarantee: robots.txt is a voluntary standard, and it only stops crawlers that choose to respect it. It is the correct first step, but it is not enforcement.