Skip to content

Robots.txt Generator

Add crawler rules per user-agent, block specific paths or specific bots like GPTBot, and attach your sitemap — with a warning before you accidentally ship Disallow: / to every crawler on the internet.

  • Per-bot rule groups
  • Catches Disallow: / for user-agent *
  • Sitemap line included
  • Notes what Google actually respects
  • Runs fully client-side

Generator

Built in your browser, never sent anywhere

    robots.txt

    User-agent: *
    Allow: /

    How to generate a robots.txt file

    1. 01

      Add a rule group

      Each group targets one or more user-agents — use * for every crawler, or name a specific one like Googlebot or GPTBot.

    2. 02

      Add Allow and Disallow paths

      Paths must start with a slash. Disallow blocks crawling of that path and anything under it; Allow carves out an exception.

    3. 03

      Add your sitemap

      A Sitemap line is not required, but it is the easiest way to point crawlers at the rest of your URLs.

    4. 04

      Read the warnings, then publish

      Fix anything flagged, copy the file, and upload it as /robots.txt at the root of your domain — it will not work from any other path.

    robots.txt controls crawling, not indexing#

    This is the single most common misunderstanding about the file. Disallowing a path tells well-behaved crawlers not to *fetch* it — it does not tell Google not to *index* it. A disallowed URL that is linked to from somewhere else can still show up in search results, just without a snippet, because Google never crawled the page to generate one.

    To actually keep a page out of search results, use a noindex meta tag or HTTP header on that page instead — and critically, do not also disallow it in robots.txt, because Google needs to crawl a page to see its noindex tag in the first place. Blocking crawling of a page you also want deindexed is a common way to accidentally keep an unwanted page indexed indefinitely.

    The mistake that deindexes an entire site#

    Disallow: / under User-agent: * blocks every well-behaved crawler from every page on the site. It is a two-character file that has taken down real production sites' search visibility — often left over from a staging environment's robots.txt that shipped to production by accident.

    This tool checks for exactly that combination and warns before you copy the output, but the check only catches what is visible in the config — always view the live file at /robots.txt after deploying, not just the version generated here.

    What crawlers actually respect Crawl-delay#

    Crawl-delay asks a crawler to wait a number of seconds between requests. Bing and Yandex honor it. Google ignores it entirely and has for years — to control Googlebot's crawl rate, the setting lives in Google Search Console instead, under crawl rate settings for the property.

    Setting Crawl-delay is harmless to include for the crawlers that respect it, but it will not do anything to Googlebot's behaviour, which surprises people who assume robots.txt is the one place all crawler behaviour is configured.

    Blocking AI crawlers specifically#

    Naming a specific user-agent, like GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (Google's AI training crawler, separate from Googlebot) or CCBot (Common Crawl), lets you block AI training crawlers without touching how search engines index the site — since Googlebot and GPTBot are different user-agents, a rule for one does not affect the other.

    This is a policy decision, not a technical guarantee: robots.txt is a voluntary standard, and it only stops crawlers that choose to respect it. It is the correct first step, but it is not enforcement.

    Frequently asked questions

    Does Disallow in robots.txt remove a page from Google?

    No. Disallow stops Google from crawling the page, not from indexing it — a disallowed page that is linked to elsewhere can still appear in search results without a snippet. To actually remove a page from search results, use a noindex tag on the page itself, and do not disallow that page at the same time, since Google needs to crawl it to see the noindex tag.

    What does Disallow: / under User-agent: * actually do?

    It blocks every well-behaved crawler from every page on the entire site. This is almost always a mistake — often a staging robots.txt that accidentally shipped to production. This tool flags the combination specifically because of how often it happens by accident.

    Does Google respect Crawl-delay?

    No, Google has ignored the Crawl-delay directive for years. Bing and Yandex do respect it. To control Googlebot's crawl rate specifically, use the crawl rate settings in Google Search Console instead.

    Can I block AI crawlers without blocking search engines?

    Yes — name the specific bot, such as GPTBot for OpenAI, ClaudeBot for Anthropic, or Google-Extended for Google's AI training crawler (a separate user-agent from Googlebot). A rule for one user-agent does not affect any other.

    Where does robots.txt need to be uploaded?

    At the root of the domain — https://example.com/robots.txt — and nowhere else. A robots.txt file placed in a subdirectory is not read by crawlers at all.

    Can I use this for a WordPress or Blogger site?

    Yes — most SEO plugins (Yoast, Rank Math) and Blogger both let you paste custom robots.txt content directly in their settings, rather than requiring you to upload a file over FTP. Generate it here and paste the result in.

    Is robots.txt actually enforced?

    No — it is a voluntary standard. Well-behaved crawlers like Googlebot respect it, but nothing stops a crawler from ignoring it entirely. It communicates intent to cooperative bots; it is not an access control mechanism.