English

robots.txt Generator — Crawler Rules

Build a robots.txt file with crawler rules and a sitemap location

Runs in your browser · nothing is uploaded

Use * for all crawlers. This tool creates one user-agent group.

One path starting with / per line. URL-encode spaces and #. Wildcards * and $ are supported.

One path starting with / per line. URL-encode spaces and #. Wildcards * and $ are supported.

robots.txt asks crawlers to follow rules. It does not control access or guarantee removal from search results. Do not use it to hide secret paths.

robots.txt generated.

How to use

  1. Keep User-agent as * for all crawlers, or enter one crawler product token.
  2. Add disallowed paths, one per line. Use /private/ to request that a directory is not crawled.
  3. Add any allowed exceptions inside blocked paths.
  4. Optionally supply the full http(s) URL of your sitemap.
  5. Copy the output or download robots.txt, then publish it at the root of the relevant site.

Rules and placement

This generator creates one user-agent group. Paths start with / and are case-sensitive. A wildcard * can match multiple characters, while $ anchors a pattern at the end. URL-encode spaces and literal hash characters before entering a path.

An empty disallowed-path list produces an empty Disallow: directive. To request that the whole site not be crawled, enter / as a disallowed path. The usual location is https://example.com/robots.txt; a different host, port or protocol has its own scope.

Example

To block /private/ while leaving /private/public/ crawlable, put the parent path in Disallowed paths and the longer child path in Allowed exceptions. The more specific rule expresses the exception.

The pattern /*.pdf$ targets URLs ending in .pdf. A URL with a query string may not match the same end condition. Test representative URLs against the intended crawler’s implementation.

Limits and security

Each path list supports up to 1,000 lines and 100,000 characters. Duplicate paths are removed, and the optional sitemap must be a full http(s) URL without credentials or a fragment.

The generated file is a set of crawl requests. It neither enforces access control nor guarantees de-indexing. Its contents are public, and a crawler can ignore them. The tool does not upload the file, visit your website or verify search engine behavior.

See the Robots Exclusion Protocol, RFC 9309 for rule syntax and matching conventions.

FAQ

Can robots.txt protect private content?

No. The file is public and asks crawlers to follow instructions. It cannot replace authentication or server access controls, and it should not be used to conceal secret paths.

Will Disallow remove a page from search results?

Not necessarily. A blocked URL can still be discovered through links and may appear in results. Crawl restrictions and search indexing controls are separate concerns.

What happens when Allow and Disallow overlap?

Standards-compliant crawlers use the more specific matching path rule, with Allow preferred for an equally specific match. Support can vary, so check the target crawler's testing tools.

Missing a tool?