SEO

Robots.txt Generator

Generate a valid robots.txt with crawler-specific rules, sitemap references and AI crawler directives, with each rule explained.

Last reviewed by the Radiatus Cloud team

Robots.txt Generator

Create robots.txt files to control crawler access.

* for all bots

Need this done properly for your business?

Radiatus delivers secure cloud, DevOps & compliance engineering.

Book a free consult

What robots.txt does and does not do

robots.txt controls crawling, not indexing. A disallowed URL can still appear in search results if other pages link to it, because the crawler never fetched the page to discover a noindex tag. That is the single most misunderstood aspect of the file. To keep a page out of the index, allow crawling and serve noindex; to keep it out entirely, require authentication.

Matching rules

Directives are prefix matches. Disallow: /api blocks /api, /api/v1 and also /api-documentation, which is rarely intended. Add the trailing slash when you mean a directory. Two wildcards are widely supported: * matches any sequence and $ anchors the end, so Disallow: /*.pdf$ blocks PDFs specifically. Where an allow and a disallow both match, the more specific rule wins, and ties go to allow.

User-agent groups do not merge

A crawler obeys exactly one group: the most specific one matching its name. If you write a User-agent: * group with several disallows and then a User-agent: Googlebot group with one rule, Googlebot follows only that one rule and ignores the wildcard group entirely. Every directive you want a named crawler to obey has to be repeated inside its own group.

AI crawlers

Retrieval crawlers such as GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot fetch pages so assistants can cite them, which is a distribution channel. Bulk corpus crawlers such as CCBot mainly feed training sets and send no traffic back. Deciding these separately is reasonable. Note that Google-Extended governs Gemini training only; it has no effect on AI Overviews or on ordinary Search, both of which follow Googlebot.

Practical rules

The file must sit at the domain root and is per host and per protocol, so subdomains need their own. Reference your sitemap with an absolute URL. Never list admin paths you actually want hidden, since robots.txt is public and reads as a map of where to look. And test before deploying: a stray Disallow: / deindexes an entire site, and it has happened to very large companies.

Related tools

  • Meta Tag Analyzer — Check a page's meta tags for the problems that actually affect indexing and click-through.
  • Redirect Chain Checker — Trace the full path of redirects (301/302) to find loops or lost link juice.
  • Robots.txt Validator — Validate robots.txt syntax and check whether a specific URL is allowed or blocked for a given crawler.
  • Sitemap Generator (Lite) — Generate a valid XML sitemap with lastmod dates and correct structure, and learn which URLs belong in it and which do not.

Frequently Asked Questions

Does robots.txt keep a page out of Google?

No. It prevents crawling, not indexing. A blocked URL can still be listed if other pages link to it, and because the crawler cannot fetch it, any noindex tag on the page is never seen. Allow crawling and serve noindex instead.

Why did Disallow: /api block another page?

Directives are prefix matches, so /api also matches /api-documentation and /apinews. Add a trailing slash when you mean a directory, and use the dollar anchor when you mean an exact path.

Do rules for a named crawler add to the wildcard group?

No. Each crawler obeys only the single most specific matching group. If you create a Googlebot group, Googlebot ignores the wildcard group entirely, so every rule it should follow must be repeated there.

Should I block AI crawlers?

Consider retrieval and training separately. GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot fetch pages to cite them, sending traffic back. CCBot mainly builds training corpora. Google-Extended affects Gemini training only, not AI Overviews or Search.

Can I hide sensitive paths with robots.txt?

No, and listing them makes matters worse. The file is publicly readable, so naming an admin path advertises it. Anything genuinely sensitive needs authentication or network-level restriction.

Privacy & Security

Generated locally in browser.

Data: None
Client-side-Side
Active
v1.0

About This Tool

This tool runs entirely in your browser. No data is sent to any server, ensuring complete privacy. Simply use the interface above to get started — no registration or login required.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.