Robots.txt Generator
Build a robots.txt with per-crawler rules and sitemap entries.
Upload to the root of your domain, at /robots.txt. Remember that Disallow prevents crawling, not indexing — to remove a page from results, allow crawling and use a noindex meta tag instead.
About this tool
Assemble a robots.txt with rules per crawler, sitemap references and optional crawl delay. Presets cover the usual patterns — block nothing, block everything, or block the common CMS admin paths.
How to use it
- Start from a preset or an empty file.
- Add allow and disallow rules, per user-agent if needed.
- Add your sitemap URL, then copy the file to your site root.
Disallow does not mean noindex
This is the single most misunderstood thing in technical SEO, and it causes real damage.
Disallow prevents crawling. It does not prevent indexing.
If another site links to a URL you have disallowed, Google can still index it — it just cannot see the content. You get a result in the search listing with your URL, no title worth reading, and the description "No information is available for this page."
Worse, because Google cannot crawl the page, it cannot see a noindex tag on it either. Blocking a page you want removed actively prevents the mechanism that would remove it.
To keep a page out of the index: allow crawling, and add <meta name="robots" content="noindex">. Once it has dropped out, you may then block crawling if you want to save crawl budget.
Use robots.txt for: crawl budget on large sites, keeping crawlers out of faceted-search URL explosions, and blocking infinite calendar pages.
Syntax rules and the AI crawler question
The file must live at the domain root — example.com/robots.txt. Nowhere else is read, and it applies per subdomain and per protocol, so blog.example.com needs its own.
Only one group applies per crawler. Googlebot finds the most specific User-agent block matching it and ignores every other group, including *. If you write a Googlebot block, it must repeat every rule you wanted to apply — they are not inherited.
Order does not matter; specificity does. For conflicting rules, the longest matching path wins, and Allow beats Disallow when equally specific.
Wildcards are supported by the major engines: * for any sequence, $ to anchor the end. Disallow: /*.pdf$ blocks PDFs.
On AI crawlers — GPTBot, CCBot, ClaudeBot, Google-Extended and others honour robots.txt, and blocking them is a genuine choice with a trade-off: less training use of your content, but also less visibility in AI answers that increasingly sit above search results. Decide deliberately rather than copying someone's block list.
It is advisory. Scrapers ignore it entirely, and the file publicly lists the paths you would rather people not visit. Never use it as a security control.
Frequently asked questions
- Does disallow remove a page from Google?
- No — and this is the most common robots.txt mistake. Disallow prevents crawling, but a blocked URL can still be indexed from external links, showing with no description. To remove a page, allow crawling and use a noindex meta tag.
- Where does the file go?
- The root of the domain, exactly at /robots.txt. Crawlers do not look anywhere else, and it applies per subdomain — a separate file is needed for each.
- Do all crawlers obey it?
- The major search engines do. It is a voluntary convention, so scrapers and malicious bots routinely ignore it. Never use it to protect sensitive paths — it advertises them.