How to write a robots.txt file (and what it cannot do)
7 min read · Updated 4 October 2026
A robots.txt file is a short text file at the root of a website that tells automated crawlers which parts of the site they may fetch. Search engines, archive services and many other bots read it before crawling. It is one of the oldest conventions on the web, and since 2022 it has been formally described in an internet standard (RFC 9309).
The file is simple, but a single wrong character can hide a whole site from search engines, or fail to hide what you meant to. This guide explains how the rules are actually matched, gives examples you can adapt, and covers the things robots.txt was never designed to do.
Where the file lives
The file must be at exactly /robots.txt on each host:
https://example.com/robots.txtapplies only tohttps://example.com.https://www.example.com/robots.txtis a separate file for thewwwhost.https://shop.example.com/robots.txtis separate again.
A robots.txt file in a subfolder, such as /blog/robots.txt, is ignored. The file should be plain UTF-8 text, served with a 200 status code.
How crawlers treat other status codes is worth knowing:
| Response for /robots.txt | Typical crawler behaviour |
|---|---|
| 200 OK | Follow the rules in the file |
| 404 or other 4xx | Treat the site as having no restrictions |
| 5xx server error or timeout | Treat the whole site as blocked, at least temporarily |
That last row surprises people: if your server returns errors for robots.txt, well-behaved crawlers stop crawling the site until it recovers. Crawlers also cache the file, typically for up to a day, so a fix may not take effect immediately.
The structure of a rule group
A robots.txt file is made of groups. Each group starts with one or more User-agent lines naming the crawlers it applies to, followed by Allow and Disallow rules:
User-agent: *
Disallow: /admin/
Disallow: /cart/
Allow: /admin/help/
Sitemap: https://example.com/sitemap.xml
User-agent: *means "any crawler without a more specific group".Disallow: /admin/blocks every URL whose path starts with/admin/.Allow: /admin/help/makes an exception inside the blocked folder.Sitemap:gives the full URL of your XML sitemap. It is not part of any group and can appear anywhere in the file. You can list several.
Lines starting with # are comments. Field names such as User-agent and Disallow are not case-sensitive, but paths are: Disallow: /Private/ does not block /private/.
How a crawler picks its group
A crawler looks for the group whose User-agent value best matches its own name. If it finds one, it follows only that group and ignores the * group completely. Rules are not merged or inherited.
User-agent: *
Disallow: /drafts/
Disallow: /internal/
User-agent: Googlebot
Disallow: /internal/
Here Googlebot follows only its own group, so it may crawl /drafts/. If you want a specific crawler to obey the general rules plus extra ones, you must repeat the general rules inside its group.
How Allow and Disallow are matched
Within a group, the crawler compares the URL path (including any query string) with every rule. The rule with the longest matching path wins, regardless of the order in the file. If an Allow and a Disallow rule match with the same length, Allow wins.
Worked example, using this group:
User-agent: *
Disallow: /shop/
Allow: /shop/sale/
| URL path | Matching rules | Result |
|---|---|---|
/shop/shoes |
Disallow: /shop/ (6 characters) |
Blocked |
/shop/sale/shoes |
Disallow: /shop/ (6) and Allow: /shop/sale/ (11) |
Allowed: longer rule wins |
/shopping-guide |
none (/shop/ needs the slash) |
Allowed |
/about |
none | Allowed |
Notice the role of the trailing slash. Disallow: /shop without a slash would also block /shopping-guide, /shop.html and /shopfront, because rules match the start of the path.
Wildcards: * and $
Two special characters are widely supported:
*matches any sequence of characters, including none.$at the end of a rule means "the URL must end here".
| Rule | Blocks | Does not block |
|---|---|---|
Disallow: /*.pdf$ |
/files/report.pdf |
/files/report.pdf?v=2 |
Disallow: /*?sort= |
/shoes?sort=price |
/shoes?colour=red&sort=price |
Disallow: /*sort= |
both of the above | /shoes |
Disallow: /search |
/search, /search?q=hat, /searching |
/blog/search-tips |
The second row is a common surprise: /*?sort= requires sort= to come straight after the ?. If a parameter can appear anywhere in the query string, /*sort= is the safer pattern, but check it does not match paths you want crawled.
Ready-to-adapt examples
Allow everything (the same as having no file, but explicit and with a sitemap):
User-agent: *
Disallow:
Sitemap: https://example.com/sitemap.xml
An empty Disallow: means "nothing is disallowed".
A typical small business or blog site:
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /search
Disallow: /*?replytocom=
Sitemap: https://example.com/sitemap.xml
This keeps crawlers out of the admin area and internal search result pages, which can generate endless near-duplicate URLs, while still allowing a file that front-end pages depend on.
An online shop with filter parameters:
User-agent: *
Disallow: /cart
Disallow: /checkout
Disallow: /account/
Disallow: /*?*sort=
Disallow: /*?*filter=
Sitemap: https://example.com/sitemap.xml
Filtered and sorted category pages can multiply into thousands of URLs showing the same products. Blocking them helps crawlers spend their time on product and category pages instead.
Opting out of a specific crawler. Some companies publish separate user-agent tokens for crawlers that collect data for AI training, for example GPTBot or Google-Extended. You can block a named crawler without affecting others:
User-agent: GPTBot
Disallow: /
Tokens change and new ones appear, so check each operator's documentation for the current name. And remember this only works for crawlers that choose to honour robots.txt.
What robots.txt cannot do
This is the part most often misunderstood.
It does not keep a page out of search results. Robots.txt controls crawling, not indexing. If other sites link to a blocked URL, a search engine can still list it, usually with no description, because it was never allowed to read the page. To keep a page out of results, let it be crawled and add a noindex directive, either as <meta name="robots" content="noindex"> in the page's HTML or as an X-Robots-Tag: noindex HTTP header. A crawler that is blocked by robots.txt never sees that noindex, so combining the two is counterproductive.
It is not security. The file is public; anyone can read it, and listing /secret-admin-panel/ in it tells curious visitors exactly where to look. Malicious bots ignore it entirely. Anything private needs a password, authentication or removal from the server.
It does not remove pages that are already indexed. Blocking a URL stops future crawling, but the old entry may stay for a long time. Use noindex (with crawling allowed) or the search engine's removal tools instead.
It does not support noindex or nofollow rules inside the file. Some old guides suggest Noindex: lines in robots.txt. Major search engines do not support them.
Crawl-delay is not standard. Some crawlers honour a Crawl-delay value; others, including Google, ignore it. If a crawler is overloading your server, rate limiting on the server is more reliable.
Testing before you publish
- Write or paste your rules into the robots.txt generator and tester.
- Test the paths that matter most: your home page, a product or article page, a category page, and each area you intend to block. Test a few URLs with query strings too.
- Test as the user agents you care about, not just
*, especially if you have crawler-specific groups. - Check that every page in your sitemap is allowed. Listing blocked URLs in a sitemap sends mixed signals. The sitemap generator builds a sitemap from a URL list you can compare against.
- After publishing, open
https://yourdomain/robots.txtin a browser to confirm it returns the new file, not a cached or error page.
To check that an individual page also has the right noindex or canonical tags, open its HTML in the SEO checker, which reports the page's robots directives.
Common mistakes
Disallow: /left over from a staging site. This blocks the entire site. It is the most damaging robots.txt mistake and surprisingly common after a site launch. Protect staging sites with a password instead.- Blocking CSS and JavaScript. Search engines render pages much like a browser. If they cannot load your stylesheets and scripts, they may misjudge the page's layout or content.
- Missing trailing slash (
Disallow: /blogalso blocks/blog-archiveand/blogroll). - Wrong case in paths.
- Expecting groups to combine. A crawler with its own group ignores the
*group. - Relative sitemap URLs.
Sitemap: /sitemap.xmlshould be the full URL, includinghttps://. - One file for several hosts. Each subdomain and each protocol needs its own file.
FAQ
Do I need a robots.txt file at all? No. Without one (a 404 response), crawlers assume everything is allowed. Many small sites only need a short file pointing to their sitemap.
How big can the file be? The standard requires crawlers to read at least the first 500 KiB, and some stop there. A file anywhere near that size usually means rules should be simplified with wildcards.
Why is a blocked page still showing in search results?
Because blocking crawling does not block indexing. Allow crawling, add noindex, and wait for the page to be recrawled.
Does the order of rules matter? Not for well-behaved modern crawlers: the longest matching rule wins. Keeping related rules together still makes the file easier to read and maintain.