robots.txt: what to block from crawlers and how not to lose traffic
How to write robots.txt: what to block (admin, cart, search), what never to block, how to declare the sitemap and why the file does not hide pages.
Checked against sources: 11 October 2026

- 500 KiBof robots.txt Google reads
- 1 filein the site root: /robots.txt
- noindexnot robots.txt removes a page from the index
What the file does
robots.txt sits in the site root and asks crawlers to skip listed addresses. It is a request, not protection: well-behaved Google and Bing crawlers obey it, malicious ones ignore it.
What to block and what not
- Block: admin, cart, checkout, the account area, internal search, sorting parameter pages.
- Do not block: product, article and category pages, nor CSS and JavaScript: without them Google cannot understand the layout.
- Never use Disallow: / on a live site. It blocks everything and drops the site from search.
User-agent: * Disallow: /admin/ Disallow: /cart/ Disallow: /search/ Sitemap: https://example.com/sitemap.xml
What the file cannot do
A page blocked in robots.txt can stay in results without a description if links point to it. To remove a page from the index you need a noindex tag, and the page must not be blocked in robots.txt so the crawler can see the tag. The robots.txt generator builds the file.
Frequently asked
How large can robots.txt be?
Google reads the first 500 KiB and ignores the rest. A normal site needs a few lines.
How do I block AI crawlers?
With separate groups for GPTBot, ClaudeBot and others. There is a dedicated generator: first decide whether you want visibility in AI answers.
How do I check the file?
Open your-site/robots.txt and make sure it returns text, not a site page. Awe Check tests the file, the sitemap and whether the home page is open to crawlers.
