Glossary
What is robots.txt?
robots.txt is a text file at the root of a site that tells crawlers which paths they may or may not fetch. It controls crawling, not indexing.
Updated
robots.txt is a plain-text file at the root of your site (/robots.txt) with instructions for crawlers: which paths they may fetch and which to stay out of.
The basics
User-agent: *
Disallow: /admin/
Allow: /
Sitemap: https://example.com/sitemap.xml- User-agent names the crawler the rules apply to;
*means all. - Disallow and Allow list paths.
- Sitemap points to your sitemap.
An important limit
robots.txt stops crawlers *fetching* a page. It does not stop the page being *indexed* if it's linked from elsewhere; to keep a page out of results, use a noindex tag, and let the page be crawled so the tag can be seen.
Mistakes to avoid
- Blocking CSS or JavaScript that pages need to render.
- Leaving a staging
Disallow: /in place after launch. - Treating it as security: it's public, and anyone can read it.
Build and test one with the free robots.txt generator and tester.
Related
- robots.txt generator and testerBuild a robots.txt file, then test whether a URL is allowed or blocked for a given crawler.
- SitemapA sitemap is a file that lists the pages of a site you want search engines to find, often with the date each last changed.
- IndexingIndexing is when a search engine stores a page in its database so it can appear in results. A page that isn't indexed can't rank.