Glossary

What is robots.txt?

robots.txt is a text file at the root of a site that tells crawlers which paths they may or may not fetch. It controls crawling, not indexing.

Updated

robots.txt is a plain-text file at the root of your site (/robots.txt) with instructions for crawlers: which paths they may fetch and which to stay out of.

The basics

User-agent: *
Disallow: /admin/
Allow: /

Sitemap: https://example.com/sitemap.xml
  • User-agent names the crawler the rules apply to; * means all.
  • Disallow and Allow list paths.
  • Sitemap points to your sitemap.

An important limit

robots.txt stops crawlers *fetching* a page. It does not stop the page being *indexed* if it's linked from elsewhere; to keep a page out of results, use a noindex tag, and let the page be crawled so the tag can be seen.

Mistakes to avoid

  • Blocking CSS or JavaScript that pages need to render.
  • Leaving a staging Disallow: / in place after launch.
  • Treating it as security: it's public, and anyone can read it.

Build and test one with the free robots.txt generator and tester.

Related