What Is a Robots.txt File? A Beginner’s Guide With Examples

A robots.txt file is a small text file that tells search engine crawlers which parts of your website they may visit. It is one of the first things Googlebot checks when it arrives at your site. Used well, it keeps crawlers focused on your important pages. Used badly, it can make an entire website disappear from Google.

Where does robots.txt live?

It must sit in the root of your domain, for example https://example.com/robots.txt. Each subdomain needs its own file. You can view any site’s robots.txt by adding /robots.txt to its address. Try it with a few big websites to see real examples.

Basic syntax

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/sitemap.xml
  • User-agent: which crawler the rules apply to. * means all crawlers.
  • Disallow: a path the crawler should not visit.
  • Allow: an exception inside a disallowed path.
  • Sitemap: the full URL of your XML sitemap.

Paths are case-sensitive, * matches any characters and $ marks the end of a URL. When rules conflict, the most specific (longest) one wins.

Ready-to-use examples

Allow everything

User-agent: *
Disallow:

Sitemap: https://example.com/sitemap.xml

WordPress site

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /?s=
Disallow: /search/

Sitemap: https://example.com/sitemap_index.xml

Online shop

User-agent: *
Disallow: /cart/
Disallow: /checkout/
Disallow: /my-account/
Disallow: /*?orderby=
Disallow: /*?filter_

Sitemap: https://example.com/sitemap.xml

Block AI training crawlers

User-agent: GPTBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

Blocking Google-Extended stops your content from being used to train Google’s AI models, but does not affect your normal Google Search rankings. You can create any of these files in seconds with our robots.txt generator.

robots.txt does not remove pages from Google

This is the most misunderstood point. Robots.txt controls crawling, not indexing. If other sites link to a blocked page, Google can still list its URL in search results, usually without a description. To keep a page out of search results:

  • Add a noindex robots meta tag to the page, and make sure the page is not blocked in robots.txt, so Google can see the tag.
  • Or protect the page with a password.

You can create a noindex tag with the meta tag generator.

Common robots.txt mistakes

  1. Disallow: / left on after launch. Development sites often block everything, and the rule gets copied to the live site. Your whole site drops out of Google.
  2. Blocking CSS and JavaScript. Google needs them to render your pages and judge mobile-friendliness.
  3. Using robots.txt to hide private pages. The file is public, so it actually advertises those paths. Use passwords instead.
  4. Wrong location or file name. It must be /robots.txt, lowercase, at the root.
  5. Forgetting the sitemap line. It is an easy way to help every search engine find your pages.

How to test your robots.txt

  1. Visit yourdomain.com/robots.txt to confirm it loads.
  2. In Google Search Console, go to Settings → robots.txt to see what Google fetched and whether there were errors.
  3. Use URL Inspection on important pages to confirm crawling is allowed.

Next step: your sitemap

Robots.txt tells crawlers where not to go; a sitemap tells them where to go. If your platform does not create one automatically, build one with the XML sitemap generator and add it to your robots.txt and Google Search Console.