Robots.txt Guide: Control Search Engines & AI Crawlers

robots.txt is the first file every crawler reads. One wrong line can hide your whole site from Google; the right lines keep crawlers focused and let you decide which AI bots may train on your content.

Close-up of a futuristic toy robot with blue eyes, showcasing modern technology indoors.
Photo by Pavel Danilyuk on Pexels

robots.txt is a plain-text file at the root of your domain (for example example.com/robots.txt) that tells crawlers which URLs they may and may not request. It controls crawling, not indexing: a blocked URL can still appear in Google if other sites link to it. Use robots.txt to keep crawlers away from low-value areas, point them to your sitemap and decide which AI crawlers may access your content.

Key takeaways: robots.txt blocks crawling, not indexing · never block CSS or JS · use noindex (on a crawlable page) to keep pages out of results · add a Sitemap line · AI crawlers can be blocked without affecting Google Search · test with the Robots.txt Checker.

Syntax in one minute

User-agent: *
Disallow: /admin/
Disallow: /search
Allow: /admin/help

User-agent: GPTBot
Disallow: /

Sitemap: https://example.com/sitemap.xml
  • User-agent starts a group of rules for one crawler (or * for all).
  • Disallow blocks paths that start with the value; Allow makes exceptions.
  • * matches any characters and $ marks the end of the URL — Disallow: /*.pdf$ blocks PDFs.
  • Sitemap can appear anywhere and should be an absolute URL.

How crawlers pick a rule

A crawler uses the most specific group that matches its name and ignores the others — Googlebot follows a User-agent: Googlebot group instead of * if both exist. Within the group, the longest matching rule wins; if an Allow and a Disallow are equally long, Allow wins. No matching rule means allowed. Paths are case-sensitive.

SEO Tutorial - The role of robots.txt files — LinkedIn Learning · Watch on YouTube

What to block

  • Admin, login and account areas.
  • Internal search results (/search, /?s=) — near-infinite thin pages.
  • Cart and checkout steps.
  • Faceted filter combinations that create millions of URLs.
  • Staging or test folders (better: protect them with a password).

What never to block

  • CSS, JavaScript and images needed to render pages — Google renders like a browser.
  • Pages you want indexed.
  • Pages carrying a noindex tag — if crawling is blocked, Google never sees the noindex and may keep the URL in results.

Blocking AI crawlers (and the GEO trade-off)

Several AI companies publish crawler names you can block:

User-agentCompany / purpose
GPTBot, ChatGPT-User, OAI-SearchBotOpenAI — training, browsing, search
Google-ExtendedGemini training (does not affect Google Search)
ClaudeBot, anthropic-aiAnthropic
PerplexityBotPerplexity search
CCBotCommon Crawl dataset
Applebot-Extended, Bytespider, meta-externalagentApple AI, ByteDance, Meta

Blocking training crawlers protects your content from being used to train models. But blocking search-style AI crawlers (OAI-SearchBot, PerplexityBot) can also stop AI assistants from citing and linking to you — the opposite of what generative engine optimisation aims for. Many publishers block training bots and allow search bots. The Robots.txt Generator lets you choose with one checkbox.

Crawl-delay

Bing and Yandex respect Crawl-delay: 10 (seconds between requests); Google ignores it. Only use it if a crawler genuinely overloads your server.

Testing before you publish

  1. Generate or edit the file with the Robots.txt Generator.
  2. Upload it to the root of every host (each subdomain needs its own file).
  3. Test key URLs for Googlebot, Bingbot and GPTBot with the Robots.txt Checker — it shows the exact rule that decided each result.
  4. Confirm the sitemap URL works with the Sitemap Checker.
  5. Re-check after every site launch: a leftover Disallow: / from staging is one of the most common causes of sudden traffic loss.

Frequently asked questions

Does robots.txt remove pages from Google?

No. It stops crawling, not indexing. To remove a page, let it be crawled and add a noindex tag, or delete it and return 404 or 410.

Will blocking GPTBot affect my Google rankings?

No. GPTBot belongs to OpenAI and has no effect on Google Search. It may reduce how often ChatGPT can use your content.

Where must robots.txt be located?

At the root of each host, for example https://www.example.com/robots.txt. Subdomains need their own file.

Is robots.txt case-sensitive?

Paths are case-sensitive, so Disallow: /Admin does not block /admin. User-agent names are matched case-insensitively.

What happens if I have no robots.txt?

Crawlers assume everything is allowed. It is not an error, but adding a file with your Sitemap line is good practice.

Shorten links & create QR codes free

Clean fvj.io links, dynamic QR codes and click analytics — plus 90 free SEO tools.

Keep reading