Robots.txt Guide: Control Search Engines & AI Crawlers
robots.txt is the first file every crawler reads. One wrong line can hide your whole site from Google; the right lines keep crawlers focused and let you decide which AI bots may train on your content.
robots.txt is a plain-text file at the root of your domain (for example example.com/robots.txt) that tells crawlers which URLs they may and may not request. It controls crawling, not indexing: a blocked URL can still appear in Google if other sites link to it. Use robots.txt to keep crawlers away from low-value areas, point them to your sitemap and decide which AI crawlers may access your content.
Key takeaways: robots.txt blocks crawling, not indexing · never block CSS or JS · use noindex (on a crawlable page) to keep pages out of results · add a Sitemap line · AI crawlers can be blocked without affecting Google Search · test with the Robots.txt Checker.
Syntax in one minute
User-agent: *
Disallow: /admin/
Disallow: /search
Allow: /admin/help
User-agent: GPTBot
Disallow: /
Sitemap: https://example.com/sitemap.xml
User-agentstarts a group of rules for one crawler (or*for all).Disallowblocks paths that start with the value;Allowmakes exceptions.*matches any characters and$marks the end of the URL —Disallow: /*.pdf$blocks PDFs.Sitemapcan appear anywhere and should be an absolute URL.
How crawlers pick a rule
A crawler uses the most specific group that matches its name and ignores the others — Googlebot follows a User-agent: Googlebot group instead of * if both exist. Within the group, the longest matching rule wins; if an Allow and a Disallow are equally long, Allow wins. No matching rule means allowed. Paths are case-sensitive.
What to block
- Admin, login and account areas.
- Internal search results (
/search,/?s=) — near-infinite thin pages. - Cart and checkout steps.
- Faceted filter combinations that create millions of URLs.
- Staging or test folders (better: protect them with a password).
What never to block
- CSS, JavaScript and images needed to render pages — Google renders like a browser.
- Pages you want indexed.
- Pages carrying a noindex tag — if crawling is blocked, Google never sees the noindex and may keep the URL in results.
Blocking AI crawlers (and the GEO trade-off)
Several AI companies publish crawler names you can block:
| User-agent | Company / purpose |
|---|---|
| GPTBot, ChatGPT-User, OAI-SearchBot | OpenAI — training, browsing, search |
| Google-Extended | Gemini training (does not affect Google Search) |
| ClaudeBot, anthropic-ai | Anthropic |
| PerplexityBot | Perplexity search |
| CCBot | Common Crawl dataset |
| Applebot-Extended, Bytespider, meta-externalagent | Apple AI, ByteDance, Meta |
Blocking training crawlers protects your content from being used to train models. But blocking search-style AI crawlers (OAI-SearchBot, PerplexityBot) can also stop AI assistants from citing and linking to you — the opposite of what generative engine optimisation aims for. Many publishers block training bots and allow search bots. The Robots.txt Generator lets you choose with one checkbox.
Crawl-delay
Bing and Yandex respect Crawl-delay: 10 (seconds between requests); Google ignores it. Only use it if a crawler genuinely overloads your server.
Testing before you publish
- Generate or edit the file with the Robots.txt Generator.
- Upload it to the root of every host (each subdomain needs its own file).
- Test key URLs for Googlebot, Bingbot and GPTBot with the Robots.txt Checker — it shows the exact rule that decided each result.
- Confirm the sitemap URL works with the Sitemap Checker.
- Re-check after every site launch: a leftover
Disallow: /from staging is one of the most common causes of sudden traffic loss.
Frequently asked questions
Does robots.txt remove pages from Google?
No. It stops crawling, not indexing. To remove a page, let it be crawled and add a noindex tag, or delete it and return 404 or 410.
Will blocking GPTBot affect my Google rankings?
No. GPTBot belongs to OpenAI and has no effect on Google Search. It may reduce how often ChatGPT can use your content.
Where must robots.txt be located?
At the root of each host, for example https://www.example.com/robots.txt. Subdomains need their own file.
Is robots.txt case-sensitive?
Paths are case-sensitive, so Disallow: /Admin does not block /admin. User-agent names are matched case-insensitively.
What happens if I have no robots.txt?
Crawlers assume everything is allowed. It is not an error, but adding a file with your Sitemap line is good practice.
Shorten links & create QR codes free
Clean fvj.io links, dynamic QR codes and click analytics — plus 90 free SEO tools.

