The five checks
- Status code — only 200 pages can be indexed.
- Meta robots / googlebot — a noindex value removes the page from Google.
- X-Robots-Tag header — the same directive sent as an HTTP header (common for PDFs).
- robots.txt — checked for Googlebot; blocked pages cannot be crawled.
- Canonical — a canonical pointing to another URL asks Google to index that URL instead.
Common reasons pages disappear from Google
- A staging "noindex" left on after launch.
- A robots.txt rule that blocks a whole folder by mistake.
- Canonical tags that all point to the home page.
Note: robots.txt blocks crawling, not indexing. A blocked page can still appear without a description if other sites link to it — use noindex (and allow crawling) to remove a page reliably.
Frequently asked questions
How do I remove a page from Google?
Add a noindex meta tag, make sure robots.txt does not block the page so Google can see the tag, and optionally use the removal tool in Search Console.
Does robots.txt Disallow stop indexing?
Not reliably. It stops crawling, but a URL can still be indexed from links. Use noindex instead.
Why does a page say "canonical points elsewhere"?
Its canonical tag names a different URL, which tells Google to index that URL instead of this one.

