Crawling & indexing controls in Google Search Console
Crawling and indexing controls do different jobs. A robots.txt rule controls fetching; a noindex directive asks a search engine to keep a page out of results. Blocking a page from crawling can prevent Google from reading the page's noindex directive, so check both controls together.
Where to start
- Check the rule that applies to the specific URL and crawler, rather than only the site's homepage.
- Inspect meta robots and HTTP X-Robots-Tag directives alongside robots.txt.
- Keep private account data protected by authentication; robots directives are not access controls.
Choose the status or metric you’re seeing
Blocked by Page Removal Tool: What It Means + How to Fix It
“Blocked by page removal tool” in Google Search Console means someone requested a temporary removal of this URL from search results. Here's how it works, how to find who requested it, and how to undo it.
Blocked by robots.txt: What It Means + How to Fix It
“Blocked by robots.txt” in Google Search Console means your robots.txt is telling Google not to crawl this page. Here's when it's intentional, when it's a problem, and exactly how to fix it.
Blocked Due to Access Forbidden (403): What It Means + How to Fix It
“Blocked due to access forbidden (403)” in Google Search Console means your server refused Googlebot with a 403. Here's why it happens — and exactly how to fix it.
Blocked Due to Unauthorized Request (401): What It Means + How to Fix It
“Blocked due to unauthorized request (401)” in Google Search Console means your server asked Googlebot to authenticate before it would serve the page. Here's why — and exactly how to fix it.
Excluded by 'noindex' Tag: What It Means + How to Fix It
“Excluded by 'noindex' tag” in Google Search Console means a noindex directive is keeping your page out of the index. Here's when that's correct, when it's a mistake, and exactly how to fix it.
Indexed Though Blocked by robots.txt: What It Means + How to Fix It
“Indexed, though blocked by robots.txt” in Google Search Console means Google indexed your page without being able to crawl it. Here's why it happens — and how to fix it.
noindex vs. robots.txt: Which One Actually Stops Indexing?
noindex and robots.txt do different jobs — one stops indexing, the other stops crawling. Using the wrong one (or both together) is a common way pages get stuck in Google's index by accident.