Skip to content
noindex vs robots.txt disallow: which one removes a page from Google? — illustrated banner

noindex vs robots.txt disallow: which one removes a page from Google?

Use noindex. A robots meta tag or X-Robots-Tag header with noindex tells Google to drop the page from its index. A robots.txt Disallow only stops crawling; the URL can still be indexed from links, without a description. Never combine them: if robots.txt blocks the page, Googlebot cannot fetch it and never sees the noindex.

Share article:

Key takeaways

  • robots.txt controls crawling; noindex controls indexing. They are different jobs.
  • A disallowed URL can still appear in Google, with no description, if other pages link to it.
  • noindex only works if the crawler is allowed to fetch the page and read it.
  • For a fast removal, use the Removals tool and then make it permanent with noindex, 404/410 or a password.

“Block it in robots.txt” is the most common answer to “how do I get this page out of Google?”, and it is the wrong one. robots.txt and noindex sound like two strengths of the same control. They are two different controls, and used together they cancel out.

What does robots.txt disallow actually do?

It asks a crawler not to request a URL. Google is explicit about the limit: robots.txt “is used mainly to avoid overloading your site with requests; it is not a mechanism for keeping a web page out of Google.” If a disallowed URL is linked from elsewhere, Google can still index the address: “its URL can still appear in search results, but the search result won't have a description.” That is the familiar result with no snippet — indexed, never read.

What does noindex do?

noindex is an indexing rule, delivered with the page. It can be a meta tag in the HTML or an HTTP response header:

<meta name="robots" content="noindex">

HTTP/1.1 200 OK
X-Robots-Tag: noindex

The header form works for PDFs, images and anything else without a <head>. When Googlebot fetches the page and sees the rule, it drops the page from search results. Google calls noindex, in the HTML or the header, “the most effective way to remove URLs from the index when crawling is allowed.”

Why does combining noindex and disallow fail?

Because the noindex is inside the page, and the disallow stops anyone reading the page. Google’s documentation puts it in one sentence: “For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file”. If it is blocked, “the crawler will never see the noindex rule, and the page can still appear in search results, for example if other pages link to it.”

# robots.txt — this keeps the noindex below from ever being read
User-agent: *
Disallow: /internal-search/

<!-- /internal-search/?q=shoes -->
<meta name="robots" content="noindex">

The order that works: remove the Disallow, leave the noindex in place, and let Google recrawl. Once the URLs have dropped out, you may add the Disallow back to save crawl requests — but understand that from then on any new noindex on those paths will go unread again.

robots.txt decides whether a crawler may read the page. noindex is something the crawler reads. You cannot ask it to read a sign you have locked away.

Adnan Arodiya, Crawlwise

Which should you use, and when?

GoalUseWhy
Keep a page out of Google’s resultsnoindex (meta or header), crawl allowedThe only rule that controls indexing
Stop crawlers wasting requests on infinite URL spacesrobots.txt DisallowControls crawling; pages may still be indexed by URL
Remove a page that no longer exists404 or 410Google treats every 4xx except 429 as “content doesn't exist”
Hide something privatePassword or loginNeither robots.txt nor noindex is access control
Remove something urgentlySearch Console Removals, then one of the aboveA temporary removal lasts about six months

One more trap: a noindex line inside robots.txt. Google retired its handling of that unsupported rule on September 1, 2019, so it does nothing in Google today. For duplicate URLs within one site, Google also advises against noindex as a canonicalization tool “because it will completely block the page from Search” — use a canonical instead, as explained in how Google treats rel=canonical. And if you are deciding what to do about AI crawlers rather than Googlebot, the same crawl-versus-use split applies; see does robots.txt control AI crawlers?

How do you remove a page from Google for good?

  1. Decide whether the page should exist. If not, return 404 or 410 (see HTTP status codes for SEO).
  2. If it should exist but not be found in search, add noindex as a meta tag or header.
  3. Make sure robots.txt does not disallow it, so the noindex can be read.
  4. For speed, submit a temporary removal in Search Console, and request a recrawl with URL Inspection.

How Crawlwise checks this

The free robots.txt tester tells you whether a specific URL is disallowed for Googlebot or any other user agent, and lists the rules in the group that names that agent — the first thing to check when a noindex “isn’t working”. The HTTP status code checker shows the X-Robots-Tag header on the final response and adds a note when it contains noindex, which catches the header a CDN or framework set without anyone noticing. Neither tool can tell you what Google has indexed; Search Console is the only source for that.

What each check reads, what counts as a failure, and what the checks cannot see are written up on the methodology page.

What a plan actually costs

The free tools stay free. You can open them, run them, and leave without an account. A plan is for the moment you want the report saved, a few sites watched, or the same checks from the API. The price on this table is the price at checkout. A yearly plan is ten months of the monthly price, so two months are on us.

PlanMonthlyCreditsA good fit when
Starter$4.9910You look after one site and check it now and then
Pro$14.9940A few sites, plus the API and a handful of watched URLs
Studio$49.99150Client work that would burn through Pro mid-month
Agency$99.99400Many locations, reported under your own name

A single-page audit is about 1.25 credits, and that includes the live probe of answer-engine crawlers. You see the estimate before anything runs. If a hold is not used, it comes back to your balance. Extra credits, when you already subscribe, are $9.99 for 20.

Frequently asked questions

Can I put noindex in robots.txt?

No. Google stopped supporting noindex as a robots.txt rule on September 1, 2019. Use a robots meta tag or an X-Robots-Tag header instead.

How long does noindex take to work?

It works the next time Googlebot crawls the page. Google says that, depending on the page’s importance, a revisit can take months; you can request a recrawl with the URL Inspection tool in Search Console.

Is the Removals tool a substitute for noindex?

No. Google says a temporary removal lasts about six months. To keep a page out after that, it needs noindex, a 404 or 410, or a password.

Sources

Documentation was read on 2026-10-05. Search engines change these pages without notice, so check the original before you act on a detail.

Crawlwise indexing checks vs checking by hand

  • Reads the robots meta tag in the served HTML

    Crawlwise
    Yes, in the page audit
    By hand
    View source
  • Reads X-Robots-Tag on the final response

    Crawlwise
    Yes, flagged when it says noindex
    By hand
    curl -I, if you remember to
  • Tests the URL against robots.txt for a given bot

    Crawlwise
    Yes, with the matching rule
    By hand
    By eye
  • Tells you whether Google has indexed the page

    Crawlwise
    No
    By hand
    Search Console URL Inspection

Meet the author

Photo of Adnan Arodiya

Related articles