SEO

Disallow

Also called robots.txt disallow, crawl block

The robots.txt rule telling a named crawler not to fetch any address beginning with a given path.

Quick facts: Disallow

Category
SEO
Also called
robots.txt disallow, crawl block
Level
Intermediate
Affects
What gets crawled, crawl budget, rendering, accidental site-wide blocks
Where to see it
Your robots.txt file, Google Search Console, a crawler such as Screaming Frog
In this article4
  1. How Disallow works
  2. Why Disallow matters
  3. Common mistakes with Disallow
  4. How to act on it

How Disallow works

Disallow is a line in robots.txt, the plain text file that sits at the root of a site. Each rule belongs to a group headed by a user-agent line naming the crawler it applies to, and the rule itself gives a path. Anything whose address begins with that path is off limits to that crawler.

Matching is by prefix, not by exact address, which is the source of most accidents. A rule blocking a folder blocks everything beneath it. A rule blocking a short path blocks every longer address that happens to start the same way. Google also understands a wildcard for any run of characters and an end-of-address marker, so patterns can be tighter than a plain folder — but they are still matched against the whole address, and paths are case sensitive.

Two limits define the rule. It is an instruction, not a barrier: well-behaved crawlers honour it and scrapers ignore it entirely, so nothing confidential is protected by it. And it controls fetching, not indexing.

Why Disallow matters

Large sites generate enormous numbers of near-worthless addresses — filtered product listings, internal search results, print versions, session parameters. Left open, a crawler spends its time on them instead of the pages you want read. Disallow is how you point it away from that noise and protect your crawl budget.

It is also, in the wrong hands, one of the fastest ways to take a site out of search. A single line blocking everything, left behind when a staging site is launched, hides the entire site from crawlers while the pages themselves look perfectly normal to anyone browsing. I check that one line before anything else on a site that has suddenly lost its visibility.

Common mistakes with Disallow

The biggest is expecting it to remove a page from Google. Blocking a page stops the crawler fetching it; it does not delete what is already indexed, and if other pages link to the address, Google may keep listing it with no description at all. The result is a bare URL sitting in the results, which is usually worse than the page it replaced.

Worse still is combining it with a noindex tag. The tag lives inside the page, and a blocked crawler never fetches the page, so it never sees the instruction. The block defeats the removal it was meant to support.

The other frequent errors are blocking stylesheets or scripts, which leaves the crawler rendering a broken version of your pages; using it as a privacy measure for documents that should be behind a login; and blocking a page that other pages link to heavily, which strands whatever value those links carried.

How to act on it

Be clear about which outcome you want first. To keep a page out of search results, leave it crawlable and use a noindex instruction. To stop a crawler wasting effort on addresses nobody should read, disallow them. To keep something private, put it behind authentication and stop treating robots.txt as security.

Write rules narrowly, test them in Search Console before publishing, and note that the file applies only to its own protocol, host and port — a subdomain needs its own. Then check it again on launch day, because the rule that blocked a staging site is invisible until traffic disappears.

Do and do not

Do

  • Block low-value addresses like internal search and filters
  • Write narrow paths and test them before publishing
  • Recheck the file on launch day

Do not

  • Use it to remove a page from search results
  • Combine it with a noindex tag on the same page
  • Block stylesheets, scripts or anything genuinely private

Questions people ask about this

Will Disallow remove my page from Google?

No. It stops the crawler fetching the address, which is a different thing from removing it. A page already indexed can stay listed, and one linked from elsewhere can be listed without ever being crawled, showing as a bare URL with no description. To keep a page out of results, allow crawling and use a noindex instruction instead.

Can I use robots.txt to hide private files?

You should not. The file is public, readable by anyone, and it works only for crawlers that choose to obey it. Listing a sensitive path in it advertises exactly where that path is. Anything genuinely private needs a login, a server-level restriction or removal from the public site altogether.

Does a disallowed page still pass link value?

Not usefully. A crawler that cannot fetch the page cannot see the links on it, so anything you meant to flow through that page stops there. If a blocked address collects links from around your site or from other sites, it is doing nothing with them, which is a reason to reconsider whether the block belongs there at all.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.