How Disallow works
Disallow is a line in robots.txt, the plain text file that sits at the root of a site. Each rule belongs to a group headed by a user-agent line naming the crawler it applies to, and the rule itself gives a path. Anything whose address begins with that path is off limits to that crawler.
Matching is by prefix, not by exact address, which is the source of most accidents. A rule blocking a folder blocks everything beneath it. A rule blocking a short path blocks every longer address that happens to start the same way. Google also understands a wildcard for any run of characters and an end-of-address marker, so patterns can be tighter than a plain folder — but they are still matched against the whole address, and paths are case sensitive.
Two limits define the rule. It is an instruction, not a barrier: well-behaved crawlers honour it and scrapers ignore it entirely, so nothing confidential is protected by it. And it controls fetching, not indexing.
Why Disallow matters
Large sites generate enormous numbers of near-worthless addresses — filtered product listings, internal search results, print versions, session parameters. Left open, a crawler spends its time on them instead of the pages you want read. Disallow is how you point it away from that noise and protect your crawl budget.
It is also, in the wrong hands, one of the fastest ways to take a site out of search. A single line blocking everything, left behind when a staging site is launched, hides the entire site from crawlers while the pages themselves look perfectly normal to anyone browsing. I check that one line before anything else on a site that has suddenly lost its visibility.
Common mistakes with Disallow
The biggest is expecting it to remove a page from Google. Blocking a page stops the crawler fetching it; it does not delete what is already indexed, and if other pages link to the address, Google may keep listing it with no description at all. The result is a bare URL sitting in the results, which is usually worse than the page it replaced.
Worse still is combining it with a noindex tag. The tag lives inside the page, and a blocked crawler never fetches the page, so it never sees the instruction. The block defeats the removal it was meant to support.
The other frequent errors are blocking stylesheets or scripts, which leaves the crawler rendering a broken version of your pages; using it as a privacy measure for documents that should be behind a login; and blocking a page that other pages link to heavily, which strands whatever value those links carried.
How to act on it
Be clear about which outcome you want first. To keep a page out of search results, leave it crawlable and use a noindex instruction. To stop a crawler wasting effort on addresses nobody should read, disallow them. To keep something private, put it behind authentication and stop treating robots.txt as security.
Write rules narrowly, test them in Search Console before publishing, and note that the file applies only to its own protocol, host and port — a subdomain needs its own. Then check it again on launch day, because the rule that blocked a staging site is invisible until traffic disappears.