How crawler blocking works
Every AI company that reads the web sends out named crawlers, and the established ones publish those names so site owners can control them. You allow or deny each one by user agent in your robots.txt file. Deny a name and a crawler that respects the file stops fetching your pages; allow it and your content stays readable to whatever that bot feeds.
The detail that decides everything is that one company usually runs several bots with different jobs. A crawler that gathers text to train a model is not the same crawler that fetches a page so it can be quoted in a live answer, and Google keeps the token governing its AI products separate from the Googlebot that builds the search index. Block the wrong name and you either lose visibility you meant to keep or keep exposure you meant to remove.
robots.txt is also a request, not a lock. A well-behaved crawler honours it; a scraper with no interest in your rules ignores it and keeps fetching. Only a block at the server, host or CDN actually refuses the connection.
Why the crawler blocking trade-off matters
Blocking looks like the safe default when the worry is that a model will absorb your writing and answer questions with it while nobody visits you. That worry is real. What it misses is that the same access is what makes you quotable. If an assistant cannot fetch your page, it answers from whoever did let it in — a competitor, a directory, or an older article repeating your facts less accurately.
You also cannot correct what you are absent from. A business blocked at every door still gets described by these systems, using whatever other people have written about it, with no page of yours in the mix to set the record straight.
Common mistakes with crawler blocking
The most common is treating “AI crawlers” as a single switch. Blocking the token Google uses for its AI products does not remove you from search results, and blocking a training crawler does not stop the separate bot that fetches pages for live answers — Google-Extended is the clearest example of one company splitting those roles. A blocklist copied from a forum post usually gets at least one of them wrong.
The second mistake is expecting robots.txt to solve an abuse problem. If a bot is hammering the server or ignoring the file, that is a hosting and firewall matter, not an SEO one. The third is blocking once and never revisiting: bot names change, new ones appear, and a file written last year no longer says what its author meant.
How to decide
Start from what the content is worth to you. Service explanations, reference pages and anything that helps a buyer choose you are usually better read than hidden, because being quoted is close to being recommended. Original research, paid material, member-only content and anything you licence are the honest candidates for a block.
Then read your server logs to see which bots actually arrive rather than guessing, and write the file crawler by crawler, with a comment beside each line saying why. Review it whenever your strategy changes. If your aim is the opposite — to be read and cited more often — that is a question of structure and clarity rather than access control, and it belongs with your wider AI search optimisation work.