How ClaudeBot works
ClaudeBot is Anthropic’s web crawler. It requests publicly reachable pages, identifies itself with its own token in the user-agent string, and honours robots.txt directives written for that token. Anthropic runs more than one agent — a scheduled crawl is a different thing from a fetch triggered when someone asks Claude to open a particular link — and it publishes the current names and address ranges. Check that documentation rather than a blog post, because the list has changed over time.
Nothing about it is exotic. It is an HTTP request like any other. It can be allowed, disallowed or rate-limited, and it leaves an ordinary line in your access log.
Why ClaudeBot matters
It matters for two practical reasons. It is one of the crawlers that determines whether an assistant can describe your business accurately when a customer asks, and it is one of the crawlers that shows up on your hosting bill — a bot that visits often is consuming capacity you pay for, which is a real consideration on a small shared plan.
It also matters because of who has been making the decision. Most site owners have never chosen anything about AI crawlers. The choice arrived through a plugin default, a hosting setting or an agency’s boilerplate robots file, and it is usually a blanket block that nobody has revisited since.
Common mistakes with ClaudeBot
The one that catches people is how robots.txt matching actually works. A crawler picks the single most specific user-agent group that names it and ignores every other group, including the general wildcard one. Add a ClaudeBot group containing nothing but a crawl-delay line and your carefully written Disallow rules no longer apply to it at all. Anything you want that crawler to obey has to be repeated inside its own group.
Two smaller traps. Tokens are matched loosely and are not case-sensitive, so a typo produces a group that matches nothing and reports no error anywhere. And the user-agent string is self-reported, so a scraper can happily call itself ClaudeBot — confirm against the published address ranges before you conclude anything about who visited.
What to do about it
Make one deliberate decision and write it down. If you want assistants to describe your business correctly, allow the crawler and keep the pages that carry your facts reachable. If your objection is to your writing being used for training, express that in rules naming the specific agents you mean, rather than blocking anything with the word bot in it.
Then check the result instead of trusting it. Fetch your robots.txt from the live domain, confirm the group structure is what you intended, and use log file analysis a week later to see whether the token still appears. If crawl volume rather than principle is the problem, a rate limit at the server is a gentler answer than a block: it protects the hosting while keeping you in the answers. The wider trade-off in blocking crawlers is worth thinking through first.