SEO

ClaudeBot

Also called Anthropic crawler

Anthropic's web crawler, which requests public pages under its own user-agent token and follows robots.txt rules written for it.

Quick facts: ClaudeBot

Category
SEO
Also called
Anthropic crawler
Level
Intermediate
Affects
How assistants describe your business, crawl load on hosting, AI visibility
Where to see it
robots.txt, server access logs, Anthropic's crawler documentation
In this article4
  1. How ClaudeBot works
  2. Why ClaudeBot matters
  3. Common mistakes with ClaudeBot
  4. What to do about it

How ClaudeBot works

ClaudeBot is Anthropic’s web crawler. It requests publicly reachable pages, identifies itself with its own token in the user-agent string, and honours robots.txt directives written for that token. Anthropic runs more than one agent — a scheduled crawl is a different thing from a fetch triggered when someone asks Claude to open a particular link — and it publishes the current names and address ranges. Check that documentation rather than a blog post, because the list has changed over time.

Nothing about it is exotic. It is an HTTP request like any other. It can be allowed, disallowed or rate-limited, and it leaves an ordinary line in your access log.

Why ClaudeBot matters

It matters for two practical reasons. It is one of the crawlers that determines whether an assistant can describe your business accurately when a customer asks, and it is one of the crawlers that shows up on your hosting bill — a bot that visits often is consuming capacity you pay for, which is a real consideration on a small shared plan.

It also matters because of who has been making the decision. Most site owners have never chosen anything about AI crawlers. The choice arrived through a plugin default, a hosting setting or an agency’s boilerplate robots file, and it is usually a blanket block that nobody has revisited since.

Common mistakes with ClaudeBot

The one that catches people is how robots.txt matching actually works. A crawler picks the single most specific user-agent group that names it and ignores every other group, including the general wildcard one. Add a ClaudeBot group containing nothing but a crawl-delay line and your carefully written Disallow rules no longer apply to it at all. Anything you want that crawler to obey has to be repeated inside its own group.

Two smaller traps. Tokens are matched loosely and are not case-sensitive, so a typo produces a group that matches nothing and reports no error anywhere. And the user-agent string is self-reported, so a scraper can happily call itself ClaudeBot — confirm against the published address ranges before you conclude anything about who visited.

What to do about it

Make one deliberate decision and write it down. If you want assistants to describe your business correctly, allow the crawler and keep the pages that carry your facts reachable. If your objection is to your writing being used for training, express that in rules naming the specific agents you mean, rather than blocking anything with the word bot in it.

Then check the result instead of trusting it. Fetch your robots.txt from the live domain, confirm the group structure is what you intended, and use log file analysis a week later to see whether the token still appears. If crawl volume rather than principle is the problem, a rate limit at the server is a gentler answer than a block: it protects the hosting while keeping you in the answers. The wider trade-off in blocking crawlers is worth thinking through first.

Do and do not

Do

  • Name the exact token in its own robots.txt group
  • Repeat every rule you want inside that group
  • Verify visits against published address ranges

Do not

  • Block everything with the word bot in it
  • Expect a block to remove existing training data
  • Rely on a plugin default you never reviewed

Questions people ask about this

Should a small business block ClaudeBot?

Usually not. Blocking it removes nothing that has already been collected, and it does reduce the chance that an assistant can describe your services correctly when a customer asks. Block it if you publish licensed or paid material, or if crawl volume is genuinely straining your hosting. Otherwise allow it and spend the effort on pages that answer clearly.

How do I write a robots.txt rule for ClaudeBot?

Create a user-agent group naming the token exactly, then put every rule you want it to follow inside that group. Crawlers obey the most specific group that matches them and ignore the rest, including your general rules, so anything left out of the group simply does not apply. Load the file from your live domain afterwards to check it.

Will blocking it remove my content from Claude?

No. A robots.txt rule governs future requests only. It cannot reach into a model that has already been trained and it deletes nothing. What it changes is whether the crawler collects your pages from now on and, depending on which agents you name, whether an assistant can look your site up when someone asks about you.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.