SEO

Crawling

Also called spidering

The process of search engine bots following links and fetching pages, which happens before anything can be indexed or ranked.

Quick facts: Crawling

Category
SEO
Also called
spidering
Level
Beginner
Affects
Whether pages get indexed, how fast updates are noticed
Where to see it
Search Console URL Inspection, crawl stats report, Screaming Frog
In this article4
  1. How crawling works
  2. Why crawling matters
  3. Common mistakes with crawling
  4. How to act on it

How crawling works

A search engine keeps a running list of addresses it intends to visit. A bot — Googlebot, in Google’s case — requests each one much as a browser would, reads the HTML that comes back, notes every link inside it and adds anything unfamiliar to the list. That loop is crawling, and it never stops.

Addresses reach the list in a few ways: links from pages already known, entries in an XML sitemap, redirects from addresses already known, and links from other websites. Before fetching anything from a site, a crawler reads its robots.txt file to see which paths it is allowed to request. Where a page assembles its content with JavaScript, the crawler has to render the page before there is anything to read, which is a second and slower step in the same process.

Why crawling matters

It is the first gate, and it is absolute. A page that is never fetched cannot be indexed, cannot rank and cannot be quoted in an answer box. Every piece of advice about writing better content quietly assumes the page was reachable in the first place.

It is also where a surprising number of invisible problems live. A product page reachable only through an on-site search box. A menu that assembles itself after the page has loaded. A page nobody ever linked to because it was built for one email campaign. None of these looks broken to a visitor handed the address, and all of them can be missing from search entirely.

Common mistakes with crawling

The classic launch error is a robots.txt file copied across from staging that tells every crawler to stay out of the whole site. It is a two-line file and it can hide an entire business. Read it on launch day, and read it again a week later.

Confusing crawled with indexed comes next: being fetched is permission to be considered, not a promise of inclusion. Blocking CSS and JavaScript is another, because the crawler then renders a version of the page your visitors never see and judges that instead. And leaning on a sitemap alone, with no links pointing at a page, may get it discovered but says nothing about whether it matters.

How to act on it

Use URL Inspection in Search Console on a page that matters. It reports whether Google has fetched the address, when it last did, what it saw and whether anything blocked it. Run it on one page from each template — home, service, article, product — rather than trying to check everything.

Then make the site easy to walk. Any page worth ranking should be reachable by following ordinary links from the home page in a small number of clicks. Keep robots.txt short and reviewed. Watch for server errors and slow responses, which quietly reduce how much gets fetched. Publish a sitemap listing the pages you want found, but treat it as a supplement to internal linking rather than a replacement — a page with nothing linking to it is an orphan whatever the sitemap claims. Keeping that plumbing sound is everyday technical SEO.

Do and do not

Do

  • Keep every page a few clicks from home
  • Read robots.txt before and after launch
  • Inspect one live URL per page template

Do not

  • Block CSS or JavaScript from crawlers
  • Assume a crawled page will be indexed
  • Rely on a sitemap instead of links

Questions people ask about this

How often does Google crawl my site?

It varies by page and by site. Addresses that change frequently, attract links and receive visitors tend to be revisited quickly, while a page that has sat untouched for a long time may be checked rarely. The crawl stats report in Search Console shows what is actually happening on your site rather than what should be.

Is crawling the same as indexing?

No. Crawling is fetching the page; indexing is deciding to store and consider it. A page can be crawled and then left out of the index because it duplicates another page, offers little of substance, or carries an instruction not to index it. Both steps have to succeed before a page can appear in results.

Why has Google not crawled my new page?

Usually because nothing points at it. Check that the page is linked from somewhere already crawled, that robots.txt does not block its path, and that it appears in your sitemap. Then inspect the address in Search Console, which will name the reason if the page was found but deliberately skipped.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.