Analytics and Tracking

Crawl Tool

Also called Site crawler, SEO spider

Software such as Screaming Frog or Sitebulb that follows a site's links and records what each page returns.

Quick facts: Crawl Tool

Category
Analytics and Tracking
Also called
Site crawler, SEO spider
Level
Intermediate
Affects
Technical SEO, site structure, indexing, internal linking
Where to see it
Screaming Frog, Sitebulb, Ahrefs Site Audit, Semrush Site Audit
In this article4
  1. How a crawl tool works
  2. Why a crawl tool matters
  3. Common mistakes with crawl tools
  4. How to act on it

How a crawl tool works

You give it a starting address. It requests that page, reads the HTML, pulls out every link it contains, adds them to a queue and repeats until no new addresses are left. For each URL it records the status code, the title, the meta description, the headings, the canonical tag, robots directives, the word count, the response time and which pages link to it.

Two settings change the results more than any others. The first is whether the crawler renders JavaScript. If your site builds its content in the browser and the crawler only reads the raw HTML, it will report pages as empty when they are not. The second is speed: a crawler asking for many pages a second can overwhelm a small shared host, so the rate has to be turned down until the site copes. Most tools can also read a list of URLs instead of following links, and connect to Search Console or analytics so each row carries clicks and sessions beside its technical detail.

Why a crawl tool matters

It is the only practical way to see a site as a machine sees it. A person browsing finds the pages the navigation offers; a crawler finds the pages nobody remembers publishing, the category that redirects three times before landing, the product template that ships the same title on every variant, and the pages no internal link points at.

It also turns opinion into a list. “The site feels messy” is not something anyone can act on. A table of URLs with duplicated titles, or a set of internal links pointing at addresses that no longer exist, is.

Common mistakes with crawl tools

Crawling a large site with default settings and then losing the run to memory or a timeout. Set the crawl to save to disk, limit the depth, and exclude the parts you do not need — faceted filters and search result pages can generate endless addresses.

Treating every flagged row as a defect is the more expensive mistake. These tools report everything they notice, and much of it is harmless. A missing meta description on a tag archive matters far less than an orphan page holding your best content. Crawling a staging copy and reporting its faults as the live site’s is a third, and it happens more often than anyone admits.

How to act on it

Match the crawl to how the site is built, throttle it so the host survives, and crawl the live site from a clean starting point. Then compare what came back with your sitemap and with Search Console: addresses in the crawl but not the sitemap, and indexed addresses the crawl never reached, are both worth explaining.

Fix by class, not by row. One template usually explains hundreds of flagged URLs, and correcting the template clears them all at once. Group the findings, order them by how much traffic sits behind each group, and hand the list to whoever maintains the site — that grouping is the substance of most technical SEO work.

Do and do not

Do

  • Throttle the crawl so a small host survives it
  • Render JavaScript when the site builds content in the browser
  • Fix by template, not one row at a time

Do not

  • Treat every flagged row as a real problem
  • Crawl a staging copy and report it as live
  • Let facets and search pages generate endless URLs

Questions people ask about this

How often should I crawl my site?

A full crawl before and after any significant change — a redesign, a migration, a new template — plus a routine crawl each quarter for a stable site. Sites that publish constantly or run an online shop with changing stock benefit from something more frequent, because broken links and duplicate titles appear as fast as pages do.

Will crawling slow down or damage my website?

It can, if you let it run at full speed against a small shared host. The crawler is making real requests, and a burst of them competes with real visitors. Lower the crawl rate, limit how many threads run at once, and where possible crawl outside your busiest hours. Done carefully it leaves no trace.

Does a crawl tool show me what Google actually indexed?

No. It shows what a crawler following your links can reach and what those pages contain. Whether Google chose to index a page is a separate question, answered in Search Console. Use the two together: the crawl explains what exists and how it is linked, and Search Console explains what Google did with it.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.