How Screaming Frog works
You give it a starting URL and it follows links the way a search engine would, requesting each page, reading the HTML and adding any new URLs it finds to the queue. What comes back is a spreadsheet of your site: every URL with its status code, title, meta description, headings, canonical, indexability, word count, response time and the pages that link to it.
It runs on your own computer rather than in the cloud, and that is the source of both its strengths and its limits. It can crawl a staging server or a password-protected site no hosted tool can reach. It can render JavaScript, so a site that builds its content in the browser can be crawled closer to the way Google sees it. And it can pull Search Console, analytics and PageSpeed data in alongside the crawl, so traffic and technical faults sit in the same row.
Why a site crawl matters
Most technical problems are invisible one page at a time and obvious in a list. A whole section sharing a single title, a category reachable only through a broken menu, a redirect chain left behind by an old migration — none of these announce themselves while you are browsing the site normally.
A crawl is also what turns a migration from a hope into a process. Crawl the old site, crawl the new one, compare the two lists, and every URL that lost its title, changed its canonical or now returns a not-found response shows up before a customer meets it.
Common mistakes with Screaming Frog
The most damaging is confusing a crawl with the index. Screaming Frog reports what your site does, not what Google decided about it. A clean crawl and an empty search presence coexist quite happily. Read a crawl next to the page indexing report, never instead of it.
The second is memory. By default the crawl is held in your computer’s RAM, which is fine for a brochure site and hopeless for a large catalogue; switching to database storage before you start saves losing a long crawl halfway through. The third is exporting everything. The tool will happily produce dozens of tabs, and a useful audit picks only the ones that answer the question you actually have.
How to act on it
Start with the response code and indexability columns, because those are absolute: server errors, not-found pages, redirect chains, and anything marked non-indexable that you expected to rank. Then move to duplication — repeated titles, repeated meta descriptions, repeated H1s — which usually points at a template rather than at individual pages.
After that, use it for the questions no other tool answers well: which pages have almost no internal links pointing at them, where depth from the homepage becomes unreasonable, and which images are missing alternative text. Fix by pattern rather than by row. When a crawl produces a long list of structural faults, that is a technical SEO project, and re-crawling after each batch of fixes is how you prove the work actually landed.