SEO

Discovery

Also called URL discovery

The first step, in which a search engine learns that a URL exists, usually through a link or a sitemap entry.

Quick facts: Discovery

Category
SEO
Also called
URL discovery
Level
Beginner
Affects
Indexing speed, crawl coverage, whether new pages rank at all
Where to see it
Google Search Console (Pages report, URL Inspection), Screaming Frog, server logs
In this article4
  1. How discovery works
  2. Why discovery matters
  3. Where discovery goes wrong
  4. How to act on it

How discovery works

Before a search engine can crawl a page it has to learn that the address exists. That first step is discovery, and it happens in a handful of ways: a link from a page the crawler already knows, an entry in an XML sitemap it has been given, a redirect pointing at the new address, a submission through Search Console or a protocol such as IndexNow, and occasionally a mention picked up somewhere else on the web.

Discovery is not the same as crawling, and neither is the same as indexing. A URL can be known and never fetched, fetched and never stored, or stored and never shown for anything anyone types. Each stage fails for its own reasons, which is why Search Console reports them separately rather than as one status.

The word is also used for two unrelated things — Google Discover, the personalised feed on mobile, and the retired Discovery ad campaign type — so check which one a report means before drawing any conclusion from it.

Why discovery matters

It is the cheapest problem in SEO to fix and the most damaging to ignore. A page nobody has told the search engine about cannot rank at any price, and no amount of content work changes that. New sites, newly launched sections and pages buried behind filters are the usual casualties.

For an ecommerce catalogue it is a structural question rather than a one-off task. Products reachable only through a faceted filter, or only from a paginated list several clicks deep, are found slowly if at all, while the same products linked from a category page are found quickly.

Where discovery goes wrong

The most common cause is a page nothing links to — an orphan page that exists in the content management system and nowhere in the navigation. Publishing it and waiting is not a plan.

Close behind are sitemaps generated once and never updated, sitemaps full of addresses that redirect or return errors, and navigation built only in JavaScript so that no real link ever appears in the HTML. Blocking a section in robots.txt also stops discovery of everything that would have been reached through it, including pages you wanted found.

How to act on it

Link to a new page from somewhere already crawled regularly — a category page, the blog index, a relevant older article — on the day it goes live. Internal links do more for discovery than any submission tool, because they also tell the crawler roughly how important the page is meant to be.

Keep an XML sitemap that is generated automatically, lists only indexable pages that return a success status, and is submitted in Search Console. Then use the URL Inspection tool to request indexing for genuinely important new pages rather than for everything you publish. Where a whole section is being missed, the cause is usually structural, and that is technical SEO work rather than a content fix.

Do and do not

Do

  • Link every new page from an already-crawled page
  • Keep the XML sitemap generated and current
  • Check the Pages report for discovered but unindexed URLs

Do not

  • Rely on a sitemap alone to get pages found
  • Publish pages that nothing on the site links to
  • Block a section in robots.txt without checking inside

Questions people ask about this

How do I get a new page discovered quickly?

Link to it from a page that is already crawled often, such as a category page, the blog index or a closely related article, and make sure it appears in your XML sitemap. Then submit the address in Search Console. The internal link matters most, because it also tells the crawler where the page sits in the site.

Is submitting a sitemap enough on its own?

Usually not. A sitemap tells a search engine that an address exists, but it carries no sense of importance, and pages found only in a sitemap are often crawled slowly and can be left out of the index altogether. Treat it as a safety net behind proper internal linking rather than as the main route into your site.

Why does Search Console say discovered but not indexed?

That status means the address is known but has not been fetched yet, or was fetched and not kept. Common causes are a server slow enough that crawling is being limited, a page that duplicates another on the site, or content thin enough that it was not judged worth storing. Improve the page and its internal links, then request indexing again.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.