SEO

Duplicate Content

Also called duplicate pages

The same or near-identical text reachable at more than one URL, which splits signals and forces search engines to pick a version.

Quick facts: Duplicate Content

Category
SEO
Also called
duplicate pages
Level
Intermediate
Affects
Which URL ranks, crawl efficiency, link value
Where to see it
Screaming Frog, Search Console page indexing report, Semrush site audit
In this article4
  1. How duplicate content happens
  2. Why duplicate content matters
  3. Where duplicate content goes wrong
  4. How to act on it

How duplicate content happens

It comes in two forms. Internally, one page turns out to be reachable at several addresses: with and without www, secure and insecure, upper and lower case, trailing slash or none, a tracking parameter on the end, a filter combination, a print version, a product filed under two categories, a blog post reproduced in full on tag and archive pages. Almost nobody creates these deliberately. A content management system produces them by default and nobody looks.

Externally it is the same text living on more than one website: a manufacturer’s description used by every reseller, a press release carried by several outlets, an article syndicated with permission, or a page copied without it. Search engines group pages they judge near-identical, choose one to show, and set the rest aside.

Why duplicate content matters

There is no penalty in the ordinary case, and Google has said so plainly. The cost is quieter than that. Links and signals earned by the content are split between copies instead of pooling on one address. Crawlers spend requests fetching the same thing repeatedly. And the version chosen for the results page may not be the one you would have chosen — a parameterised address, an old subdomain, or in the external case somebody else’s site entirely.

For an online shop the effect is sharper. Variant pages differing by a single colour word, plus category pages generated by filters, can easily outnumber the pages you actually want people to find.

Where duplicate content goes wrong

The commonest false move is assuming a penalty and reacting hard: deleting pages, padding text to make it artificially different, or spinning copy until it reads badly to everyone. Manual action for duplication is reserved for deliberate scraping and mass-generated pages, not for a shop that offers two routes to the same product.

The second is reaching for the wrong instrument. A noindex where a canonical belonged removes a page that should have been consolidated. A canonical where a redirect belonged leaves an address alive that nobody needs. And location pages built by swapping one city name through a single template are the version most people never recognise as duplication, until the pages simply refuse to rank.

How to act on it

Crawl the site and group pages by title and by how similar their text is; near-identical titles are the fastest tell. For each group, decide which address you want to keep. Where the others must stay reachable, point their canonical tag at the one you kept. Where they need not stay reachable, redirect them. Where the pages ought to be genuinely different, make them different or merge them into one stronger page.

For text that appears elsewhere, ask the publisher to canonicalise to your version, or at minimum to link to it. If you sell products described by the manufacturer, write your own descriptions for the lines that actually sell rather than trying to do all of them at once. This grouping and deciding is standard work inside an SEO audit, because the fix is rarely writing and almost always a decision about addresses.

Do and do not

Do

  • Pick one preferred address for each group
  • Canonicalise copies that must stay reachable
  • Redirect copies nobody needs to reach

Do not

  • Rewrite text merely to make it different
  • Build location pages by swapping the city name
  • Assume a penalty and delete pages hastily

Questions people ask about this

Is there a duplicate content penalty?

Not in the normal sense. Google groups near-identical pages and picks one to show, which costs you consolidation rather than triggering a punishment. Penalties in this area are aimed at sites built from scraped or automatically spun text with no value of their own. A shop with two paths to the same product is not in that category.

Does using the manufacturer's product description hurt me?

It rarely earns a penalty, but it makes ranking harder, because dozens of shops are offering identical words and search engines have no reason to prefer yours. Original descriptions, real photographs, delivery detail and answers to the questions buyers actually ask give a page something to be chosen for. Start with the products that make you money.

How similar can two pages be before it is a problem?

There is no fixed threshold, and chasing one is the wrong approach. The practical test is whether the two pages serve different search intents and would give a visitor different answers. If they would not, you have one page split across two addresses, and merging them will serve readers and search engines better than keeping both.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.