SEO blog · Technical SEO

Canonical Tags and Duplicate Content: When and How to Use Them

Key takeaways

  • A canonical tag tells Google which URL is the preferred version when several pages show the same or very similar content.
  • Rel=canonical is a strong signal, not a command: Google may choose a different URL if your signals contradict each other.
  • Duplicate content does not trigger a penalty by itself; the real costs are diluted signals and Google showing the wrong page.
  • Every indexable page should carry an absolute canonical in the head, pointing to itself or to its preferred equivalent.
  • Use URL Inspection to compare the canonical you declared with the one Google selected.

A canonical tag is a signal that tells Google which URL you consider the main version when the same or very similar content appears at several addresses. You declare it with rel="canonical" in the page <head> or in an HTTP header, and redirects and sitemaps send related signals. Google treats the tag as a strong signal rather than a command, and it makes the final decision on which address appears in results.

Duplicates appear almost without you noticing: the same page with and without www, with ?utm_source=, with a different sort order, in a print version or inside several categories. Without guidance, Google picks a version on its own, and its pick may not be the one you want. This guide covers how Google judges duplicates, when to use the canonical tag, what the code looks like, what happens in online stores and how to check the result. It is a central piece of technical SEO.

What is duplicate content, and why is it not a penalty?

Duplicate content is identical or substantially similar text reachable at more than one URL, within one site or across sites. The usual causes are technical, not deliberate:

  • http and https, www and non-www, trailing slash or none;
  • tracking or sorting parameters (?utm_medium=, ?sort=price);
  • the same product page inside several categories;
  • archive, tag or print pages that repeat articles;
  • content republished on another site, as happens with advertorials.

Contrary to a popular belief, Google does not run a “duplicate content penalty” in normal cases; its documentation says some duplicate content on a site is normal and not a spam policy violation. It clusters similar pages, picks a representative one and shows it. Action is reserved for duplication done to deceive, not a page reachable through three addresses.

The real problems are different:

  1. Diluted signals. Links and engagement are split across variants instead of adding up on one.
  2. The wrong choice. Google may show the parameter version or a secondary page.
  3. Wasted crawling. Google spends requests on copies instead of new pages.

Do not confuse a technical duplicate with keyword cannibalization: there the pages are different, but they compete for the same search.

How does rel=canonical work?

Google accepts several ways of saying which address you prefer, with different strength, according to its documentation:

MethodHow to set itStrength
Permanent redirect301 or 308 from the duplicate to the main pageStrong signal
rel="canonical" in <head><link rel="canonical" href="...">Strong signal
HTTP Link headerLink: <url>; rel="canonical"Strong; useful for PDFs and files
Sitemap inclusionOnly the canonical address in the sitemapWeak signal

Notice that a redirect makes the duplicate disappear for users, while the canonical tag keeps both pages reachable. If the variant no longer needs to exist, choose the redirect; types and setup are covered in our guide to 301 and 302 redirects. If both pages must stay reachable (filters, parameters, variants), choose the canonical tag.

Each method can be combined with the others, as long as they all point to the same address. Contradictory signals lower the chance that Google follows your preference.

How do you implement a canonical tag correctly?

The rules from Google’s documentation:

  1. Use an absolute URL, with protocol and full domain, not a relative path.
  2. Put one canonical per page, in the <head>. A tag that ends up in the <body> is not considered valid.
  3. Add the canonical to the preferred page too, pointing to itself (self-referencing).
  4. Point to an accessible page that returns 200, not one that redirects, is blocked or does not exist.
  5. Be consistent. Do not declare different canonicals through HTML, header and sitemap.
  6. Put it in the served HTML, not injected with JavaScript, whenever possible.

An HTML example, on a category page that can also be reached with parameters in the address:

<head>
  <link rel="canonical" href="https://www.example.com/running-shoes/" />
</head>

An example for a PDF, in the HTTP header (Apache):

<Files "guide-copy.pdf">
  Header add Link '<https://www.example.com/guide.pdf>; rel="canonical"'
</Files>

On WordPress, SEO plugins add a self-referencing canonical by default. Check a few page types (posts, categories, paginated pages, archives) to confirm the tags are what you expect. Platform basics are in our WordPress SEO guide.

What should you do in common scenarios?

This is where most time is lost, so here is a table of typical decisions:

SituationWhat to do
http/https, www/non-www301 redirect to the chosen version, plus a self-referencing canonical
UTM and session parametersCanonical to the address without parameters
Sorts and filters with no search valueCanonical to the base category or noindex, not both without a reason
Filter with real demand (category plus brand)Own indexable page with content and its own canonical
Paginated pages (?page=2)Each page has its own canonical; do not point all to page 1
Product in several categoriesOne canonical product URL, used in internal links
Content republished on another siteCanonical to the original source, or noindex at the partner
Language versionsEach version canonical to itself, with hreflang annotations

For stores, remember the pagination rule: pointing page 1 as the canonical for the whole set hides the products on later pages. Category page structure is covered in our guide to category pages in a store.

For republished content, such as an article that also runs on a partner site, ask for a canonical pointing to your source. If you work with external publications, the difference between publishing formats is explained in our press release vs advertorial comparison. For languages, the canonical must stay within the same language; see the hreflang guide and the hreflang generator.

Example: one category, five addresses

Suppose a store has a category called “Running shoes.” The same product list can be opened like this:

  1. https://www.example.com/running-shoes/
  2. https://www.example.com/running-shoes/?sort=price
  3. https://www.example.com/running-shoes/?utm_source=newsletter
  4. https://www.example.com/running-shoes/?page=2
  5. https://www.example.com/running-shoes/nike/

The decision, address by address:

  • Address 1 is the preferred version: it carries a canonical pointing to itself.
  • Addresses 2 and 3 show the same products, just ordered differently or tracked. They get a canonical to address 1, and internal links should never use them.
  • Address 4 is a different page in the series, with different products. It carries a canonical to itself, not to address 1.
  • Address 5 is category plus brand, with real search demand. It deserves its own content (title, intro text, relevant products) and a canonical to itself.

Without these decisions, a crawler can discover hundreds of combinations and treat them as separate pages. With them, Google gets the same story from several sources: HTML, internal links and sitemap all pointing to the same addresses.

Canonical, noindex, redirect or robots.txt: which one?

These tools get mixed up, though they do different things:

ToolEffectUse it when
CanonicalConsolidates signals on one address; both pages stay reachableDuplicates that must exist
301 redirectMoves users and signals; the duplicate disappearsThe old page should no longer be reached
noindexRemoves the page from the index, with no consolidationPages with no search value (thank-you pages, internal search)
robots.txtBlocks crawling, not indexingAreas that should not be crawled; not for consolidation

Google’s documentation explicitly advises against robots.txt for canonicalization: if you block access, Google cannot read the canonical tag. Do not combine noindex with a canonical to another page either, because you send opposite messages. For robots.txt, our generator helps you avoid accidental blocks.

How do you check which canonical Google chose?

Your canonical is a suggestion. Check what Google picked:

  1. In Search Console, open URL Inspection for the page.
  2. Read “User-declared canonical” and “Google-selected canonical.” If they differ, your signals conflict.
  3. Open the Pages report and look for canonical-related exclusion reasons: “Duplicate without user-selected canonical,” “Duplicate, Google chose different canonical than user” and “Alternate page with proper canonical tag.”
  4. The last one is not necessarily a problem: it means the duplicate page was correctly associated with the original.

If Google picks a different address than yours, check internal links (do they point to different variants?), the sitemap, redirects and content differences between the pages. How to read these reports is explained in the Search Console guide, and the causes of exclusions in the guide to pages not indexed.

Common canonical tag mistakes

  • Relative canonicals. href="/running-shoes/" can be misread; use full addresses.
  • A canonical pointing to a redirected or 404 page. Google will likely ignore the signal.
  • Several canonicals on one page. Usually one comes from the theme and another from a plugin.
  • A canonical in the <body>. The tag must be in the <head>.
  • All paginated pages pointing to the first. You lose the products or articles on later pages.
  • A sitewide canonical to the homepage. A faulty template can point every page at the homepage, which pushes them out of the index.
  • A sitemap with non-canonical addresses. It weakens the message; see our XML sitemap guide.
  • Canonicals between pages with different topics. The tag is for equivalent content, not a way to “move” authority.

When a canonical tag will not solve the problem

Be realistic about the limits:

  • It is a hint: Google may ignore it if the pages are too different or if other signals (internal links, redirects) point elsewhere.
  • It does not fix thin or templated content. If different pages say practically the same thing, the answer is to merge them or differentiate the content.
  • It does not replace architecture: if a site generates thousands of variants through filters, you must also limit how those URLs are created.
  • It cannot rescue a page that is blocked or not indexable.

Conclusion and next steps

A simple plan:

  1. Choose a single version of the domain (HTTPS, with or without www) and redirect the rest.
  2. Check that every indexable page has an absolute, self-referencing canonical in the <head>.
  3. Handle parameters and filters with the rule “does it have search demand or not.”
  4. Watch Search Console for differences between the declared canonical and Google’s choice.
  5. Keep internal links and the sitemap on canonical addresses.

If you want someone to run these checks on your site, from canonicals to indexing, see our technical SEO service or reach us through contact.

Sources and further reading

Frequently asked questions

Frequently asked questions

Does duplicate content cause a Google penalty?

Usually not. Google groups similar pages and picks one to show. Penalties apply to duplication done to manipulate results, not to normal technical situations such as parameters or product variants. The real risk is that Google shows a different page than the one you want.

Should the original page also have a canonical tag?

Yes, a canonical pointing to the page itself, called self-referencing, is recommended. It protects against parameters added by campaigns or external links and makes the preferred address clear. Many SEO plugins add it automatically; it is worth checking that they do it correctly.

Can I use canonical tags across two different domains?

Yes, Google accepts cross-domain canonicals, for example when content is republished on another site. It remains a hint, not an obligation. It works better when the pages are nearly identical and the partner actually adds the tag pointing to your original.

How long until Google respects a canonical change?

There is no fixed timeline. Google has to recrawl the pages involved and re-evaluate the duplicate cluster, which can take anywhere from a few days to a few weeks. You can request indexing of the key page in URL Inspection, but that does not guarantee an immediate change.

Do canonical tags help with filtered category pages?

They can, if a filter creates variants nearly identical to the base category. If a filter matches a real search, such as category plus brand, it deserves its own indexable page with distinct content. The decision rests on actual search demand, not on a technical rule.

Related service

Technical SEO & speed

See the service →

Let’s grow your site’s organic traffic

Send us your website address and we’ll reply with a free initial analysis and a concrete SEO strategy — no strings attached.