Canonical Tags and Duplicate Content: When and How to Use Them
Canonical tag explained: how Google handles duplicate content, when to use rel=canonical, code examples, store scenarios and mistakes to avoid.
Key takeaways
Pages not indexed are URLs that Google knows about but has not added to its search index, so they cannot appear in results. They are not always a problem: some are excluded on purpose. What matters is the pages that should be visible and are missing. The report in Google Search Console tells you the reason for each one, and below we go through the common ones with fixes.
In Search Console, open Page indexing. The report splits URLs into two groups: indexed and not indexed. For the non-indexed ones, Google shows a reason for each group of URLs, with a list of examples.
Three things to keep in mind:
If you have not worked with the tool yet, the Google Search Console beginner's guide explains the basic reports. For issues of this kind, which belong to technical SEO, follow the steps below.
The wording in the interface may vary slightly, but the meaning is the one in Google's documentation.
| Status | What it means | Frequent cause | What you do |
|---|---|---|---|
| Crawled – currently not indexed | Google fetched the page but did not index it | Thin or redundant content, weak signals | Improve the page, link to it internally, then wait |
| Discovered – currently not indexed | Google knows the URL but has not crawled it yet | New site, weak signals, slow server | Internal links, sitemap, check server speed |
| Excluded by 'noindex' tag | The page explicitly asks not to be indexed | Noindex set on purpose or by mistake | Remove noindex if the page should be indexed |
| Blocked by robots.txt | Crawling is forbidden | An overly broad Disallow rule | Fix robots.txt |
| Alternate page with proper canonical tag | The page points to another URL as canonical | Page variants, parameters | Usually nothing: it is correct |
| Duplicate without user-selected canonical | Duplicate with no declared canonical | Same content on several URLs | Declare your preferred canonical |
| Duplicate, Google chose different canonical than user | Google picked a different canonical | Content is not similar enough | Check for contradictory signals |
| Soft 404 | The page looks empty or like an error but returns 200 | Pages with no content, error messages | Return a 404 or add content |
| Not found (404) | The URL does not exist | Broken link, deleted page | Redirect if there is an equivalent |
| Page with redirect | The URL redirects elsewhere | Intentional redirect | Usually nothing: it is correct |
| Server error (5xx) | The server returned an error | Overload, configuration | Fix the server |
| Redirect error | A redirect loop or a chain that is too long | Redirect chain | Fix the chain |
Read the first column as a diagnosis, not as blame. Below we expand on the groups that raise the most questions.
These are the most common statuses and the most commonly misunderstood.
Discovered – currently not indexed means Google learned of the page (from a sitemap or a link) but has not crawled it yet. The documentation says crawling is usually postponed to avoid overloading the site. On a small, new site, it is often just a matter of time. What you can do: link the page from pages that are already indexed, add it to a clean sitemap and check that the server responds quickly.
Crawled – currently not indexed is more annoying: Google visited the page and decided, at least for now, not to keep it. The documentation notes the page may be indexed later without you resubmitting it. The reason usually lies in the page, not in the technology:
The fix is to improve the page: a direct answer, original information, examples, clear structure, internal links from strong pages. The principles are in the article on internal linking and site architecture. If you have many similar pages, consolidate them into one better page.
The tag <meta name="robots" content="noindex"> or the header X-Robots-Tag: noindex excludes the page. It often appears after launches, when test configuration stays behind. Check the page source, not only what the SEO plugin displays.
An overly broad Disallow forbids crawling. Note the difference: robots.txt prevents crawling, it does not guarantee exclusion from the index. If you want a page to stay out of results, use noindex and leave the page crawlable, otherwise Google cannot read the tag. For a correct file you can use the robots.txt generator.
If a page declares another URL as canonical, Google treats it as an alternate and does not index it. That is right for parameter variants; it is wrong when the canonical was copied from another template. The full explanation is in the article on canonical tags.
URLs returning 404 are not indexed (and should not be). The problem appears when internal links or the sitemap point to them. 5xx errors indicate server problems; if they repeat, Google crawls less often. Clean up broken links and, where an equivalent exists, use a 301 redirect; details are in the redirect guide.
Pages that read like "nothing found" or are nearly empty but return 200. Either make them return 404 or give them real content. It often happens with empty category pages and internal search pages.
"Page with redirect" is normal for old URLs. The problems are chains and loops ("Redirect error"): each old address should lead straight to its final destination.
If the text only appears after rendering, Google may see an empty page. See the article on JavaScript SEO.
On a site with hundreds or thousands of URLs, you cannot fix everything at once, and you do not need to. Order them by business impact:
Say the report shows 120 URLs under "Crawled – currently not indexed." You open the list and see that 90 are tag pages with two articles each, 20 are very short old articles, and 10 are important service pages. You do not submit all 120 for indexing. The rational decision: services first (improve them and link them from the homepage), merge or expand the old articles, and set tag pages to noindex if they have no value of their own. The result is not a report with zero non-indexed URLs, but a site where what matters is indexed and what does not matter is excluded on purpose.
This kind of decision is easier when you have defined what you want to measure; see the article on SEO KPIs.
A "clean" report does not mean every URL is indexed. These are normal, among others:
noindex that you set (thank-you pages, cart, account, internal search results);On an online store, filters and parameters often produce thousands of correctly excluded URLs. What you need to watch is that the important category, product and article pages are indexed. As for crawl budget, just remember that crawl budget problems appear mainly on very large sites.
Google states explicitly that it does not guarantee indexing of all pages. Even a technically flawless site may have pages Google does not consider useful enough. Also:
A non-indexed page is a symptom, and the reason is in the report. Proceed like this:
If you have hundreds of URLs in unclear situations, or important pages that will not get into Google, a technical SEO audit identifies and prioritizes the causes. You will find more guides on the blog.
Frequently asked questions
For a new page, a few days to a few weeks is normal, especially on a young site. If an important page is still not indexed after two to three weeks, check the reason in URL Inspection. Do not request indexing dozens of times: it does not speed things up.
Only for important or recently changed pages, after confirming they have no technical problems. The request does not guarantee indexing; it only adds the URL to the crawl queue. For a whole site, a correct sitemap and internal links to new pages work better.
Sometimes, but it is not a goal in itself. Thin, duplicate or useless pages can be merged, improved or removed with a 404 or 410. But intentionally excluded pages, such as filters or canonicalized variants, should not be deleted just to make the report look cleaner.
They can come from URL parameters, filters, internal search pages, variants with and without a trailing slash, or old addresses. Google finds them through links or the sitemap. If they have no value, control them with canonicals, noindex, or by removing the links that lead to them.
Google indexed the URL even though robots.txt blocks crawling, usually because it found the address through internal or external links. If you want it out of results, use noindex and allow crawling so Google can read the tag. Robots.txt alone does not prevent indexing.
Related service
Keep reading
Canonical tag explained: how Google handles duplicate content, when to use rel=canonical, code examples, store scenarios and mistakes to avoid.
301 redirect or 302? The difference, how Google treats each, Apache and nginx code, a redirect map and the mistakes that cost you traffic.
Structured data schema in JSON-LD: which types are worth adding, code for an organization, article and product, how to validate and what Google retired.
Send us your website address and we’ll reply with a free initial analysis and a concrete SEO strategy — no strings attached.