Why does Search Console report fewer indexed pages than my sitemap lists?
Because a sitemap is a suggestion and the index is a decision. Submitting a URL tells Google the page exists. It does not oblige Google to crawl it, and crawling does not oblige Google to index it. The gap between those two numbers is normal, and its composition is what matters.
This question comes up on almost every large content site I audit. Somebody exports a sitemap, counts the URLs, opens the page indexing report, and finds hundreds of pages missing. The instinct is panic, then a request to index everything.
That instinct is usually wrong. Most of the gap is working as intended, and the useful work is separating the part that is fine from the part that is not.
What is a sitemap actually for?
A sitemap helps Google discover URLs. Google's documentation is direct about the requirement behind it: Google needs a way to find a page in order to crawl it, which means the page must be linked from a known page or from a sitemap. Discovery is the job. Nothing beyond that is promised.
That single sentence reframes the whole exercise. A sitemap is a discovery aid, not an index request and not a quality signal. A page listed in your sitemap and linked from nowhere else on your site has been discovered and abandoned, which is a real problem and not one a sitemap solves.
So the first question for any missing page is not why Google refused it. It is whether the page has any other route in. Pages that exist only in a sitemap are the ones most likely to sit uncrawled.
Which report should you be reading?
The page indexing report for patterns, and the URL Inspection tool for individual pages. Google's own help says the page indexing report is not used to investigate the index status of specific pages, and points to URL Inspection for that. People routinely use it the other way round and get confused.
There is a scale note worth knowing too. Google's documentation says that if your site has fewer than 500 pages, you probably do not need this report at all, and suggests using site queries to check whether key pages are indexed instead. Small sites can spend weeks studying a report designed for problems they do not have.
The report also has a display limit. Google notes that the list of example URLs in the report is limited to 1,000 items. On a site with thousands of pages, what you are looking at is a sample, so treat proportions as meaningful and individual absences as unremarkable.
What do the most common not-indexed reasons really mean?
Four reasons cover most of the gap, and Google defines each of them precisely. Read the definitions rather than guessing from the label, because two of them sound like errors and are not, and one of them sounds patient when it is actually a signal about your site's capacity.
Crawled and currently not indexed means, in Google's words, that the page was crawled but not indexed, that it may or may not be indexed in the future, and that there is no need to resubmit the URL. This is a judgement about the page, and resubmitting does not change a judgement.
Discovered and currently not indexed is different and more interesting. Google's description says the page was found but not crawled yet, typically because Google wanted to crawl the URL but expected that to overload the site, so it rescheduled the crawl. The documentation adds that this is why the last crawl date is empty on the report.
The two canonical states complete the picture. Duplicate without user-selected canonical means Google found a duplicate with no preferred version declared, chose another page as canonical, and will not serve this one. Google's help explicitly frames that as working as intended rather than an error, because duplicate pages are not served.
Which of these gaps are actually fine?
The canonical ones, almost always. Google's documentation says that a page marked as an alternate with a proper canonical tag correctly points to the canonical page, which is indexed, so there is nothing you need to do. Seeing a large number in that bucket is a sign your canonicals are working, not failing.
Paginated archives, tag pages, and filtered views often land here too, and that is the right outcome. You do not want fifty near identical category pages competing with the articles they list. I covered the canonical mechanics on Webflow in canonical tags and duplicate content.
Old 404s are also less alarming than they look. Google's help notes that Googlebot will probably keep trying a 404 URL for some period, that there is no way to tell Googlebot to permanently forget a URL, and that it will crawl it less and less often. A list of 404s from a migration two years ago is history, not a live problem.
Which gaps are worth fixing?
Three. Pages you care about sitting in discovered and not crawled. Pages you care about sitting in crawled and not indexed. And anything in your sitemap that should not be there at all, which quietly teaches Google that your sitemap is unreliable.
The discovered bucket is the one I treat as urgent on large sites, because Google's stated reason is load related rescheduling. That points at site performance and at how many low value URLs you are asking Google to work through, not at the individual page.
The crawled and not indexed bucket is a content problem wearing a technical costume. If Google crawled a page and declined it, the honest questions are whether the page says anything the rest of the web does not, and whether it is reachable from anywhere a reader would actually travel. I unpacked the distinction between the two stages in crawling versus indexing.
How do you diagnose a single URL properly?
Use URL Inspection, one page at a time, and read the canonical Google selected rather than the one you declared. Those two disagreeing explains a large share of mysterious absences, and no amount of studying the summary report will surface it.
Then check the route in. Count the internal links pointing at the page from pages that are themselves indexed. A page with one link from a paginated archive on page eleven is technically linked and practically invisible.
Finally check the response the crawler gets rather than the response your browser gets. Redirect chains, parameter variants and inconsistent trailing slashes produce sitemap entries that quietly resolve to something other than the URL you submitted.
What does this look like on a large CMS blog?
On a blog with a thousand or more posts, expect a permanent gap and manage the composition rather than chasing zero. My working target is that every page I would defend in a meeting is indexed, and that the unindexed remainder is made up of duplicates, alternates, and pages that genuinely do not deserve a slot.
The practical control is the sitemap itself. Keep out anything you would not want a reader to land on from a search result, because every junk URL in there competes for the same crawl attention as the posts you care about. My rules for that on Webflow are in controlling what gets indexed.
Review it on a schedule rather than in a crisis. A monthly glance at the proportions in each bucket tells you whether something changed. A yearly panic tells you nothing, because you have no baseline to compare against.
What should you do next?
Open the page indexing report, note the size of each not-indexed bucket, and sort them into fine and not fine using Google's own definitions rather than the names. Then take the ten pages you most want indexed and inspect them individually. That is a morning of work and it usually ends the mystery.
If the answer turns out to be discovered and not crawled across a large share of your site, stop optimising individual pages and look at what you are asking the crawler to do overall. That is a site architecture conversation, not a page conversation.
If you want a second pair of eyes on a coverage report that does not make sense, send me the proportions and a handful of example URLs and I will tell you which part is worth your time.
Get found, cited and the back office automated
Let's make your site the source AI engines quote and wire up the systems behind it.
Read more blogs
Let's get your website found and cited by AI
Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.