Feature

A sitemap submits candidates, not commands

Why URLs in a successful sitemap can remain unindexed, how to identify the actual stage, and which fixes match each Search Console status.

Impetuous · · 4 Min Read

A URL in an XML sitemap is a candidate for discovery and crawling. It is not an instruction to add that URL to Google’s index.

That distinction explains a common operational mismatch: the sitemap is valid, Search Console reports Success, and some listed pages remain unindexed. “Success” means Google fetched and read the sitemap without errors. Even URLs counted as discovered are not guaranteed to be crawled or indexed, according to Google’s Sitemaps report documentation.

A sitemap only acts at the discovery layer

Google describes Search as three stages: crawling, indexing and serving results. A sitemap helps Google discover URLs before those later decisions occur. Google may then choose whether and when to crawl a discovered URL; after crawling it, Google may decide not to index it. Even an indexed page is not guaranteed to appear for a particular query. Google explicitly says it does not guarantee crawling, indexing or serving, even when a page follows its technical guidance (how Google Search works).

Sitemap inclusion therefore means:

  • the publisher wants Google to know about the URL;
  • the publisher considers the URL important enough to crawl;
  • Google has another route by which to discover it.

It does not mean:

  • Googlebot has fetched the page;
  • the response is technically indexable;
  • Google accepts it as the canonical URL;
  • the content warrants a separate index entry.

Sitemap inclusion is also only a weak canonicalization signal. Redirects and rel="canonical" annotations are stronger, and Google can select a different canonical when it considers multiple URLs duplicates (Google’s canonicalization guidance).

Why a submitted URL can remain unindexed

The exclusion usually occurs after sitemap processing.

Google has not crawled it yet

A Discovered – currently not indexed status means Google knows the URL but has not crawled it. Google says this typically occurs when crawling was expected to overload the site and was rescheduled. On large or rapidly changing sites, crawl demand also depends on factors including site size, page quality, relevance, update frequency and the volume of duplicate or unimportant URLs (crawl budget guidance).

A sitemap does not compensate for an inflated URL inventory, unstable servers or weak site architecture. Important pages should also have crawlable internal links; Google recommends that every page a publisher cares about be linked from at least one other page on the site (link guidance).

Google crawled it but did not index it

Crawled – currently not indexed means the fetch happened but Google did not add the page to its index. Google says the page may or may not be indexed later and that resubmitting it for crawling is unnecessary (Page indexing report).

At this point, investigate the page rather than the sitemap. Compare it with pages targeting the same intent. Check whether the primary content appears in Google’s rendered output, whether the page offers meaningful information of its own, and whether templated pages differ only by a few fields.

The URL is not technically indexable

For basic eligibility, Google says Googlebot must not be blocked, the page must return HTTP 200, and it must contain indexable content (technical requirements). A sitemap cannot override:

  • a noindex meta tag or X-Robots-Tag header;
  • authentication or access failures;
  • 4xx and 5xx responses;
  • redirects to another URL;
  • content Google cannot process as an indexable page.

A robots.txt block creates a particularly confusing case: it can stop Googlebot from fetching the page and therefore from seeing a page-level noindex rule. Google must crawl a page to read that rule (noindex documentation).

Google chose another canonical

Duplicate and alternate URLs should not all be indexed. If a sitemap lists parameter variants, HTTP and HTTPS versions, or near-identical landing pages, Google may index one representative URL and exclude the others. Align the sitemap, internal links, redirects and canonical annotations around the same preferred URL. For a fuller choice between exclusion and consolidation, see canonical tags versus noindex.

Diagnose the stage before changing the page

Use this sequence for a representative sample of affected URLs:

  1. Confirm sitemap processing. In the Sitemaps report, distinguish a successfully read file from the indexing status of its URLs.
  2. Filter the Page indexing report by sitemap. Group exclusions by reason instead of treating every unindexed URL as the same defect.
  3. Inspect individual URLs. URL Inspection shows the latest indexed state, crawl details, indexing permission, user-declared canonical and Google-selected canonical. The live test checks current accessibility and potential indexability, but it cannot predict canonical selection or guarantee indexing (URL Inspection documentation).
  4. Apply the matching fix. Repair access and response errors; remove accidental noindex; consolidate duplicates; improve pages that lack distinct value; add crawlable internal links; or reduce low-value URL generation.
  5. Measure by cohort. Keep separate sitemaps for useful groups such as articles, products or publication periods. Track the share indexed and the lag from publication to first crawl. This makes a systemic problem visible without implying that every submitted URL deserves indexing.

Do not repeatedly request indexing as a substitute for diagnosis. Google says repeated recrawl requests for the same URL do not make crawling faster (recrawl guidance). For a broader operating approach, use the content publishing and crawl budget workflow.

The practical rule is simple: use a sitemap to declare a clean inventory of preferred URLs, then evaluate discovery, crawling, technical eligibility, canonicalization and content selection as separate stages. Sitemap health is evidence that the submission channel works—not that every submitted page belongs in the index.