Skip to content
Searcle Book a demo

Build a Store Sitemap That Search Engines Can Actually Use

Nina Okonkwo

An ecommerce sitemap should not be a database export disguised as an SEO file. Its job is to identify the canonical, indexable URLs that a store genuinely wants search engines to discover—not every URL the platform can generate.

The strongest setup is automatically maintained, divided into useful groups, technically valid, and monitored for conflicts. It supports discovery and crawling, but it cannot compensate for weak internal linking, duplicate pages, incorrect canonicals, blocked URLs, thin content, or poor catalog architecture. Nor does it guarantee indexing, rankings, traffic, or sales.

What an ecommerce sitemap is—and which type you actually need

“Ecommerce sitemap” can describe three different things:

  1. An XML sitemap is a machine-readable file that identifies important store URLs for search engines.
  2. An HTML sitemap is a browsable webpage containing organized links for visitors and crawlers.
  3. A visual sitemap is a planning diagram used by designers, developers, marketers, and SEO teams to map the store’s architecture.

These formats are complementary, not interchangeable.

A visual sitemap is useful before development, redesign, or migration. It can group product pages, categories, customer-account areas, payment steps, editorial content, and marketing landing pages so teams can assess how the store should fit together. It is an internal planning artifact; it is not submitted to search engines. Visual ecommerce templates commonly organize pages around functions such as products, account administration, payments, and marketing, as illustrated by this ecommerce architecture template from Moqups.

An XML sitemap identifies the pages and other files the store considers important. It may also provide information such as when a page was last modified, where an associated image is located, or which alternate-language versions exist. Search engines read this structured file to discover and crawl site content more efficiently.

An HTML sitemap is a normal webpage. It can give visitors another route to important categories, brands, buying guides, policies, or other content that does not fit comfortably in the main menu. It may also strengthen internal navigation, but it is optional. A store with clear menus, categories, breadcrumbs, contextual links, and search may not need one.

Type Primary audience Format Typical location Maintenance Primary use
XML sitemap Search engines XML file or sitemap index Usually /sitemap.xml or another declared XML URL Preferably automatic Identify preferred URLs and related metadata
HTML sitemap Visitors and crawlers Standard HTML webpage Often linked from the footer Manual or CMS-generated Provide an additional navigation route
Visual sitemap Internal teams Diagram, flowchart, or planning board Project documentation Updated during planning and redesign Design information architecture and page relationships

XML sitemaps are especially useful for:

  • Large product catalogs
  • New stores with few external links
  • Stores with deep, complicated, or imperfect internal linking
  • Multilingual or multi-country storefronts
  • Sites containing many product images, videos, or other media
  • Catalogs that add, update, retire, or redirect URLs frequently

That does not mean every small store technically requires one. Google says a comprehensively linked site may not need a sitemap, particularly when it is small and has little media content. In that guidance, Google uses approximately 500 index-worthy pages as a working definition of a small site—not as a universal threshold or a point at which sitemap rules suddenly change. Google also emphasizes sitemaps for large, new, complex, and media-heavy sites in its overview of when sitemaps are useful.

The central limitation should remain clear: an ecommerce sitemap can make discovery and crawling easier to manage. It does not guarantee that a listed URL will be crawled, indexed, ranked, visited, or converted into a sale.

The ecommerce URL decision matrix: include, exclude, or review

Use one governing test for every sitemap entry:

List a URL when it is fully qualified, canonical, indexable, returns HTTP 200, and is intentionally eligible to appear in search results.

“Fully qualified” means including the complete protocol and hostname, such as:

https://store.example.com/products/wool-coat

It does not mean:

/products/wool-coat

The sitemap should represent preferred search-result URLs. It should not automatically contain every address that a crawler can technically request.

URL type Default decision Conditions or reason
Canonical product detail page Include Indexable, returns 200, and intended for search
Category or collection page Include Canonical, useful, and intentionally indexable
Editorial article or buying guide Include Important, canonical, and eligible for search
Shipping, returns, or warranty page Include Useful to searchers and intentionally indexable
Brand page Review Include only if substantial, canonical, and independently useful
Product variant Review Include only if independently canonical and intended to appear as its own result
Curated facet or filter landing page Review Must have distinct value, a canonical URL, and intentional indexability
Paginated category URL Review Depends on architecture, canonical treatment, and product reachability
Temporarily out-of-stock product Review May remain if useful, canonical, indexable, and expected to return
Seasonal product page Review Depends on whether the page will return and retain long-term value
Discontinued product Review Retain, redirect, or retire according to its continuing usefulness
Cart Exclude Transactional system page
Checkout Exclude Transactional system page
Customer account or login area Exclude Private or utility content not intended for search
Internal-search result Exclude Often dynamic, duplicative, or low-value
Tracking-parameter URL Exclude Duplicate route rather than the preferred destination
Redirected legacy URL Exclude List the final canonical destination instead
Error or server-error URL Exclude Does not return a valid page
noindex URL Exclude Sitemap inclusion contradicts the indexing directive
Noncanonical duplicate Exclude The canonical version should be listed instead
Staging or test URL Exclude Not intended for public search

For example, include:

https://store.example.com/products/wool-coat

Exclude:

https://store.example.com/cart
https://store.example.com/checkout
https://store.example.com/search?q=coat
https://store.example.com/products/wool-coat?utm_source=newsletter
https://store.example.com/products/old-wool-jacket

The final example should be excluded if it redirects to the current wool-coat page. The sitemap should list that final destination, not use a redirected legacy URL as a substitute.

Filtered and parameterized URLs are usually excluded because faceted navigation can produce large numbers of overlapping combinations:

/category/coats?color=black&size=m&material=wool

Many such combinations show substantially the same products and have little independent value. An exception may be appropriate when a filtered view has deliberately been turned into a stable landing page—for example, a canonical collection for “women’s black wool coats” with useful content and genuine independent search value.

Google may attempt to crawl sitemap URLs exactly as provided, which is why malformed URLs, redirects, and duplicate alternatives undermine file quality. Google recommends using fully qualified preferred canonical URLs in its technical sitemap construction guidance.

A clean sitemap is therefore a declaration of intent: these are the URLs the store believes deserve consideration for search results.

Variants, filters, pagination, and changing inventory: make conditional decisions

Variants, facets, pagination, and product-lifecycle URLs cannot be handled reliably with one universal rule. For each URL, evaluate canonical status, indexability, usefulness, technical response, internal accessibility, and independent search value.

Product variants

When color, size, or configuration variants consolidate to one canonical product detail page, list only that primary page.

For example, these variant routes might all canonicalize to the main product:

/products/wool-coat?color=black
/products/wool-coat?color=camel
/products/wool-coat?size=medium

In that case, the sitemap should contain:

https://store.example.com/products/wool-coat

A variant may deserve its own entry when all of the following are true:

  • It has its own canonical URL.
  • It has sufficiently distinctive content or product characteristics.
  • It is intentionally indexable.
  • It offers independent value as a search result.
  • Internal links treat it as a meaningful destination rather than a temporary state.

A separately marketed “red leather travel bag” might qualify if it has distinctive imagery, availability, copy, and product characteristics. A size selector that changes only a query parameter usually will not.

Faceted navigation

Ordinary filter combinations should generally stay out of the sitemap when they duplicate categories or generate low-value permutations.

Curated facets deserve individual review. A facet can qualify when it has:

  • A stable, preferred URL
  • A self-referencing canonical
  • Indexable content
  • A useful product set
  • Distinctive supporting copy where appropriate
  • Internal links from relevant parts of the store
  • A reason to exist independently in search

Do not include a facet merely because keyword research found a phrase. The page must also function as a coherent landing page rather than a thin filtered state.

Pagination

One common practitioner approach is to exclude paginated category URLs from the sitemap and let crawlers discover them through internal links. That can be sensible for some implementations, but it is not a universal search-engine requirement.

The important architectural question is whether crawlers can reach deeper products through crawlable links. If page two, page three, and later category states are the only routes to older or less prominent products, removing or blocking those routes can make discovery harder. Sitemap treatment should therefore be decided alongside category links, “load more” behavior, infinite scrolling, canonicals, and product-link accessibility.

A sitemap is not a substitute for crawlable pagination. Conversely, a paginated URL should not be included merely to conceal an internal-linking problem.

Temporarily out-of-stock products

Stock status alone does not decide sitemap inclusion.

An out-of-stock product can remain eligible when its page:

  • Returns HTTP 200
  • Is canonical and indexable
  • Still gives shoppers useful information
  • Is expected to return to inventory
  • Remains a reasonable search destination

If the page becomes a thin dead end with no useful information and no expected return, reconsider whether it should remain indexable and listed.

Discontinued products

Use a lifecycle decision rather than deleting every discontinued product automatically:

  1. Retain the URL when the page still serves a useful search purpose—for example, documentation, compatibility information, replacement guidance, or support for previous buyers.
  2. Redirect it when a genuine replacement exists and satisfies substantially the same need.
  3. Remove it from the sitemap when it no longer deserves indexing, redirects elsewhere, or no longer returns a valid indexable response.

Do not redirect every discontinued product to a generic category merely to avoid a missing page. A redirect should represent a meaningful replacement.

Seasonal products

A temporarily dormant holiday product may keep its URL if the page is designed to return each season and remains part of the long-term canonical strategy. A permanently retired seasonal line should be treated like any other discontinued product.

Preserving a stable annual page can be practical when the same product or collection returns.

Conditional-URL review checklist

For every variant, facet, paginated page, unavailable product, or seasonal URL, ask:

  • Is it canonical?
  • Is it indexable?
  • Does it return HTTP 200?
  • Is it useful now or reasonably expected to become useful again?
  • Does it have independent search value?
  • Can crawlers reach it through internal links?
  • Does its sitemap treatment agree with its long-term lifecycle strategy?

The sitemap cannot repair poor lifecycle handling. Canonical tags, redirects, status codes, robots controls, internal links, and on-page usefulness must be managed separately.

Design a sitemap structure that scales with the catalog

A single sitemap may contain no more than 50,000 URLs and no more than 50 MB of uncompressed data. Both constraints apply, so a file must comply with both documented limits (Google Search Central).

A sitemap index solves this scaling problem by pointing to multiple child sitemap files. Beyond protocol compliance, segmentation can make reporting and troubleshooting much easier.

Useful divisions include:

  • Products
  • Categories or collections
  • Brands
  • Editorial articles and buying guides
  • Images or other media
  • Languages
  • Countries or regional domains
  • Product creation dates or catalog ranges

For a modest store, the structure might be:

/sitemap.xml
  /products.xml
  /categories.xml
  /articles.xml

Here, sitemap.xml acts as the sitemap index and points to the three child files.

A large catalog could use:

/sitemap-index.xml
  /products-1.xml
  /products-2.xml
  /products-3.xml
  /categories.xml
  /brands.xml
  /articles.xml

Every child file must contain no more than 50,000 URLs and no more than 50 MB uncompressed. If the product catalog expands, the generator can create another product file without changing the category or editorial files.

This segmentation should aid diagnosis. If discovered category URLs suddenly fall while product counts remain stable, a separate category sitemap narrows the investigation. If one language file starts returning canonical conflicts, the issue can be examined without mixing it with every market.

A multilingual store might use:

/sitemap-index.xml
  /en-us-products.xml
  /en-gb-products.xml
  /fr-fr-products.xml
  /de-de-products.xml

Split by language or country only when it improves maintenance or reporting. The file division does not establish alternate-language relationships by itself. Hreflang relationships still need to be explicitly declared as complete, reciprocal groups.

There is no documented requirement to keep sitemap files below 10,000 URLs for better performance. Smaller operational files can be useful, but that is a management choice—not a search-engine rule. Use the documented limits unless monitoring, deployment, or troubleshooting needs justify narrower internal thresholds.

Google supports XML, RSS, Atom, and plain-text sitemap formats. XML remains the most versatile ecommerce choice because it can represent localized versions and media information, while plain-text files contain URLs only and feeds generally emphasize recent content.

Generate valid XML and automate catalog updates

Manual sitemap maintenance becomes fragile as soon as a store has more than a few dozen URLs or changes frequently. Google recommends automatic generation for sites beyond that scale, preferably through the website software or CMS.

Automation matters because products, collections, articles, images, redirects, and availability states are continuously added or revised. The sitemap should reflect those changes without relying on someone to edit XML by hand.

Use the platform’s native generator when it provides enough control. Otherwise, use a maintained CMS plugin, ecommerce module, or custom catalog process.

A platform-neutral implementation sequence is:

  1. Identify the platform’s native sitemap generator.
  2. Retrieve the main sitemap and every child file.
  3. Inspect which page types and URL patterns it includes.
  4. Define exclusions for system, duplicate, redirected, and nonindexable URLs.
  5. Confirm what events trigger regeneration.
  6. Test that new, updated, redirected, and retired products are handled correctly.
  7. Validate the XML and crawl samples from every child file.
  8. Establish monitoring for retrieval failures, URL counts, and signal conflicts.

A minimal product sitemap can look like this:

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://store.example.com/products/wool-coat</loc>
    <lastmod>2026-08-15</lastmod>
  </url>
</urlset>

The URL must be fully qualified, and XML values must be entity escaped. For example, an ampersand inside a URL must be represented as &amp;, although parameterized tracking URLs generally should not be listed in the first place.

A minimal sitemap index could be:

<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap>
    <loc>https://store.example.com/products.xml</loc>
    <lastmod>2026-08-15</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://store.example.com/categories.xml</loc>
    <lastmod>2026-08-12</lastmod>
  </sitemap>
</sitemapindex>

Validate the following:

  • URLs are fully qualified.
  • Every URL is the preferred canonical destination.
  • Listed pages return HTTP 200.
  • Listed pages are indexable.
  • Files use UTF-8 encoding.
  • XML values are properly escaped.
  • Files comply with the URL-count and uncompressed-size limits.
  • Child sitemap URLs are accessible.
  • Generated dates are accurate.
  • Redirects, errors, and noncanonical alternatives are absent.

Do not spend time manipulating priority or changefreq for Google. Google says it ignores both fields. A generator may still output them for compatibility, but they should not distract from URL quality and accurate modification dates.

Shopify implementation

Shopify automatically creates a root-level /sitemap.xml for applicable store domains. The main file links to separate product, collection, blog, and webpage sitemap files. Shopify updates the generated files when relevant store content changes and includes primary product images, according to the platform’s official sitemap documentation.

Do not assume every ecommerce platform behaves like Shopify. PrestaShop may rely on a module and scheduled task. Other systems may use plugins, custom routes, cron jobs, or application-level generation. Inspect the current documentation and actual output for the installed platform and version.

Avoid manually editing a generated file until you know how the generator works. A platform update, cache refresh, catalog event, or scheduled process may overwrite the change. Where exclusions are needed, configure the generator or upstream URL logic rather than repeatedly patching its output.

Use lastmod, image data, and hreflang without creating false signals

Sitemap metadata is useful only when it remains trustworthy. More fields do not automatically make a sitemap better.

Use lastmod for meaningful, verifiable changes

Google may use lastmod when the value is consistently and verifiably accurate. It should represent the page’s meaningful last modification, not the time at which the sitemap happened to regenerate.

Changes that may justify a new lastmod value include:

  • Material price changes
  • Availability changes
  • Substantial product-copy revisions
  • New or materially changed product images
  • Meaningful structured-data changes
  • Significant specification or compatibility updates

That does not mean every tiny inventory fluctuation, punctuation correction, review, or templated layout change must reset the date. Define meaningful change according to the store’s content model, then implement that definition consistently.

A weak implementation assigns the current timestamp to every URL whenever the generator runs. That creates fabricated freshness and makes the field less informative.

As noted above, Google ignores priority and changefreq. Accurate URLs and honest modification dates deserve more attention than tuning ignored values.

Represent product images when useful

XML can identify images associated with product pages. A basic image entry uses image:loc for the image location:

<?xml version="1.0" encoding="UTF-8"?>
<urlset
  xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
  xmlns:image="http://www.google.com/schemas/sitemap-image/1.1">
  <url>
    <loc>https://store.example.com/products/wool-coat</loc>
    <image:image>
      <image:loc>https://store.example.com/images/wool-coat-black.jpg</image:loc>
    </image:image>
  </url>
</urlset>

An image-heavy store can attach image extensions to product entries or organize images separately. A separate image sitemap is an operational option, not something that can be assumed to outperform image entries in the main product files.

Whichever structure you choose, image locations should be valid and accessible. The sitemap cannot compensate for broken image URLs, blocked resources, weak image context, or poor product pages.

Implement complete hreflang groups

Alternate-language or country URLs can be represented in XML sitemaps or declared in HTML. Sitemap-based hreflang can be useful when centralized generation is easier than editing page templates.

A simplified group might connect:

https://store.example.com/en-us/products/wool-coat
https://store.example.com/en-gb/products/wool-coat
https://store.example.com/fr-fr/produits/manteau-laine

The group must be complete and reciprocal. Every applicable version should participate in the same alternate set, and each relationship should be confirmed by the corresponding versions. Do not place isolated or one-directional alternates into different files and assume the file names establish the relationship.

Shopify adds multilingual store languages to sitemap files for applicable domains, and international domains receive their own sitemaps. That automation still needs to be checked against the store’s actual domain, language, canonical, and redirect setup.

Use this metadata checklist:

  • lastmod reflects meaningful, verifiable changes.
  • Sitemap regeneration does not fabricate freshness.
  • Image locations are valid and accessible.
  • Hreflang groups are complete and reciprocal.
  • Alternate URLs agree with the canonical and domain strategy.
  • No process relies on priority or changefreq to control Google.

Publish, submit, and connect the sitemap to the rest of technical SEO

Place the main sitemap or sitemap index at the root when the platform permits it:

https://example.com/sitemap.xml

Root placement gives the file the broadest natural scope. When a sitemap is discovered outside Search Console, Google generally treats its scope as the descendants of the directory containing it; submitting it through Search Console removes that directory-scope constraint.

A sitemap can also be referenced in robots.txt so crawlers can locate it explicitly:

Sitemap: https://example.com/sitemap.xml

If there are multiple indexes, list each applicable location or reference the index that connects the child files.

To submit a sitemap in Google Search Console:

  1. Verify the correct website property.
  2. Open the Sitemaps report.
  3. Enter the sitemap or sitemap-index URL.
  4. Submit it.
  5. Review processing status, last-read information, discovered counts, and retrieval errors.

For a Shopify store using international domains, verify each relevant domain and submit that domain’s sitemap separately. The storefront must also be accessible without active password protection for Google to crawl it. Shopify documents these domain and accessibility requirements in its instructions for finding and submitting store sitemaps.

Submitting a sitemap is only one step in a longer sequence:

  1. Discovery: The search engine knows the URL exists.
  2. Crawling: The search engine requests and fetches the URL.
  3. Indexing: The search engine evaluates the page and may select it for its index.
  4. Ranking: An indexed page may be shown in a particular position for a particular query.

Submission does not guarantee any later stage, and Google does not promise a particular crawling or indexing timeline.

XML sitemaps should complement—not replace—internal navigation. Every important page should still have a logical route through elements such as:

  • Main and secondary menus
  • Categories and subcategories
  • Breadcrumbs
  • Related products
  • Brand hubs
  • Buying guides
  • Contextual editorial links
  • Crawlable pagination where needed

A listed URL should send consistent signals across the system:

Signal What should agree
Sitemap Lists the preferred URL
Canonical tag Identifies the same preferred URL
HTTP status Returns 200 for an active indexable page
Redirects Do not send the listed URL elsewhere
Robots directives Permit crawling and indexing as intended
Hreflang References valid canonical alternates
Internal links Point mainly to the preferred version
Content Provides a useful, distinct destination

A sitemap cannot fix duplicate content, thin product descriptions, poor architecture, blocked resources, contradictory canonicals, or weak internal linking. It can expose those problems during an audit, but the corrections must happen in the underlying store.

Audit and troubleshoot an ecommerce sitemap

A sitemap audit should test both the XML files and the pages they declare. A file can be syntactically valid while containing hundreds of redirects, noncanonical pages, or obsolete products.

Use this repeatable workflow.

1. Retrieve the complete sitemap set

Find:

  • The sitemap referenced in robots.txt
  • The submitted Search Console sitemap
  • Every sitemap index
  • Every child sitemap
  • Language- or country-specific files
  • Image or media files, where applicable

Confirm that each file is accessible to a normal unauthenticated request. Flag empty files, authentication barriers, server errors, redirect loops, or references to missing children.

2. Validate the XML

Check:

  • Valid XML structure
  • Correct namespace declarations
  • UTF-8 encoding
  • Properly escaped values
  • Fully qualified URLs
  • Valid date formats
  • No truncated or malformed entries

A browser displaying the XML does not prove that the file is valid. Use an XML validator or a crawler capable of parsing sitemap indexes and child files.

3. Count URLs and measure file sizes

Record the compressed and uncompressed sizes where relevant, plus the number of URLs in each child file. Confirm that every file contains no more than 50,000 URLs and no more than 50 MB uncompressed.

Compare current counts with expected catalog behavior. A sharp drop may indicate a broken generation job; a sudden increase may indicate that filters, parameters, or duplicates have entered the file.

4. Crawl every listed URL

For each URL, record:

  • HTTP response
  • Redirect destination
  • Indexability
  • Canonical target
  • Robots accessibility
  • Meta robots directive
  • X-Robots-Tag where applicable
  • Internal-link count or reachability
  • Page type
  • Language or country
  • Sitemap lastmod value

The desired baseline is a 200 response, an indexable page, and canonical alignment with the URL listed in the sitemap.

5. Find conflicting and stale entries

Common defects include:

  • Empty or inaccessible sitemap files
  • Malformed XML
  • Relative URLs
  • HTTP URLs where HTTPS is canonical
  • Hostname inconsistencies
  • Noncanonical duplicates
  • noindex pages
  • Robots-blocked URLs
  • Redirects
  • 404 responses
  • 5xx responses
  • Deleted products that remain listed
  • Staging and test URLs
  • Tracking parameters
  • Internal-search URLs
  • False lastmod dates
  • Child files exceeding protocol limits

Prioritize systemic errors. If every product URL redirects from a trailing-slash version to a non-slash version, fix the generator rather than handling entries individually.

6. Compare the files with Search Console

Review available sitemap information such as:

  • Submission status
  • Last-read date
  • Discovered URL counts
  • Retrieval failures
  • Parsing or processing errors
  • Differences between page-type groups

A submitted and valid sitemap URL can still remain unindexed.

For discovered-but-not-indexed URLs, investigate:

  • Can the page be crawled reliably?
  • Does the canonical point elsewhere?
  • Is Google selecting a different canonical?
  • Is the content duplicated across products, variants, or markets?
  • Does the page receive meaningful internal links?
  • Is it buried behind scripts or inaccessible pagination?
  • Is it useful and distinct enough to merit indexing?
  • Does it contradict robots or redirect rules?
  • Does the page remain stable, or does it frequently disappear?

No single correction guarantees inclusion. The objective is to remove contradictions and improve the page’s accessibility and distinct value.

7. Review lastmod accuracy

Compare sitemap dates with actual content changes. Watch for these patterns:

  • Every URL receives today’s date.
  • Dates change whenever the XML regenerates.
  • Deleted products continue showing recent dates.
  • Significant product changes never update the date.
  • Child sitemap dates change even when their contents do not.

The right update logic depends on the catalog, but it must be explainable and consistently applied.

8. Use partitions diagnostically

Separate files make mismatches easier to isolate:

  • Product files: inventory lifecycle, variants, redirects, and stale products
  • Category files: canonical and faceted-navigation problems
  • Brand files: thin or nonindexable brand destinations
  • Editorial files: publication and content-management issues
  • Language files: translation, hreflang, and localized canonical conflicts
  • Country files: domain, redirect, and regional-indexing inconsistencies

If all page types are mixed together, an aggregate discovered count may hide the fact that one group has failed completely.

Sitemap health metrics

Measure sitemap health using operational indicators:

  • Successful file retrieval
  • Valid XML
  • Expected and reasonably stable discovered counts
  • Percentage of listed URLs returning HTTP 200
  • Percentage of entries aligned with their canonical
  • Absence of blocked or noindex URLs
  • Absence of redirects and error responses
  • Accurate lastmod values
  • Correct child-file counts
  • Internal-link reachability of listed pages

Do not use rankings, traffic, revenue, or sales changes as direct proof that the sitemap worked. Those outcomes depend on many factors beyond URL discovery.

Pre-launch and migration checklist

Before launching or migrating a store:

  • Generate sitemaps containing the new canonical URLs.
  • Remove staging, preview, and development URLs.
  • Remove old-domain and legacy URLs from the new files.
  • Validate old-to-new redirects.
  • Confirm that active destination pages return 200.
  • Confirm crawler access in robots.txt.
  • Remove accidental noindex directives.
  • Test canonical tags on every major page template.
  • Test hreflang groups across languages and countries.
  • Update the sitemap reference in robots.txt.
  • Verify the correct production-domain property.
  • Submit the new sitemap or index.
  • Monitor retrieval, redirects, discovered counts, and canonical mismatches after launch.

Repeat the audit after large catalog imports, URL-rule changes, international expansion, platform migrations, redesigns, or major product-retirement projects.

Ecommerce sitemap FAQ

Does every ecommerce website need an XML sitemap?

No. A small store whose important pages are comprehensively linked may be discoverable without one. Nevertheless, an XML sitemap is usually a sensible operational tool for ecommerce because catalogs change frequently and often contain deep product, category, media, or localized URLs.

Its value generally increases as the catalog grows, internal linking becomes more complicated, or the store adds markets and media. It should still be treated as a discovery aid rather than a technical requirement that guarantees indexing.

Where can I find the sitemap for a Shopify store?

For an applicable Shopify store domain, check:

https://your-domain.example/sitemap.xml

Shopify generates the file automatically at the domain root. The main file links to child sitemaps for products, collections, blogs, and webpages. If the store uses international domains, check each applicable domain separately.

Should out-of-stock products remain in an ecommerce sitemap?

They can. Stock status alone does not determine eligibility.

Keep an out-of-stock product listed when it remains canonical, indexable, useful, returns HTTP 200, and is expected to return. Remove it when the page no longer deserves indexing, redirects to a genuine replacement, or no longer returns a valid indexable response.

Does submitting a sitemap improve rankings or guarantee indexing?

No. Submission tells a search engine where the sitemap is and helps it discover the URLs listed there. It does not guarantee crawling, indexing, rankings, traffic, or sales.

The page still needs to be accessible, canonical, internally linked, technically consistent, and useful enough to be selected for indexing. Sitemap submission cannot override those considerations.

Why are sitemap URLs discovered but not indexed?

“Discovered” means the search engine knows the URL exists. It does not mean the page has been crawled or selected for the index.

Investigate whether the URL is accessible, internally linked, canonical, technically stable, and distinct from other product or category pages. Check for duplication, conflicting canonicals, robots directives, weak content, soft-error behavior, rendering problems, and inaccessible product routes. Resubmitting an unchanged sitemap does not resolve those underlying issues.

The operating principle is straightforward: generate the sitemap automatically, include only preferred indexable URLs, divide large catalogs into diagnostic groups, keep metadata honest, submit and monitor the files, and ensure every listed URL agrees with the store’s canonicals, status codes, robots rules, and internal links. A clean sitemap makes discovery easier to manage, but the quality and accessibility of the underlying pages still determine whether they are suitable for indexing.