XML Sitemap
An XML sitemap is a file that lists all the important URLs on a website, helping search engines discover and prioritise content for crawling and indexing.
Reviewed by Alexander Yarovenko · Updated: 2026-09-18
What is an XML Sitemap?
An XML sitemap is a structured file (in XML format) that lists the important URLs on your website, optionally including metadata such as when each page was last modified, how frequently it changes, and its relative priority. It acts as a roadmap for search engine crawlers, helping them discover content they might otherwise miss.
XML Sitemap Structure
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://example.com/</loc>
<lastmod>2025-01-15</lastmod>
<changefreq>weekly</changefreq>
<priority>1.0</priority>
</url>
<url>
<loc>https://example.com/about/</loc>
<priority>0.8</priority>
</url>
</urlset>
Types of Sitemaps
- Standard XML sitemap — web pages
- Image sitemap — helps Google discover images
- Video sitemap — video metadata for Google Video Search
- News sitemap — required for Google News inclusion (articles from last 2 days)
- Sitemap index — points to multiple child sitemaps for very large sites
Best Practices
- Only include canonical, indexable URLs — no noindex, no redirect targets, no 404s
- Keep each sitemap file under 50,000 URLs and 50MB uncompressed
- Submit your sitemap in Google Search Console
- Reference the sitemap in
robots.txt:Sitemap: https://example.com/sitemap.xml - Update the sitemap automatically when content changes
Sitemap vs robots.txt
These two files work together but serve opposite purposes: robots.txt tells crawlers what NOT to crawl; a sitemap tells them what IS important to crawl. Both should be consistent — don't include in your sitemap URLs that are disallowed in robots.txt.
How to use this concept in SEO
An XML sitemap is a prioritized inventory, not a dump of every generated URL. Include canonical, indexable 200 pages you genuinely want in search and provide truthful lastmod values.
Audit checklist
- Confirm that the implementation matches the page purpose and the user task.
- Check the HTTP response, rendered HTML, canonical, robots rules, and internal links together.
- Use Search Console and analytics to verify the result instead of relying on one crawler signal.
- Test a representative URL on mobile and desktop after deployment.
- Document the expected outcome so regressions can be detected in the next audit.
Practical example
A site removes redirects, noindex pages, parameters, and thin translations from the sitemap; Search Console then compares submitted URLs with indexed URLs more meaningfully.
How to validate the result
Record the current crawl, indexation, performance, and traffic state before changing anything. Recheck the affected template and a representative URL after deployment, then monitor Search Console over the following crawl cycle. A technical change is complete only when the live response and Google’s observed state agree.