DK logodavorkarafiloski
AboutCase StudiesSpeakingGlossary Work With Me
Glossary/XML Sitemap
Core / Technical SEO

XML Sitemap

FoundationsPractitionerSenior lens
Quick definition

A machine-readable list of a site’s important URLs, submitted to help search engines discover and prioritise them.

01
Foundations
New to SEO? Start here.

An XML sitemap is a structured file listing the URLs you want search engines to know about, optionally with a lastmod date signalling when each page changed. You submit it in Search Console. It does not force anything to rank — it is a discovery and prioritisation aid that helps crawlers find pages efficiently.

02
Practitioner
Doing the work day to day.

A sitemap should contain only indexable, canonical, 200-status URLs — the exact set you want ranked. Do not include redirects, noindexed pages, or non-canonical duplicates; that erodes Google’s trust in the file. On large sites, split into multiple sitemaps under a sitemap index (by content type or section) so the Search Console coverage report becomes a diagnostic: indexed vs submitted per segment tells you where quality problems live.

03
Senior lens
Strategy, trade-offs, judgement.

Treat the sitemap as a reconciliation tool between what you think you publish and what Google actually indexes. Accurate lastmod dates genuinely help recrawl priority on large, frequently-updated sites — but only if you keep them honest; fake freshness gets ignored. For news, video, and image assets, the specialised sitemap formats unlock discovery you cannot get any other way.

DKDavor’s take

A sitemap full of redirects and noindex URLs is worse than no sitemap — you are actively teaching Google that your own list of "important pages" cannot be trusted. Keep it pristine.

Common mistakes
Dumping every URL in, including redirects, noindex, and non-canonical pages.
Letting the sitemap drift out of sync with the site after a migration.
Faking lastmod dates to chase recrawls — Google learns to ignore them.
In practice
A blog splits into sitemap-posts.xml and sitemap-pages.xml under a sitemap index, so Search Console reports indexation coverage per content type.
Related terms
Indexation
The process by which a search engine stores and makes a crawled page eligible to appear in results.
Robots.txt
A root-level file that tells crawlers which parts of a site they may or may not request.
Canonical Tag
A signal (rel="canonical") that tells search engines which URL is the master version of duplicate or near-duplicate pages.
Keep learning · Core / Technical SEO
Crawl BudgetCore Web VitalsStructured Data (Schema Markup)Internal Linking
← Previous
Robots.txt
Next →
Structured Data (Schema Markup)
← Back to the full glossary

Want this applied to your pipeline?

I turn concepts like these into quarterly roadmaps and measurable organic revenue for SaaS teams.

Work with me →
davorkarafiloski

Proven SEO systems for SaaS teams that refuse to fall behind in AI-era search.

© 2026 Davor Karafiloski. All rights reserved.
Skopje, North Macedonia · SEO Director at SmartClick