# Sitemap checker

> Free sitemap checker. Find the XML sitemap a crawler finds, count its URLs, see whether robots.txt declares it and what the lastmod dates say.

Type a domain and this checker looks where a crawler looks: the Sitemap lines of robots.txt, then /sitemap.xml, then /sitemap_index.xml. It reads the tree under whatever answers. The report counts the URLs, reads the lastmod dates and names what to fix.

Page: https://seodraft.app/tools/sitemap-checker

## Run it from your agent

The check on this page is also a JSON endpoint and a tool on a public MCP server. Neither needs an account or a key.

### JSON endpoint

```http
POST https://seodraft.app/api/tools/sitemap-checker
content-type: application/json

{
  "url": "https://seodraft.app"
}
```

Limits: 4 calls per minute per IP, 20 per hour per checked site. Over a limit the answer is 429 with Retry-After.

### Free MCP server

Add this URL to Claude, ChatGPT, Cursor or any MCP client. It exposes every free tool, with no sign-in.

- URL: https://seodraft.app/tools/mcp
- Tool: `check_sitemap`

Finds the sitemap a crawler would find for a site and reads the tree under it. url takes a site URL or a sitemap URL. verdict is ok, issues, not-found or unreachable. sitemapUrl is the file the crawl started from, foundVia says how it was reached (given, robots, conventional), and inRobots is true when a Sitemap line in robots.txt names that file. sitemapsRead, urlCount, sampleUrls (20 at most) and contentPrefix describe what was read. lastmod counts URLs with a readable date, missing, invalid and dated after now, plus newest and oldest as ISO. skipped counts what was left out as offHost, limits or errors. findings is ordered by severity (critical, warning, info), each with a code (no-sitemap, unreachable, no-urls, not-in-robots, off-host-urls, lastmod-missing, lastmod-future, lastmod-invalid, sitemap-errors, partial-read), a message saying what to fix and the count it is about. Reads at most 10 sitemap files, 5,000 URLs and 15 seconds, and answers partial-read when it stopped at one of those.

## What a good XML sitemap does

A sitemap is the site's own list of its URLs: one <loc> per page, with an optional <lastmod> saying when that page last changed. A crawler reads it as a claim about what exists on the site.

Four properties make that claim usable. The file answers at a URL robots.txt names, it parses as XML, every URL sits on the same domain as the file, and every lastmod is a real date in the past.

The sitemaps.org protocol caps one file at 50,000 URLs and 50 MB uncompressed. A bigger site splits the list and publishes a sitemap index, which is a sitemap of sitemaps, and gzip is allowed for both.

## How this sitemap finder works

It follows the order a crawler follows. robots.txt comes first and every Sitemap line in it is tried in file order. Then /sitemap.xml, then /sitemap_index.xml.

Paste a sitemap URL and the check starts there. A path ending in .xml or .xml.gz, or a last segment containing the word sitemap, is read as the file itself and the rest of the URL as the site.

Whatever answers gets parsed. When it is an index, the files under it are opened too. Everything stays on the host you typed, so a URL pointing at another domain is counted and left unread.

## What each finding means

Findings arrive critical first, then warnings, then information, and each one carries the count it is about.

no-sitemap says nothing answered at robots.txt, /sitemap.xml or /sitemap_index.xml. not-in-robots says the file exists while robots.txt stays silent about it, which leaves a crawler guessing the path. no-urls says the file parsed and lists nothing.

off-host-urls counts URLs on a domain other than the one serving the file. lastmod-invalid counts values that fail to parse as a date. lastmod-future counts dates after today, which come from a build stamp. lastmod-missing counts URLs with no date at all.

sitemap-errors counts files inside an index that answered with something other than readable XML. partial-read says the crawl stopped at a limit, so read the counts as a floor.

## Where this check stops

It reads sitemap files and stops there. No page is downloaded, so the report stays quiet about status codes, canonicals and indexing.

One call reads at most 10 sitemap files, 5,000 URLs and 15 seconds of wall time, and it samples 20 URLs back. A bigger site gets the partial-read finding and counts that cover the part it read.

Whether a URL in the sitemap made it into the index is a question for Search Console's page indexing report.

## Questions

### How do I find my sitemap URL?

Type your domain in the box above and the check goes looking for it. It reads robots.txt for a Sitemap line, then tries /sitemap.xml and /sitemap_index.xml. WordPress publishes its own at /wp-sitemap.xml and plugins move it elsewhere, which is what the robots.txt line is for.

### Does a sitemap get my pages indexed?

A sitemap tells a search engine which URLs exist and when they changed. Indexing is decided per URL after a crawl, and Search Console's page indexing report is where the outcome shows up.

### What should lastmod say?

The date the content of that page changed. Google's documentation says it uses lastmod when the value is consistently accurate, so a date restamped on every URL at every deploy teaches a crawler to ignore the field.

### Can I check a sitemap index?

Yes, and the files listed inside it are opened too, up to 10 of them. Paste the index URL and the report covers the URLs across every file it read.

### Why does the report say off-host-urls?

Your sitemap lists URLs on a domain other than the one serving the file. A sitemap covers the host it sits on, and URLs on another domain belong in that domain's own sitemap, declared in that domain's robots.txt.

### Does it read a gzipped sitemap?

Yes, a .xml.gz file is decompressed and parsed like any other. That is how most large sitemap indexes are served.

## The full process runs in seodraft

seodraft keeps the same crawl as a workspace inventory: import_sitemap stores the pages the site already published, and add_topics withholds a question one of them already answers, naming the URL it clashes with.

MCP: https://seodraft.app/mcp
