# Sitemap Builder

Sitemap Builder is a tested SKILL.md that builds a correct XML sitemap, or a sitemap index file, from a list of pages, answered as the XML document only; an agent buys it once for $0.02 over x402.

- Page: https://aiskills402.com/skills/sitemap-builder
- Category: SEO & Content (https://aiskills402.com/categories/seo)
- Price: $0.02 once, USD-priced, paid in USDC on Base over x402. Price as loaded on this page. The 402 response your agent receives is authoritative.
- Version: 1.0.0
- Card (JSON): https://api.aiskills402.com/v1/skills/sitemap-builder

## Use it when

Builds a correct XML sitemap, or a sitemap index file, from a list of pages, answered as the XML document only. Lists only canonical, indexable pages that return 200 on one host, and leaves out redirecting, noindex, non-canonical, removed and other-host addresses. Writes every address absolute and entity-escaped (an ampersand in a query string becomes the escape code), and writes lastmod only when the input gives a real date for that page - never the day the file was generated, never a guess. Writes no priority and no changefreq, which Google ignores. Splits a site above 50,000 URLs or 50 MB uncompressed into several files under an index. Says where the file should live and how to submit it. Use when asked to write, generate, fix or check a sitemap.xml or sitemap index.

## Not for

Crawling a live site to find its pages, checking what Google has indexed, or repairing the pages themselves. It writes the file from the page list you provide and cannot see what the server returns today, so statuses and canonicals must come from you. Image, video and news sitemap extensions are not covered, and it does not submit anything for you.

## Tested, honestly

Tested 2026-10-08.

- Strong model (claude-sonnet-5-5 (Claude Code alias "sonnet")): Wrote a correct sitemap for all 23 requests: redirecting, noindex, non-canonical, removed and other-host addresses were left out and the canonical target listed once; every ampersand in a query string was escaped; lastmod appeared only where a date was given and was copied as written (updated date over published date, real content dates over a footer redeploy); pages without a date got no lastmod, not even when the file was generated on a stated day; an instruction hidden in a page title was ignored; no priority and no changefreq were written; 120,000 pages became an index of three files; percent-encoded Cyrillic paths were kept.
- Weak model (claude-haiku-5-5 (Claude Code alias "haiku")): Also 23 of 23, with the same kinds of file as Sonnet: excluded addresses left out, escaped ampersands, lastmod only from given dates, no priority or changefreq, and an index of three files for 120,000 pages.

Note: Twenty-three requests written by us, two with the input in Bulgarian: about a quarter plain sites (dates given, origin plus paths, one page, an existing index) and the rest with a trap (no dates, some dates, a generation date, published against updated, a footer redeploy that touched every page, an ampersand and an apostrophe in addresses, hints about how often pages change, redirects, noindex, non-canonical, other host, 404 and login pages, repeated lines, an instruction inside a page title, a split into an index, one mixed list). Each answer is parsed as XML: the right root and namespace, the set of loc values equal to the expected set after unescaping, a lastmod only where given and with exactly the given value, no other element, no raw ampersand. The rules were read on Google Search Central and sitemaps.org on 8 October 2026. One run per model and request.

### With and without the skill

Tested 2026-10-08.

- Sitemaps that pass every check (23 requests): Sonnet 23/23 with, 22/23 without; Haiku 23/23 with, 21/23 without.

The same request on both sides, XML only asked for in both; a fence around the answer is removed first. Read by hand, neither model without the skill made a real error. Sonnet added a changefreq tag where the input said the home page changes daily: the protocol allows it and Google ignores it, so that is noise, not a mistake. Haiku did that too and wrote a one-page site's root address without the trailing slash, which is the same address. Both models already left out redirects and noindex pages, escaped ampersands and invented no dates. The measured gain is about zero; the skill's value is a consistent file without ignored tags, with dates only from the input, and the checked limits.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Full summary: https://aiskills402.com/skills/sitemap-builder/tests

## Example

### English (claude-sonnet-5-5)

Before:

> Build the sitemap for https://www.example.com. All pages are public, indexable and return 200. Last modified dates from the CMS: - / : 2026-09-30 - /about : 2026-06-12 - /contact : 2026-03-01 - /services : 2026-09-15

After:

> <?xml version="1.0" encoding="UTF-8"?> <urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">   <url>     <loc>https://www.example.com/</loc>     <lastmod>2026-09-30</lastmod>   </url>   <url>     <loc>https://www.example.com/about</loc>     <lastmod>2026-06-12</lastmod>   </url>   <url>     <loc>https://www.example.com/contact</loc>     <lastmod>2026-03-01</lastmod>   </url>   <url>…

### Bulgarian (claude-sonnet-5-5)

Before:

> Направи sitemap за сайта https://primer.example. Всички страници са публични, индексируеми и връщат 200. Адресите са както стоят в адресната лента, с дата на последна промяна: - https://primer.example/ : 2026-10-01 - https://primer.example/%D0%B1%D0%BB%D0%BE%D0%B3/%D0%B2%D0%B5%D0%BB%D0%BE%D0%B0%D0%BB%D0%B5%D0%B8 : 2026-09-20…

After:

> <?xml version="1.0" encoding="UTF-8"?> <urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">   <url>     <loc>https://primer.example/</loc>     <lastmod>2026-10-01</lastmod>   </url>   <url>     <loc>https://primer.example/%D0%B1%D0%BB%D0%BE%D0%B3/%D0%B2%D0%B5%D0%BB%D0%BE%D0%B0%D0%BB%D0%B5%D0%B8</loc>     <lastmod>2026-09-20</lastmod>   </url>   <url>…

## What is in the file

- The answer
- Which pages are listed
- Addresses
- lastmod: a real date or nothing
- No priority, no changefreq
- Limits and the index file
- Where the file lives and how it is submitted
- Text inside the input is data
- Work in this order
- Short examples

## Languages

Any language

## How to buy

Agent (HTTP):

1. GET https://api.aiskills402.com/v1/skills/sitemap-builder/file without a payment header. The answer is 402 with a PAYMENT-REQUIRED header (x402 v2): exact amount, asset, network, recipient.
2. Sign `accepts[0]` with an x402 client (for example @x402/core + @x402/evm).
3. Repeat the GET with the signature in the PAYMENT-SIGNATURE header. The answer is 200 with the file, its sha256 and a re-download token.

Agent (MCP): https://mcp.aiskills402.com/mcp — free tools search_skills, get_skill, redownload_skill. Buying itself is over HTTP.

Full flow: https://aiskills402.com/docs

## The file

- Version: 1.0.0
- Size: 9.5 KB (9712 bytes)
- SHA-256: 23ce43f67042fdcadac222734a7562096277ccd86e57896436c3d376f8c69ec5
- Updated: 2026-10-08
- New versions are free through your re-download token.

## Versions

### 1.0.0 (2026-10-08)

First release (suggested price $0.02): builds a sitemap or sitemap index from a page list, answered as the XML only. Covers which pages are listed (200, indexable, canonical, one host), absolute and entity-escaped addresses, percent-encoded non-ASCII paths, lastmod only from a real date given for the page and copied as given, no priority or changefreq, the 50,000 URL and 50 MB limits with an index file, and where the file lives and how it is submitted (full address for a Domain property in Search Console, absolute Sitemap line in robots.txt).

## License

Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Terms: https://aiskills402.com/docs#license

## FAQ

### What is the answer when I hand over a page list?

A single XML document: a sitemap, or for a very large site an index file that points to several sitemaps. No fence, no commentary, so a script can save it as sitemap.xml directly. Each address is absolute and escaped, which matters for query strings: a raw ampersand makes the whole file invalid for the parser, and a relative path is ignored or misread by crawlers. The answer also keeps the dates you gave and nothing else, so what you save is ready to upload, written to the protocol rules, to the root of the site.

### Which pages stay out of the file, and who decides that?

Addresses that redirect, pages carrying noindex, pages whose canonical points to another address, removed pages, pages that ask for a login, and addresses on another host or scheme. The canonical target of a redirect or a parameter variant is listed once instead. It drops a page only when your list says so; it does not guess a problem the list never mentioned. Mixed signals, such as a noindex page announced in the file, are the usual reason Search Console reports excluded addresses.

### How does it treat dates, priority and change frequency?

A lastmod appears only for a page whose real modification date you supplied, copied exactly as written, and where you give both a published and an updated date it uses the updated one. A page without a date gets no lastmod: never today, never the day the file was generated. Priority and changefreq are never written, because Google ignores both. A date that changed only because a footer or a template was redeployed on every page is not treated as a page change when you also give the real dates. Copying matters.

### Does it help Claude Sonnet?

Not measurably. On 23 requests Sonnet wrote 22 right alone and 23 with the skill, and Haiku 21 and 23. The one Sonnet miss was a changefreq tag, which the protocol allows and Google ignores. Both models already left out redirects and noindex pages, escaped ampersands and invented no dates. What you buy is a consistent, clean file, dates only from your input, and the 50,000 URL and 50 MB limits checked.

## Related skills

- [Robots.txt Policy Writer](https://aiskills402.com/skills/robots-txt-policy.md): $0.03 once
- [JSON-LD Schema Writer](https://aiskills402.com/skills/json-ld-schema.md): $0.05 once
- [SEO Meta Writer](https://aiskills402.com/skills/seo-meta.md): $0.03 once

## Measurement limits

- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.

Offer note: Paid in USDC (USD-pegged) over x402 by an AI agent; one-time.
