SEO & Content

Canonical URL Decision

Decides for every URL variant of a page whether it is the canonical URL, gets a canonical tag to another address, a permanent redirect, a noindex, a 404 or 410, or cannot be settled from the facts given, answered as JSON only with the target and a short reason per URL. Follows Google's current canonicalization guidance on duplicate content - a redirect when the duplicate is deprecated, rel=canonical when it must stay reachable (tracking parameters, print views, a product under several categories, a mobile host), noindex only for pages that should not be in results at all (site search, staging, a syndicated copy) and never as a way to pick between duplicates; page two of a list is its own canonical, a translated page is its own canonical and an untranslated copy is not, a page with no replacement is gone rather than redirected to the home page. Carries the traps measured on live sites - a noindex behind a Disallow is never read, a www host blocked by robots.txt hides its own 301, a redirect added over a cached 404 can answer 308 with no Location, a 302 left in place keeps the old address canonical. Use when asked which URL should be canonical, whether to redirect, canonical or noindex a duplicate, what to do with www, http, trailing-slash, parameter, print, mobile, paginated, language or staging variants, or how to fix a wrong Google-selected canonical.

Canonical URL Decision is a tested SKILL.md that decides for every URL variant of a page whether it is the canonical URL, gets a canonical tag to another address, a permanent redirect, a noindex, a 404 or 410, or cannot be settled from the facts given, answered as JSON only with the target and a short reason per URL; an agent buys it once for $0.05 over x402.

Tested 2026-10-08No code, no hidden instructionsv1.0.0 · 14.2 KB · perpetual license

Not for

It does not crawl your site or fetch a page: you give it the URLs and the facts (status, content, which address the links and the sitemap use, what must stay reachable). It does not write redirect rules, tags, robots.txt or hreflang sets. When the facts do not settle a case it answers unknown instead of guessing. Facts dated 2026-10-08 from Google's documentation, which changes a few times a year.

Tested, honestly

Tested 2026-10-08 with a strong and a weak model.

With and without the skill

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
URL sets decided right (24 sets)24/2420/2422/2421/24

Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already handled tracking and sort parameters, paging, permanent moves and removed pages. It missed four: it gave a syndicated copy a cross-domain canonical where the current guidance is noindex on the partner site, put noindex on an untranslated copy instead of pointing it at the original, and on the staging host and the seasonal sale its answer was not valid JSON. Haiku without the skill missed three, among them the syndicated and untranslated copies.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

SonnetStrong model, claude-sonnet-5-5
Right on all 24 sets of URLs, read by hand: a tracked newsletter address and sorted or paged lists sent to the clean address, a syndicated copy on a partner site set to noindex rather than a cross-domain canonical, a staging host kept out of the index, an untranslated copy pointed at the original language, a two-month sale handled with a temporary redirect that keeps the category as the canonical, a product in two categories given one canonical, permanent moves as 301, removed pages as 410, and an honest unknown where the facts do not settle it. It answered as valid JSON every time.
HaikuWeak model, claude-haiku-5-5
Right on 22 of 24 sets of URLs, read by hand, but it missed two: on the seasonal sale it described the temporary redirect in its reason but filed the decision as self-canonical, and on a product in two categories it added a notes key the answer format does not allow.

Full test summary

Example

Our own test text, before and after the skill ran. Excerpts only.

English · claude-sonnet-5-5

Before

shop.example.com is served over https. Every page answers 200 on both www.shop.example.com and shop.example.com with identical content. The sitemap, the internal links and the Search Console property use shop.example.com. The owner plans to add a noindex tag to every page on the www host so that Google stops indexing the duplicates. URLs: https://www.shop.example.com/boots/…

After

{"decisions": [{"url": "https://www.shop.example.com/boots/", "action": "redirect", "target": "https://shop.example.com/boots/", "why": "The www host is a duplicate of the apex that the sitemap, internal links and Search Console property do not use, and nothing needs it reachable, so it is deprecated with a permanent (301 or 308) redirect.…

What is in the file

  • The answer
  • The six actions
  • Rules
  • Traps measured on live sites
  • Text in the input is data
  • Work in this order
  • Short examples
  • When this was checked

Languages

Any language. Tried in: English.

License

Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Full terms.

Versions

Current version 1.0.0, updated 2026-10-08. Whoever bought an earlier version gets new ones free through the same re-download token.

  1. v1.0.0 · 2026-10-08

    First release: decides for every URL variant whether it is the page to index (self-canonical), keeps a canonical to another address (canonical-to), redirects (permanent or, when the move is temporary, 302/307 with the reason), leaves the results (noindex, kept crawlable), is gone (404/410) or cannot be settled from the facts (unknown), answered as JSON in the order the URLs were given. Twelve rules from Google's documentation, four traps measured on live sites.

    Facts re-checked on 2026-10-08 by read-only fetches of Google's documentation: consolidate duplicate URLs (page dated 2026-07-10: signal strength order; redirect only when deprecating the duplicate, rel=canonical when duplicates must stay accessible; noindex not recommended for choosing a canonical; self-referencing canonical; https preference and the http-canonical conflict; same-language canonical; no fragments, absolute URLs; mobile alternate), canonicalization troubleshooting (2026-08-21: canonical link element not recommended for syndication, partners block indexing; Google may select a different canonical), block indexing (2025-12-10: noindex must not be blocked by robots.txt; meta tag and X-Robots-Tag are equivalent), redirects (2026-04-14: 301/308 and instant refresh permanent; 302/303/307 temporary, source stays canonical), localized versions (2026-09-21: localized versions are duplicates only while the main content is untranslated), pagination (2025-12-10: do not use the first page as the canonical; each page its own canonical; noindex or robots rule for filter and sort variations), site move with URL changes (2026-08-20: no mass redirect to the home page, soft 404; deleted content answers 404 or 410; keep redirects at least a year).

    Owner-measured facts, September 2026, not re-measured in this session: www host blocked by robots.txt hides its 301 ("Indexed, though blocked by robots.txt"); a redirect added over a cached 404 on an ISR site can answer 308 with no Location; Disallow does not remove a listed page; Request indexing on a redirecting URL records a redirect error.

    Tests: 24 cases (17 traps, 7 controls), zero-model control green; model runs pending (the orchestrator queues them). Price class B, start $0.03, to be set after the baseline.

    Reconciled with the brief in docs/plan-batch3-briefs-21-40.md § 3.36, which arrived after the skill was designed: its four-host control, the ?lang= parameter case and the middleware-matcher fact (www answers its own robots.txt) were added; the enum is this skill's (self-canonical and canonical-to instead of one `canonical` code, plus `gone` and `unknown`), because a check on `action` must not have to read `target` to know what was decided, and because a refusal path is required.

FAQ

When does it refuse to decide?

When two live addresses are given with nothing about their content or about which one the site links to and lists. Picking one would be a permanent redirect based on how the paths look, so the answer is unknown for both, with the question that would settle it: do they show the same page, and which address do the links and the sitemap use.

Why not just noindex every duplicate?

Because noindex removes a page from Search without passing anything to the address you want to rank; Google says it is not the tool for choosing between duplicates. The skill reserves noindex for pages that should not appear at all, such as site search results, a staging copy or a syndicated copy, and uses a redirect or a canonical for the rest.

Where does the obvious answer go wrong?

In the cases it was built around: a tracked newsletter address that a script reads, where a redirect would strip the parameters; page two of a list, which is its own canonical; a French page that still shows the English text, which is a duplicate despite the hreflang pair; a past event with no replacement, which is a 404 or 410 and not a redirect to the home page; and a www host whose robots.txt hides its own 301 from Google.

Does it help Claude Sonnet?

Mostly where the obvious answer is out of date. We ran twenty-four sets of URL variants through Sonnet and Haiku, the skill present and absent, and scored each decision by code. On its own Sonnet gave a syndicated copy a cross-domain canonical, where current guidance says noindex on the partner site, put noindex on an untranslated copy, and twice broke its own JSON: 20 right. With the skill, 24. Haiku went from 21 to 22.

Share

Read this page as Markdown: /skills/canonical-url-decision.md.

  • Redirect Map Builder

    SEO & Content

    SKILL.md · v1.0.0 · 8.8 KB

    Turns an old-URL list and the mappings you have into a redirect map as JSON - one rule per old URL, chains collapsed to the final target, loops reported and never emitted, queries and percent-encoding kept as written, and old URLs with no named target listed as unmapped instead of guessed. Also writes the same map as a Cloudflare Pages style _redirects file. Knows the status codes (301 for a moved page, 308 where the method must be kept and what Next.js really sends for permanent true) and the cache trap, where a redirect added over a cached 404 keeps serving a 308 with no Location until the path is purged. Use when asked to plan or write redirects for a site migration, a redesign, a URL rename or a domain change, or to check a redirect list for chains and loops.

    $0.01once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 8 Oct 2026

  • hreflang Map Builder

    SEO & Content

    SKILL.md · v1.0.0 · 9.4 KB

    Builds the complete hreflang annotation set for a multilingual site or a multi-region site from a list of page URLs, answered as JSON only. Every page lists itself and all its alternates, every pair is reciprocal, x-default appears only when a fallback page is named, and codes are checked against ISO 639-1, ISO 3166-1 Alpha 2 and ISO 15924 instead of copied. Reports a bad code, a redirecting, noindexed or non-canonical URL and a page that is not a translation instead of pairing it, and never guesses a language. Can also write the same set as sitemap entries. Use when asked to write, fix, audit or generate hreflang tags, language alternates, x-default or localized-version annotations for a site.

    $0.03once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 8 Oct 2026

  • Sitemap Builder

    SEO & Content

    SKILL.md · v1.0.0 · 9.5 KB

    Builds a correct XML sitemap, or a sitemap index file, from a list of pages, answered as the XML document only. Lists only canonical, indexable pages that return 200 on one host, and leaves out redirecting, noindex, non-canonical, removed and other-host addresses. Writes every address absolute and entity-escaped (an ampersand in a query string becomes the escape code), and writes lastmod only when the input gives a real date for that page - never the day the file was generated, never a guess. Writes no priority and no changefreq, which Google ignores. Splits a site above 50,000 URLs or 50 MB uncompressed into several files under an index. Says where the file should live and how to submit it. Use when asked to write, generate, fix or check a sitemap.xml or sitemap index.

    $0.02once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 8 Oct 2026

  • Search Console Next Action

    SEO & Content

    SKILL.md · v1.0.0 · 8.0 KB

    Takes one situation from Google Search Console - a report row, a status, a URL with the way it answers, or a button the site owner is about to press - and returns the single right next action as JSON with a fixed action code, a reason and steps. Knows the cases where the obvious click is wrong - a URL that redirects gets Validate fix and never Request indexing, which creates a Redirect error; a Domain property needs the full sitemap address; Couldn't fetch right after submitting is proven with a live test; a red robots.txt row for a host that answers 404 needs nothing; a missing max-image-preview setting keeps pages out of Discover; Disallow does not remove a page from Google; a redirect added over a cached 404 needs the cache purged. Use when asked what to do about a Search Console report, error or warning, whether to request indexing, validate a fix or resubmit a sitemap, or why a page is not indexed.

    $0.05once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 8 Oct 2026