Skip to content

Sitemap URLs

Paste a site URL. We’ll find its sitemap and list every URL without duplicates.

  • Respects robots.txt
  • Finds all sitemap files
  • Fast and free

Try an example: microlink.io·vercel.com·stripe.com

Why list sitemap URLs?

A sitemap is the site’s own map of pages it wants found, including orphan pages no internal link points to. Listing those URLs is faster than crawling, and it is the same starting point search engines already use when they spend crawl budget.
You get every URL the site declared, not a sample guessed from internal links.

What can I use this sitemap tool for?

Paste a site URL to get every page the site declared in its sitemap, without crawling. That list is useful for:
  • SEO audits: see the URLs the site wants search engines to find
  • Content inventories: export a checklist of published pages
  • Coverage checks: compare the sitemap against a crawl or CMS export
  • Migrations: snapshot the public URL list before you move a site
  • Competitor research: see the public URL surface another site declared
  • Seed a crawl: start from the sitemap instead of guessing links
  • AI agents: give a pipeline the full URL inventory before it fetches pages
  • Data pipelines: start an export from the public map, not a homepage walk
Copy the list, or download it as a text file, when you need it in a spreadsheet or script.

How do I list every URL from a website?

Paste a site URL and submit. The helper reads the origin from that URL, fetches robots.txt, collects every sitemap, and walks nested indexes. Copy the list, or download it as a text file, when you need it in a spreadsheet or script.
URLs the site never declared will not appear. We don’t guess sitemap.xml if robots.txt has no sitemap, and we don’t filter PDFs or images out of the list. If they’re in the sitemap, they show up.

How does it find the sitemap?

The helper reads the origin from the URL you pasted, fetches robots.txt, and collects every sitemap. Nested indexes are expanded from there.
We don’t try sitemap.xml, sitemap_index.xml, or other common paths. If there is no sitemap, the list is empty.

What’s the difference between this and crawling a site?

This tool reads the URLs the site already declared in its sitemap. A crawler walks links from a homepage and can find pages the sitemap never listed. It will miss orphan pages that no link points to.
Use this when the site publishes a sitemap in robots.txt. Pages not in the sitemap will not appear. If you need the rendered page, that’s screenshot, HTML, or markdown, not this list.

Does it follow nested sitemap indexes?

Yes. Large sites (shops, publishers, docs) usually publish a sitemap index, a sitemap of sitemaps, that points at many child files. xml-urls walks those indexes and returns the page URLs.
Discovery still starts at robots.txt. If there is no sitemap, the list is empty. We don’t guess sitemap.xml.

What does the list include?

The loc URLs from the XML sitemaps robots.txt pointed at, the same locations search engines read. lastmod, changefreq, and priority are not in the list.
HTML sitemaps (the human-readable page in a footer) are not read. PDFs or images show up if they’re listed as loc. We don’t filter them out.

Why not parse the sitemap myself?

You’d still fetch robots.txt, follow nested indexes, and host a parser. This page is that script, already running as a Microlink Function: require() the packages, send the function, get the URL list back.
Open the card in the editor and change it. That’s the same helper, not a private endpoint.

Is this sitemap tool free?

Yes. This page needs no login and no credit card. The same helper from your code runs as a Microlink Function on the free plan: 25 requests per day.
Need more? See pricing.

Other questions?

We’re always available at [email protected].