Website Page Counter
How many pages does a website have? Enter a domain and get the declared total from its sitemap, broken down by section, in seconds.
Short answer
A website page counter reports how many pages a site has by reading its sitemap: the counter finds the sitemap from the domain, walks the index through up to 50 child files, and returns the declared total with a per-section breakdown. When no sitemap exists it falls back to a live crawl of up to 100 pages and labels the result accordingly. The number it returns is the site's declared count, which is related to, but not the same as, the number of pages Google has indexed.
What the counter counts
"How many pages does this site have?" sounds like one question, and the counter answers the most useful version of it: the number of URLs the site declares to search engines. Here is exactly what happens on a run:
- 01
Find the sitemap
From a bare domain, robots.txt is checked first, then nine common sitemap paths. A pasted sitemap URL skips the discovery.
- 02
Walk the index
Sitemap indexes are followed through their child files, up to 50 of them, including gzipped ones.
- 03
Count and de-duplicate
Every declared URL counts once. Image and video sitemap entries that repeat a page URL per asset are de-duplicated.
- 04
Break the total down
By first path segment, so the count arrives with a shape: how many pages live under /blog/, /products/, and so on. On very large sites the breakdown is computed from the first 5,000 URLs and labelled as sampled.
- 05
Fall back to a crawl when there is no sitemap
Up to 100 pages by following internal links, clearly labelled, because a crawl count is a floor rather than a total.
The three page counts, and why they differ
A site does not have one page count. It has three, and the gaps between them are where the findings live:
| Count | Where it comes from | What it means |
|---|---|---|
| Declared pages | The sitemap (this tool) | What the site tells search engines it has and wants indexed |
| Indexed pages | Search Console's Pages report | What Google actually holds. The site: search estimate is the rough public version |
| Linked pages | A crawl following internal links | What a visitor or crawler can actually reach from the homepage |
On a healthy site the three run close together. The gaps are diagnoses:
- Declared far above indexed: Google is declining pages you are offering. The Pages report's exclusion reasons say why, and thin or duplicative content is the usual answer.
- Declared far above linked: the sitemap lists pages your own site never links to, which are orphan pages. The sitemap comparison tool finds them by diffing the two lists.
- Linked far above declared: the site generates URLs the sitemap ignores, usually parameters, pagination, or forgotten sections quietly spending crawl budget.
What your page count is telling you
The total is a headline; the section breakdown is the story. Three reads worth doing on your own result:
- Where the weight sits. If one section holds most of your pages, that is where your crawl attention goes, whether or not it is where your value sits. A 4,000-page tag archive dwarfing a 60-page product catalogue is a decision somebody should have made on purpose.
- Growth nobody chose. Sections grow by generation: tags, dates, filters, feeds. Any section whose count surprises you is worth opening, because surprise pages are rarely good pages.
- The earning share. Put the count next to how many pages actually earn clicks in Search Console. On most mature sites a small fraction of pages does almost all the earning, and knowing your ratio is the starting evidence for a content pruning pass.
Run it on a competitor next
A competitor's page count with its section breakdown is their content strategy in one table: how many product pages, how many posts, which sections they keep feeding. It is public, it takes ten seconds, and it is more honest than their marketing.
Why the count matters for SEO
Page count is not a score, and more is not better. It matters as a denominator:
- Crawl attention is finite. Every declared page asks Google to spend crawl time on it. Thousands of pages that earn nothing are thousands of requests not spent on the pages that do.
- Quality is judged in aggregate. A site where most crawled pages are thin or generated reads as that kind of site. The count tells you how much surface you are asking Google to judge.
- Maintenance scales with pages. Every page is a future stale fact, broken link, and decayed ranking. The honest question is not "how many pages can we have?" but "how many can we keep good?"
From counting to knowing
The counter tells you how many pages exist. The follow-up questions, how many are indexed, how many earn traffic, and which ones deserve to stay, need your Search Console and analytics data joined to a crawl, which is exactly the join CrawlRaven runs: every page's search trend, engagement and technical findings in one ranked plan. The count is the denominator; the join fills in the numerator.
Frequently asked questions
How does the website page counter work?
It finds the site's sitemap (checking robots.txt and nine common paths), walks any sitemap index down through its child files, and counts every declared URL, broken down by section and by sitemap file. If no sitemap exists, it falls back to a live crawl of up to 100 pages and says so, because a crawl-based count answers a different question than a sitemap-based one.
How many pages does my website have?
Enter your domain above and the counter reads your sitemap's declared total in seconds. Note what the number means: it is the count of pages your site declares to search engines, which is usually close to, but not the same as, the number of pages that exist or the number Google has indexed. The three counts and their differences are explained below the tool.
Why is the page count different from a site: search on Google?
A site: search estimates how many of your pages Google has indexed, and Google itself describes that number as a rough estimate that fluctuates. The sitemap count is what you declare; the site: estimate is what Google admits to holding. A large gap between them, in either direction, is a finding worth investigating in Search Console's Pages report, which is the accurate version of the indexed count.
Why is my page count different from my CMS's page count?
The CMS counts content records; the sitemap counts URLs the site declares for indexing. Tag archives, pagination, media attachment pages and generated URLs can push the sitemap far above the CMS count, while noindexed or excluded content pulls it below. When the two disagree wildly, the sitemap generator's settings are usually the explanation.
What is the maximum number of pages it can count?
The counter reads up to 50 sitemap files and reports the full declared total it finds within its time budget. The per-section breakdown is computed from the first 5,000 URLs, so on very large sites the total is complete but the breakdown is labelled as sampled.
What if my site has no sitemap?
The counter falls back to a live crawl of up to 100 pages by following internal links from your homepage, and labels the result as crawl-based. That is a floor, not a total, for anything but small sites. It is also a prompt: generate a sitemap, because a site without one is making Google guess at exactly the question you just asked.
Does the page count include images and PDFs?
It counts what the sitemap declares. Most sitemaps declare pages only, but image and video sitemap entries reference page URLs per asset, which the underlying reader de-duplicates. PDFs appear only when a site deliberately lists them, so a sitemap-based count is best read as a page count, not a file count.
Can I count the pages on a competitor's website?
Yes. Sitemaps are public files, and a competitor's page count with its per-section breakdown is a fast read on the size and shape of their content investment: how many product pages, how many blog posts, which sections they keep growing. Run it on two competitors and the strategy differences are visible in the section table.
Does more pages mean better SEO?
No, and past a point the opposite. Pages that earn nothing still spend crawl budget and dilute the site's quality signals, which is why mature sites prune. What matters is the share of your pages that deserve to exist: a 200-page site where every page earns beats a 5,000-page site where 200 do.
How do I find out how many of my pages Google has indexed?
Search Console's Pages report is the accurate answer: it shows indexed and not-indexed counts with the reason for every exclusion, for your verified property. The site: search estimate is a rough public approximation. Comparing your sitemap count against the indexed count, and reading the exclusion reasons, is one of the fastest health checks in SEO.
Other sitemap tools
XML Sitemap Generator
Crawl any website and generate a valid XML sitemap you can download and submit. No signup, no software to install.
Use toolXML Sitemap Validator
Test any XML sitemap against the sitemaps.org protocol. Catches namespace errors, bad lastmod dates, size limits, and cross-domain URLs.
Use toolSitemap Checker
Check a sitemap end to end: the file against the sitemaps.org protocol, then a live check of the URLs inside it. One pass, one health score.
Use toolSitemap Crawler
Crawl every URL your sitemap declares and find the dead links, redirects, noindex pages, and canonical conflicts hiding inside it.
Use toolSitemap URL Extractor
Pull every URL out of a sitemap or sitemap index and download the list as CSV, TXT, or JSON. Handles nested indexes and gzipped files.
Use toolSitemap Index Generator
Paste your sitemap URLs and get a valid sitemap index file back, with optional lastmod dates. For sites past the 50,000 URL limit that split their sitemaps.
Use toolSitemap Comparison
Compare two sitemaps and see exactly which URLs were added, removed, and kept. Catch the pages a migration or a regenerated sitemap silently dropped.
Use toolSitemap Finder
Find where any website hides its sitemap. Checks robots.txt directives plus nine common paths, including WordPress and gzipped variants.
Use toolVisual Sitemap Generator
Turn any site into a visual tree of its structure. See how deep your pages sit and which sections carry the most URLs.
Use toolRead up on sitemaps
A tool tells you what is wrong. These explain what to do about it.
How to Add Google Search Console to Shopify (So It Survives Your Next Theme Update)
Read guideHow to Perform a Technical SEO Audit in 2026 (Step-by-Step Guide)
Read guideTechnical SEO Audit Checklist 2026 (Free Template, 50+ Checks)
Read guideTerms this tool checks
- XML Sitemap
An XML sitemap is a file (typically at /sitemap.xml) that lists all the important URLs on your website along with metadata like last modification date, change frequency, and priority.
- Crawl Budget
Crawl budget is the number of pages a search engine will crawl on your site within a given timeframe, determined by crawl rate limit (how fast Googlebot can crawl without overloading your server) and crawl demand (how much Google wants to crawl based on popularity and freshness).
- Indexation
Indexation is the process by which search engines add web pages to their searchable database (index).
Counting pages is the easy half
CrawlRaven joins the same pages to your Search Console impressions, GA4 engagement, and a 200-point crawl, so you know how many pages earn, not just how many exist.