Skip to content

SEO

How many pages can Google index?

There is no page count that decides whether a site succeeds. What decides it is how many of your URLs are distinct, useful and reachable, which is a different question wearing the same clothes.

THE SHORT ANSWER

The number of pages Google will index is not a fixed quota, and Google does not publish one. Its own documentation says plainly not to expect every URL on a site to be indexed, because some are duplicates and some contain nothing meaningful. So an 8,000 page site is not inherently a problem, and 8,000 substantially repetitive pages is a problem at any size. The useful question is never how many pages you can publish. It is how many of them represent a genuinely distinct thing somebody searches for.

The direct answer

There is no published maximum, and no useful universal number. Google does not issue sites a page allowance, and very large sites index and rank perfectly well.

What Google does say, in its own documentation on the Page Indexing report, is this: "Don't expect every URL on your site to be indexed. Some URLs might be duplicates or might not contain meaningful information. Just be sure that the key pages on your site are indexed."

That sentence is the whole subject. Indexing is not a capacity you fill. It is a series of judgements, made per URL, about whether that URL is worth storing and showing. Publishing capacity and indexing are different problems, and conflating them is what produces the question in the first place.

Nine different things people mean by "pages"

Almost every argument about site size is two people using one word for different states. A URL passes through all of these, and it can stall at any of them for unrelated reasons. Knowing which one you are stuck at tells you what to fix.

  • URLs generated. What a build process produced. Costs nothing and means nothing on its own.
  • URLs published. Live and reachable by a browser. Still nothing to do with Google.
  • Crawlable URLs. Not blocked by robots.txt, not behind auth, not dependent on rendering Google cannot do, and reachable by a link.
  • URLs submitted in sitemaps. What you have told Google you would like considered. Google is explicit that this is a hint.
  • URLs discovered. Google knows the URL exists, from a sitemap or, more durably, from a link.
  • URLs crawled. Google has actually fetched the page.
  • Canonical URLs. Where duplicates were consolidated, the one Google selected to represent the group.
  • Indexed URLs. Stored and eligible to appear. Eligible is not the same as appearing.
  • URLs earning impressions. Actually shown to somebody. This is the only one that has any relationship to outcomes.

Where large sites actually stall, in order

Each gap between two of those states has its own cause, and they are not interchangeable.

  • Published but not crawlable. Orphan pages with no internal links, blocked paths, or content that only exists after a script runs. A page nothing links to is a page your own site declined to vouch for.
  • Crawlable but not discovered. Usually a linking problem rather than a sitemap problem. Sitemaps help Google find things; internal links tell it what matters.
  • Discovered but not crawled. Google has the URL queued and has not fetched it. On large sites this is where crawl capacity genuinely shows up.
  • Crawled but not indexed. Google fetched the page and decided against storing it. This is a quality and distinctness judgement, and it is the one people misdiagnose as a technical fault.
  • Indexed but no impressions. The page is eligible and nobody is searching for what it covers, or others answer it better. Nothing is broken. The page was not worth making.

The two Search Console statuses everybody confuses

These sound alike, mean opposite things, and are the single most common source of wrong conclusions about large sites. Google defines both in the Page Indexing report documentation.

"Discovered, currently not indexed" means, in Google's words, that the page was found but has not been crawled yet, and that typically Google wanted to crawl it but expected that to overload the site, so it rescheduled the crawl. Nobody has judged your content. It has not been read.

"Crawled, currently not indexed" means Google fetched the page and did not index it, and that it may or may not be indexed later.

The practical difference is total. The first is a capacity and scheduling signal, and it is one of the situations where Google says crawl budget is worth thinking about. The second is a verdict on the page. Rewriting content to fix the first, or upgrading a server to fix the second, is a common and expensive mistake.

What crawl budget really is, and who Google says should care

Crawl budget is the interaction of two things Google describes separately. Crawl capacity limit is how much crawling your server can take without being harmed, and crawl demand is how much Google wants to crawl your site.

Google scopes the topic explicitly. Its guide is aimed at large sites of a million or more unique pages with content changing about weekly, at medium or larger sites of ten thousand or more unique pages with very rapidly changing content, and at sites where a large share of URLs sit in "Discovered, currently not indexed".

Google states directly that those numbers are a rough estimate to help you classify your site, and that they are not exact thresholds. It also says that if your pages seem to be crawled the same day they are published, you do not need the guide at all. Treat any article that presents those figures as a hard limit with suspicion, including on the day somebody quotes this one back at you.

Sitemaps: useful, and not what most people think

A sitemap tells Google which URLs you consider canonical and would like shown. It is a discovery aid.

Two facts worth having exactly right. First, the format limits: a single sitemap file is capped at 50MB uncompressed or 50,000 URLs, which is why large sites use a sitemap index pointing at several files. Second, and more important, Google states that submitting a sitemap is merely a hint, and does not guarantee that Google will download the sitemap or use it for crawling URLs on the site.

So submission is not indexation, and a sitemap full of weak URLs does not become strong by being submitted. What a sitemap genuinely earns you is faster discovery of pages that were worth discovering.

Duplicates and canonicalisation, which quietly eat large sites

When several URLs carry substantially the same content, Google consolidates them and picks one to represent the set. It chooses "the page that, based on the factors the indexing process collected, is objectively the most complete and useful for search users."

Two consequences people find surprising. Your canonical tag is a hint, not a rule. Google says indicating a canonical preference is a hint and that it may choose a different page than you did. And duplicates are crawled less frequently, which Google describes as a way of reducing crawling load. So a large site with a duplication problem does not only waste the duplicate pages. It slows down the crawling of everything.

This is why near-duplicate detection matters more as a site grows, and why it has to be insensitive to a swapped noun. Two pages that differ only by a place name are the case a naive comparison calls different and a reader calls identical.

Why 8,000 pages is fine and 8,000 of the wrong pages is not

Nothing above sets a ceiling. A site with 8,000 URLs, each answering a genuinely different question, is a large useful site, and Google indexes large useful sites every day.

The same 8,000 URLs become a problem when they are substantially repetitive or have no justification for existing separately. Then several things happen at once: duplicates are consolidated so most of them never appear anyway, duplicate crawling is throttled so the rest of the site is discovered more slowly, and the set starts to resemble things Google names in its spam policies. Scaled content abuse is defined as many pages generated for the primary purpose of manipulating search rankings rather than helping users. Doorway abuse covers pages created to rank for specific, similar queries that lead users to intermediate pages less useful than the destination.

Neither definition mentions how the pages were produced. Both are about what the pages are for. That is the entire distinction, and it does not move with the page count.

The position worth holding

Scale is not the strategy. Distinct search intent is the strategy, and scale is the consequence when the market supports it.

Read the sequence backwards and it decides everything. Start with a page count and you will produce URLs to hit it, which means producing pages that are not distinct, which means duplicates, throttled crawling and a set that looks like the thing the policies describe. Start with the genuinely distinct questions your buyers ask and count them, and the size of the site is an output. Sometimes that output is forty pages. Sometimes it is several thousand. Both answers are correct and neither was chosen.

Which also means the honest version of "how many pages can Google index" is a question about your market, not about Google.

What to measure instead

Counting indexed URLs is a weak measure on its own, because indexed only means eligible. A site can grow its indexed count while nothing improves.

  • How many published URLs have ever earned an impression. The gap between published and earning is the honest picture of a large build, and it is the number nobody volunteers.
  • The count of distinct queries the site appears for. It moves earlier than clicks and is hard to inflate without doing real work.
  • Movement between states over time. Discovered rising while crawled stays flat is a capacity story. Crawled rising while indexed stays flat is a quality story.
  • What has been merged or removed. On a large site, pruning and consolidating is ordinary maintenance rather than an admission of failure.

Common questions

How many pages can Google index from one site?

Google does not publish a per-site limit and there is no useful universal number. Its documentation says not to expect every URL on a site to be indexed, because some are duplicates and some contain nothing meaningful, and advises making sure the key pages are indexed. Large sites are indexed routinely. What varies is how much of a given site is judged worth indexing.

What does "Discovered, currently not indexed" mean?

That Google found the URL but has not crawled it yet. Google says this typically happens when it wanted to crawl the URL but expected that to overload the site, so it rescheduled the crawl. Nothing about your content has been assessed, because the page has not been read. It is different from "Crawled, currently not indexed", which means Google did fetch the page and did not index it.

Does submitting a sitemap get pages indexed?

No. Google states that submitting a sitemap is merely a hint and does not guarantee that Google will download the sitemap or use it for crawling URLs on the site. A sitemap tells Google which URLs you consider canonical and speeds up discovery of pages worth discovering. It does not make a weak page eligible.

Do I need to worry about crawl budget?

Probably not. Google aims its crawl budget guidance at sites of around a million or more unique pages changing about weekly, at sites of around ten thousand or more pages changing very rapidly, and at sites with a large share of URLs sitting in "Discovered, currently not indexed". It also states those numbers are a rough estimate rather than exact thresholds, and that if your pages are crawled the same day they are published you do not need the guide.

Will removing pages help the rest of the site?

It can. Google says duplicates are crawled less frequently in order to reduce crawling load, so a large set of near-identical URLs slows down discovery across the site as well as failing on its own terms. Merging overlapping pages into one stronger page, and removing pages that have never earned an impression, is ordinary maintenance on a large site rather than a sign something went wrong.

Is a page being indexed the same as it ranking?

No. Indexed means stored and eligible to appear. Whether it appears depends on whether anybody searches for what it covers and whether other pages answer that better. This is why counting indexed URLs is a weak measure on its own, and why the more useful number is how many published URLs have ever earned an impression.

KEEP READING

The page count was never the question.

Count the genuinely distinct things your buyers search for and answer those. If that number is small, a small site is the right site. If it is genuinely large, the size takes care of itself, and none of the limits people worry about will be what stops you.

Start now