SEO
How to build thousands of SEO pages safely
Google does not have a page limit. It has a threshold for whether a page was worth making. Those are different problems, and only one of them is solvable by writing more.
THE SHORT ANSWER
Building SEO pages at scale is defensible when every URL represents a materially distinct search intent and tells the reader something a page about a different intent could not. It stops being defensible the moment page count becomes the goal. The failure mode Google penalises is not volume, it is producing many pages that are substantially the same page, which is why a hundred genuinely different pages are safer than ten near-identical ones.
The contradiction worth addressing first
Elsewhere on this site we tell service businesses, plainly, not to build a page for every nearby city. That advice is correct and we are not walking it back. Google's spam policies name doorway pages explicitly, and a set of pages produced by swapping a place name into a template is the textbook example.
So how does the same company sell campaigns that publish thousands of pages?
Because page count was never what made doorway pages a problem. What made them a problem is that forty city pages built from one template are forty copies of a single page wearing different hats. They serve one intent, badly, forty times. The number is a symptom. Sameness is the disease.
A large site can be built the other way round: start from thousands of genuinely different questions people ask, and produce one page per question that actually answers it. That site is large for the same reason an encyclopedia is large, and it is not doing anything Google objects to.
The test every URL has to pass
Before a page is worth building, it has to survive a set of questions that have nothing to do with how long it will be. If a proposed page fails any of these, producing it anyway makes the site worse, not bigger.
- Is the intent materially distinct? Would somebody searching this be satisfied by a page you already have? If yes, you do not have a new page, you have a section of an existing one.
- Is there anything true and specific to say? A page needs information that could only appear on that page. If the only difference between two drafts is a noun, the second one has nothing to add.
- Is the business actually behind it? A page about a service you do not sell, or an area you do not work in, is a promise you cannot keep. It also tends to attract enquiries you have to turn down, which costs more than the page earned.
- Does it fit the entity you already are? Pages that sit outside what your site is demonstrably about are the hardest to get traction for, because nothing else on the site supports them.
- Would you be comfortable if a customer read it? This is not a soft criterion. It is the fastest available proxy for whether a page was made for a person or for a crawler.
Distinct intent, in practice
Distinct intent is easy to say and easy to fudge, so it helps to look at where the line actually falls.
Genuinely distinct: what a fault code means, what it costs to fix, whether it is safe to drive with, and which part usually causes it. Four questions, four answers, four pages. Somebody searching one of them is not served by the others.
Not distinct: the same fault code page written once per city. The searcher does not want a local page about a fault code. There is no local version of the answer.
The useful habit is to write the question the page answers, in the searcher's words, before writing the page. If two questions collapse into the same sentence, they are one page.
Cannibalisation, which is the failure nobody plans for
When two pages target the same intent, they do not each get half the traffic. Something worse happens: Google picks one, inconsistently, and often picks the weaker one. Links and relevance that should have accumulated on a single strong page are split across two mediocre ones, and both underperform the page you would have had if you had written one.
This is the specific mechanism by which building more pages produces less traffic, and it is why the qualification step above is not bureaucracy. At scale, cannibalisation is not an edge case, it is the default outcome of generating pages from a list without checking them against each other and against what the site already has.
- Check new pages against the existing site, not only against each other. The page you are about to duplicate is usually one you already published.
- Near-duplicate detection has to be insensitive to the swapped noun. Two pages that differ only by a place name should be caught as duplicates precisely because a naive comparison says they are different.
- When two pages do overlap, the fix is usually to merge them into the better one and redirect, not to keep both and hope.
Architecture matters more than the pages do
A large set of good pages with no structure connecting them is a pile, not a site. The structure is what lets Google understand that the pages relate to each other and to the thing you actually do.
- Internal links have to mean something. Pages should link to the pages a reader would genuinely want next. A block of links to every other page is not architecture, it is furniture, and it is trivially recognisable as such.
- Nothing should be orphaned. A page reachable only from the sitemap is a page telling Google that nothing on your own site considered it worth linking to.
- Canonicals have to be honest. If two URLs really are the same page, say so. Canonical tags are a description of reality, not a lever for hiding a duplication problem you chose not to fix.
- Indexability should be deliberate. Some pages exist for users and do not need to be in the index. Deciding that on purpose is healthy. Discovering it by accident is not.
- Related pages should form clusters. A group of pages around one subject, linked to a page that introduces the subject, is legible to a crawler in a way a flat list never is.
Publish in waves, not all at once
Dropping several thousand URLs into a sitemap on a Tuesday is a bad idea even when every page is good. It gives you no way to learn anything, and no way to stop.
Controlled publishing means releasing a portion, waiting for Google to work through it, and reading what happened before releasing the next. The first wave tells you whether the pages are being crawled at all, whether they are being indexed, and whether the ones that do get indexed attract anything. That is information you cannot buy any other way, and it is worthless if you have already published everything.
It also gives you an exit. If the first wave performs badly, you have a problem with a few hundred pages instead of a problem with the whole site.
The part almost everybody skips: pruning
A large site is not finished when it is published. Some proportion of any large build will not earn its place, and the honest thing to do is remove or merge those pages rather than leave them.
Pages that attract nothing over a reasonable period are not neutral. They are the pages that make the rest of the site look thinner than it is. Merging several weak pages into one strong one is frequently the single highest-return action available on a large site, and it is unpopular because it feels like undoing work.
This is also the honest answer to "how many of these pages will work?" Nobody knows in advance. What you can commit to is measuring it and acting on the answer.
- Judge a page on whether it earns impressions and clicks, not on whether it exists.
- Give it long enough to be judged fairly. Weeks, not days.
- Merge before you delete. The content is usually fine; it was the separate URL that was wrong.
Where automation genuinely helps, and where it does not
Production is the part of this that scales well. A system can draft, structure, link and validate pages far faster than a person can, and for a build of any size that is the only way it happens at all.
Qualification does not scale the same way. Deciding that an intent is real, that the business can stand behind the page, and that it is not a version of something you already published, is where the judgement lives. A system can surface candidates, flag overlaps and refuse obvious duplicates. It cannot decide what your business is willing to promise.
The honest framing is that automation moves the constraint. It does not remove it. A pipeline that generates ten thousand pages from a keyword list without a qualification step is not a faster version of good SEO, it is a faster version of the doorway problem.
What good looks like from the outside
If you are evaluating somebody else's large-scale SEO work, or your own, these are the things worth checking. None of them require access to the system that built it.
- Pick three pages at random and read them side by side. If you can tell what is different without hunting, that is a good sign.
- Search for two of the intents yourself. If the same page from the site answers both, they should have been one page.
- Look at how many of the published pages have ever received an impression. This is visible in Search Console and it is the number nobody volunteers.
- Ask what has been removed or merged since launch. An answer of "nothing" on a large build means nobody has looked.
Common questions
Does Google have a limit on how many pages a site can have?
No. There is no published page quota, and large sites rank perfectly well. What is finite is Google's willingness to keep crawling and indexing pages that have not shown themselves to be worth it. The constraint is quality-shaped rather than count-shaped, which is why adding pages to a site with a thin-content problem tends to make it worse rather than better.
How is this different from doorway pages?
A doorway page exists to catch a search and funnel the visitor somewhere else, and doorway sets are usually the same page repeated with a swapped location or keyword. The distinguishing feature is that the pages are not meaningfully different from each other. A large set of pages that each answer a genuinely different question is not a doorway set, however many of them there are. Google's spam policies describe the behaviour, not the volume.
So should a local service business build a page for every city?
No, and nothing here changes that. A page per area you genuinely work in, with detail only somebody who works there could write, is worth having. A page per city you would like to serve is a doorway page. Most local service businesses have a few dozen genuinely distinct pages available to them, not thousands, and building to the real number is the whole point.
Will AI-produced pages get the site penalised?
Not for being AI-produced. Google's position is about scaled content abuse, meaning content produced primarily to manipulate rankings rather than to help anyone, and it applies the same way to content produced by people. Production method is not the test. Whether the page was worth making is.
How long before you know whether it worked?
Months rather than weeks, and the shape of the answer arrives before the size of it. Crawling and indexing of a first wave is visible fairly quickly. Whether the pages accumulate impressions and clicks takes longer, and the pages that eventually perform are usually not the ones anyone predicted at the start, which is the argument for publishing in waves and measuring rather than committing to a list up front.
KEEP READING
Most sites do not have thousands of good pages available to them.
The useful question is not how many pages you could publish. It is how many genuinely distinct things people search for that you could answer better than whoever currently ranks. That number is knowable before anybody writes anything, and it is frequently much smaller, or much larger, than expected.