The Test That Separates Programmatic SEO From Spam Takes One Question
← Back to Blog

The Test That Separates Programmatic SEO From Spam Takes One Question

August 5, 2026 · 5 min read

Programmatic SEO
Featured book
Programmatic SEO
$3.99Free on Kindle Unlimited
Amazon

A large share of the pages you land on from a search were not written by a person. They were generated — one template, one dataset, thousands or millions of URLs.

Every job board, every travel aggregator, every comparison site, every "X near me" directory works this way. So do a great many of the sites that dominate long-tail search in their categories. This is not a fringe tactic; it is how a substantial part of the commercial web is built.

It is also, in a slightly different configuration, exactly what search engines penalise.

The technique is identical in both cases. The difference is narrow and it is worth being precise about, because "generate pages at scale" describes both a real business and a way to get your domain removed from results.

The One Question

Here is the test, and it survives every policy update because it is what the policies are trying to approximate:

Would this page be worth existing if search engines did not?

If a user arrived at your page by any route — a link from a friend, a bookmark, a direct URL — would they find something that answered their question?

A page listing every flight between two specific cities, with current prices and times, passes. Somebody genuinely wants that, and a human could not reasonably compile it for every city pair. A page that says "Flights from Denver to Austin. Looking for flights from Denver to Austin? We have flights from Denver to Austin," padded to four hundred words, does not. It exists solely to be indexed.

Both were generated from a template. One has a reason to be there.

What Thin Actually Means

The most common failure is thin content, and the term is used loosely enough that people miss what it means.

It is not about length. A short page that answers the question precisely is not thin. It is about whether there is anything on the page that the user could not have guessed from the URL.

The recognisable failure patterns:

Data with no interpretation. A page whose entire content is a number. The population of a town, a stock price, a distance. The data may be accurate and it is not an answer to anything; nobody searching has a question that terminates in an unexplained figure.

Templates where only the variable changes. Every page reads identically except the city name. Search engines are extremely good at spotting this, and so are readers, within about four seconds.

Insufficient depth for the query. "Best restaurants in Boise: here are some options," followed by three names and no addresses, no prices, no reason any of them are on the list. The query implies a standard of usefulness that the page does not meet.

Keyword repetition. Awkward restatement of the target phrase, which reads as badly to a person as it scores to an algorithm.

Notice that each of those is a description of a page that fails the one question. They are not separate rules; they are symptoms of the same thing.

What the Winners Actually Have

The sites that succeed at this all have something the spam versions do not, and it is nearly always one of three things.

Proprietary or hard-to-assemble data. The reason nobody else has this page is that gathering the underlying information took real work — scraping, licensing, aggregating, cleaning, maintaining. The moat is the dataset, and the pages are just how it gets exposed.

Genuine aggregation. The value is in the collection, not any individual item. A user would otherwise have to check eleven sources. You checked them, and the page is the saving.

Real per-page variation. Each page contains substantively different information because the underlying entity is different, not because a variable was substituted. If two of your pages would be equally useful with their contents swapped, the template is doing the work and there is nothing underneath it.

If you have none of the three, you do not have a programmatic SEO strategy. You have a page factory, and the economics of that have deteriorated sharply — search engines have gotten much better at detecting scaled low-value content, and generative tools have made producing it so cheap that the supply overwhelmed whatever advantage there was.

The Part That Changed Recently

For years the trade-off was manageable: produce a large volume of mediocre pages, accept that some fraction ranks, profit on the aggregate.

Two things broke that. The cost of generating plausible text fell to nearly zero, so everybody could do it and the tactic stopped being differentiating. And search engines responded by targeting scaled content produced primarily for rankings rather than for people — which is aimed precisely at the aggregate-mediocrity strategy.

The result is that the middle has been squeezed out. There is still a viable business in generating pages at scale from data that is genuinely worth having, and there is essentially no business left in generating pages at scale from nothing.

Which is a healthier position than the one before it, and it means the interesting work moved. It is no longer about how to produce pages efficiently. It is about what you know that nobody else has bothered to assemble.

The Practical Version

Before building anything, answer three things.

What is the dataset, and why is it hard to get? If the answer is "I generated it," stop.

What does one page look like when it is genuinely good? Build that page by hand, first, and be honest about whether you would be pleased to land on it. Everything after that is a scaling problem.

And how many of these can you make good? Frequently the answer is a few hundred rather than a few hundred thousand, and a few hundred pages that answer a real question outperform a hundred thousand that do not — by a margin that has been widening every year.

Programmatic SEO: Build Traffic at Scale covers the whole method — the anatomy of a programmatic page, search intent at scale, how the aggregators and data publishers do it, finding your opportunity, data collection, template design, technical implementation, and staying out of the spam trap.

From the Catalog

Browse all
New
Wu Zetian
Wu Zetian
China's Only Female Emperor and How She Got There
New
Rapa Nui
Rapa Nui
What Really Happened on Easter Island
$3.99KU🎧
New
Thera
Thera
The Volcano That Shattered the Minoan World
$3.99KU🎧
New
Göbekli Tepe
Göbekli Tepe
The Temple Before Farming
$3.99KU🎧