All posts

Programmatic SEO for Directories in 2026: What Still Works

Programmatic pages survived March 2026 only where the data was genuinely differentiated. How to decide which pages to generate, the quality floor per URL, and what to deindex.

Your AI chat can build this directory.

Describe the niche, watch the agent design the fields and fill the catalogue. Free plan, no card.

Start free

Programmatic SEO is not dead. The version of it that most directories were running is.

Google's March 2026 core update enforced the scaled content abuse policy hard, and sites generating thousands of pages by swapping variables into a template lost 60 to 90 percent of their rankings. Sites generating thousands of pages from genuinely distinct structured data were largely untouched.

Same technique. Opposite outcome. Here is where the line sits.

The line, stated precisely

Google's policy targets pages produced primarily to manipulate rankings rather than help users. In practice the operative question per URL is:

Does this page answer a query no other page on this site answers, using data that genuinely differs from the neighbouring pages?

What passed: property listings with real inventory, comparison pages with live pricing, business directories with verified typed entries, availability-driven travel pages.

What failed: "Best [service] in [city]" across four hundred cities where the only differences are a proper noun and a reordered paragraph.

Volume was never the violation. Volume without differentiation is.

Why directories are the natural fit, and the natural risk

A directory generates pages from data by design. That is the whole model.

Whether it lands on the safe side depends entirely on one thing: how much your entries actually differ from each other in structured terms.

Two hundred entries carrying name, description and a link produce two hundred pages that are 90 percent identical. Two hundred entries carrying eight typed attributes produce two hundred genuinely different documents, with different facet values, different comparisons and different reasons to exist.

The automation is the same. The data model is not.

Which pages to generate

The answer is fewer than you want to.

Generate:

  • Category pages where you have real depth. These are your primary targets.
  • Facet combinations that a buyer would actually search for, and only where the combination has entries behind it.
  • Collections grouping entries by a meaningful shared trait, such as gunsmiths offering partner-armoury delivery.
  • Entry pages, always. Each entry is an asset and its page is intrinsically unique if the schema is real.

Do not generate:

  • Location pages for every settlement in a country when you have entries in nine of them
  • Every possible combination of every facet, which produces combinatorial thin pages
  • Comparison pages between entries that have nothing meaningful to compare
  • Pages that exist because the URL pattern would be nice to own

The quality floor per URL

Set this as a hard rule and enforce it in the generation logic, not in your intentions.

Five to ten real entries minimum. Below that the page has nothing to show and should not exist.

A unique introduction that says something true about this specific page. Not a template sentence with a variable slot. If you have forty category pages, that is forty short pieces of writing, once. It is a day of work and the highest-return day available.

Meaningfully varying attribute data. If every entry on the page has the same values, the page is not a comparison, it is a list with extra steps.

Appropriate schema markup. ItemList on collection pages, the right entity type on entry pages, BreadcrumbList everywhere.

A distinct query it answers. If two of your generated pages would satisfy the same search, merge them.

If a candidate page fails any of these, the correct action is not to generate it. That is the entire discipline.

What to do about pages you already generated

If you scaled before the update and are now looking at thousands of thin URLs:

Audit by template, not by page. Search Console will show which patterns lost. The losses cluster.

Fix the data model first. Every cosmetic improvement on top of a weak schema is wasted work.

Deindex what cannot be improved. Noindex the pages that might earn substance later, 410 the ones that never will. Leaving them indexed dilutes the quality signal for the whole site.

Enrich the survivors. Real introductions, complete attribute data, proper markup.

A site with 400 substantial pages outperforms the same site with 4,000 pages of which 3,600 are thin. Removing pages to gain traffic is counterintuitive and it works.

Where the search traffic actually is now

AI Overviews appear on roughly half of all queries, with the heaviest coverage on informational searches. Explainer pages that answered "what is X" have largely stopped earning clicks.

Comparative and transactional intent held up much better. Which is convenient, because that is what a directory does.

The practical adjustment: shift your programmatic effort from informational patterns toward comparison and shortlisting patterns. "Gunsmiths certified for X who deliver to a partner armoury" is a page a filterable dataset can serve and an answer box cannot.

Also worth building for: being cited. Structured data, clear headings, direct answers in the first sentence of a section, and question-shaped subheadings all raise the odds of appearing inside AI-generated answers. Citations without clicks still build branded search, and the visitors who do click from those surfaces convert far better than average.

Doing this without engineering

Most of the discipline above is data work, not code. It is enforced by having a real schema and refusing to generate below the floor.

On DirectoryFast the schema is designed for the niche before entries are filled, and every entry is validated against it, which makes the differentiation structural rather than something you have to police. Collections give you indexable grouping pages under your control rather than automatic combinatorial generation.

Whatever the platform, the rule is the same: the page count is an output, not a target.

FAQ

Is programmatic SEO safe in 2026?

Yes, when each page carries distinct data and clears a real quality floor. The technique was never the problem.

How many pages is too many?

There is no number. There is a ratio: what proportion of your pages would a reader call substantial. If it is low, the count is too high.

Was AI-generated content penalised?

Not as such. Google permits AI assistance. Mass output with no added value was penalised, and AI made that easy to produce.

Should I delete thin pages or noindex them?

Noindex if the URL could earn substance later, 410 if it never will. Both beat leaving them indexed.

How long does recovery take?

Months, and it tends to arrive with subsequent core updates rather than on a schedule. Fix the data model, wait, resist the urge to generate more.

Describe your niche and see the typed schema an agent designs →

Related reading

Stop reading, start one

Everything above is easier to do than to read about. Describe a niche in your AI chat and see what the agent proposes.

Start free, no card