+353 863834528

Rankdough treats deduplication in programmatic SEO as a ranking problem, not a tidiness task: on dentaltourismalbania.com a page at position 1 in July 2026 fell to position 9 by September because a second URL on the same site answered the same question, and five more position-1 pages vanished unnoticed while total traffic rose sixfold.

✓Why you can trust this article▼
RS

Roman Sadowski · Co-Founder & SEO Lead, Rank Dough

SEO and AI visibility strategist, previously at iProspect. Has run site migrations for Smyths Toys, theaa.ie and ProPlayerTeam, and builds the content and citation-testing systems Rank Dough uses with its clients.


Sources used in this article

  • ✓ developers.google.com
  • ✓ Ahrefs Site Explorer (client data)

Editorial policy. Every figure is traced to its source: client data from Google Search Console and Ahrefs, Rank Dough’s own test records, or public documentation such as Google Search Central. AI tools help with research and drafting. A person checks the finished page, including the title, meta description, structured data and image alt text, before it is published.

✓ Human verified by Roman Sadowski

Last reviewed: October 2026

TL;DR

Deduplication in programmatic SEO is a ranking problem that shows up three months after launch, not a tidy-up before publishing. On dentaltourismalbania.com “can I eat chocolate after tooth extraction” fell from position 1 on 1 July 2026 to position 9 on 22 September because a second page answers the same intent, and 5 more position-1 pages stopped returning a URL. My system now runs three dedupe passes, and the one that matters cannot run until the pages have ranked.

How this was researched

The work behind the numbers, so you can judge them for yourself.

EFFORT

Page by page, two dates

Ahrefs top pages for one client site compared on 1 July and 22 September 2026, every page at position 1 checked for what happened to it.

ORIGINALITY

A loss most reports hide

The pages that disappeared under a growing traffic total. Our own client data, not a case study someone else published.

SKILL

Run by an SEO, not a tool

Found by Roman Sadowski in a routine monthly check, and now built into the three dedupe passes.

ACCURACY

Limits

Visit figures are Ahrefs estimates, not analytics, and it’s one site.

What does cannibalisation look like in the data?

Two chocolate-after-extraction pages on dentaltourismalbania.com: position 1 fell to 9 and position 10 to 11 between 1 July and 22 September 2026, both losing traffic
Figure 1. The duplicate pair lost together. Ahrefs, 1 Jul vs 22 Sep 2026.

Ahrefs top pages, compared 1 July against 22 September 2026, dentaltourismalbania.com:

  • “Can I eat chocolate after tooth extraction”: position 1 to 9, estimated visits 30 to 1
  • “When can I eat chocolate after tooth extraction”: position 10 to 11, visits 28 to 4
  • “What soft food to eat after tooth extraction”: position 1 held, visits 25 to 3, with two near-duplicate soft-food URLs on the same site
  • “When can I eat McDonalds after tooth extraction”: position 12 to 9, visits 89 to 13, which is a volume drop rather than a ranking loss
  • Five pages with July traffic of 141, 64, 41, 25 and 23 visits that return no URL in September
Page (question) Position 1 Jul Position 22 Sep Visits 1 Jul Visits 22 Sep
can I eat chocolate after tooth extraction 1 9 30 1
when can I eat chocolate after tooth extraction 10 11 28 4
what soft food to eat after tooth extraction 1 1 25 3
when can I eat McDonalds after tooth extraction 12 9 89 13
five pages with no URL in September (141, 64, 41, 25, 23 visits in July) 1 none 294 0

Table 1. Ahrefs top pages, dentaltourismalbania.com, 1 Jul vs 22 Sep 2026. Visits are Ahrefs estimates.

This data was compiled from Ahrefs Site Explorer top-pages exports taken on 1 July and 22 September 2026, matched by URL.

The two chocolate pages are the clean case. Same site, same question, two URLs with different word order.

In July one ranked first and the other tenth. By September both had lost, because the search engine resolved them as one intent and demoted the pair rather than picking a winner.

That is the pattern to look for: not one page falling, but a cluster of near-identical pages falling together.

Why did the spreadsheet dedupe pass both pages?

The keyword deduplicator in my system was built on 13 March 2026. It compares keyword lists, strips exact duplicates, and surfaces unique terms per list. In May it gained URL-derived comparison: it reads titles, H1s, H2s and meta descriptions from a URL list, derives keywords from them, and compares those against a candidate list.

Both chocolate questions passed every version of it. They are different strings.

“Can I eat” and “when can I eat” share four of six words but are not identical, and no string rule would merge them without also merging questions that deserve separate pages. The clustering layer was supposed to catch it: group keywords the way a reader would, so “dental implant cost” and “how much are dental implants UK” land in one cluster.

It did that at small scale. At 1,600 keywords on 3 June it put baseball terms in a dental cluster.

At 5,000 to 11,000 keywords it failed outright until batching with a cross-batch pass was added on 9 May, and it was still unreliable a month later.

Source: Rankdough content system build record, deduplicator (13 Mar 2026), URL comparison (17 to 18 May), clustering (9 May, 3 Jun)

The lesson is not that the tool was bad. It is that intent overlap is not visible in the keywords.

It is visible in which pages the search engine returns for them. Two keywords belong to one page if their results pages overlap, and that information does not exist until something ranks.

What are the three dedupe passes, in order?

  1. At intake, on keywords. Exact and near-exact string duplicates removed. Lists compared against each other so a term already covered by an earlier batch is not briefed twice. Cheap, mechanical, catches perhaps a third of eventual duplicates.
  2. Before briefing, on URLs. Candidate keywords compared against terms derived from the site’s existing titles and headings. Stops the system briefing a page that already exists under a different working title. This is the pass most content operations skip, and it is why migrated sites end up with 2019 and 2024 versions of the same article.
  3. After ranking, on SERP overlap. Once pages have positions, pull the queries each page ranks for from Search Console. Where two URLs rank for the same queries, or where one URL’s position fell as a sibling’s rose, merge them: keep the stronger URL, fold in any unique content from the weaker, and 301 the weaker. This is the pass that would have caught the chocolate pair, and it cannot run on day one.

Rankdough runs the third pass monthly. It reads two reports: pages whose position fell while a sibling with overlapping queries exists, and pages that held a top-3 position last month and return no URL this month. Both reports came out of the September comparison above.

Why does “publish wide, then merge” beat “dedupe first”?

dentaltourismalbania.com keywords at position 11 or worse fell from 884 to 59 while top-3 keywords rose from 12 to 593, Sep 2025 to Sep 2026
Figure 2. Publish wide, then merge: the long tail was replaced, not optimised.

The conventional advice is to resolve duplicates before publishing. In practice that means guessing which of two questions the search engine will treat as the same intent, and guessing wrong in both directions: merging questions that deserved separate pages, and keeping pairs that did not.

The data from the same site argues for the opposite order. In September 2025 dentaltourismalbania.com had 884 keywords ranking at position 11 or worse and 12 in the top 3.

In September 2026 it had 59 at position 11 or worse and 364 in the top 3, after peaking at 593 in July. The long tail did not get optimised; it got replaced. Publishing wide produced the ranking data; the ranking data showed which pages to keep.

Source: Ahrefs keywords history by position band, monthly; Google documents canonical selection and redirects as the consolidation tools: Google Search Central, Consolidate duplicate URLs

The cost of that order is what the chocolate pages show: a period where duplicates compete and both lose. The cost of the other order is pages never written because a spreadsheet rule thought they were duplicates.

On this site the first cost is a few visits a month for one quarter. The second is unmeasurable, which is the problem with it.

What is the missing-URL report?

Five dentaltourismalbania.com pages at position 1 on 1 July 2026 with 141, 64, 41, 25 and 23 estimated visits returned no URL on 22 September
Figure 3. Five position-1 pages gone, about 294 visits a month, hidden by a green top line.

Five pages that held position 1 in July and 294 estimated visits a month between them return no URL in Ahrefs in September. Nobody noticed, because the site’s total traffic was up 6.1 times year on year (284 to 1,734 estimated visits). A green top line hides individual losses.

That is the second monthly check: every page that held a top-3 position last month, matched against this month. A page that drops out gets a status check (removed, redirected, noindexed, or deindexed) before anything new is written. Restoring a page-one page is cheaper than earning a new one.

dentaltourismalbania.com monthly organic traffic peaked at 2,266 in July 2026 and fell 23 percent to 1,734 by September
Figure 4. The dip after the peak is where competing URLs and missing pages show up.

What should you run on your own site this month?

  • Search Console, Pages, 28 days, sorted by position change. For every page that fell, search the site for a URL answering the same question. If one exists, you have a pair.
  • For each pair, compare the queries both rank for. More than a third shared means merge.
  • Pull last month’s top-3 pages from Ahrefs or Search Console. Check each still returns 200 and is indexed.
  • Only then brief new pages. Nothing new goes out while a known pair is still competing. That order is how Rankdough runs every programmatic site it manages.

WORK WITH RANK DOUGH

Want this done on your own site?

The three-pass dedupe runs inside our content service, and the monthly missing-URL check is part of every SEO and AI audit.

Get your free snapshot

FAQ

Why not merge duplicates before publishing?

Because intent overlap is only visible in the results pages, which do not exist until the pages rank. Pre-publish dedupe catches string duplicates and misses intent duplicates such as “can I” versus “when can I”.

How do I know two pages are cannibalising rather than one just dropping?

Both pages lose position in the same period, and they rank for overlapping queries in Search Console. One page falling alone is a different problem.

Which URL do you keep when merging?

The one with more clicks and backlinks over the last 90 days. Unique sections from the weaker page are folded in, then the weaker URL is 301-redirected.

How often should the SERP-overlap pass run?

Monthly, alongside a check that last month’s top-3 pages still return a URL. Both reports came out of a July-to-September comparison that found two cannibalising pairs and five missing pages.

Does this apply at small scale?

Less often, but yes. Any site with more than one page per topic can have a pair. The test is the same: overlapping queries and a shared fall.

Sources