Rankdough treats deduplication in programmatic SEO as a ranking problem, not a tidiness task: on dentaltourismalbania.com a page at position 1 in July 2026 fell to position 9 by September because a second URL on the same site answered the same question, and five more position-1 pages vanished unnoticed while total traffic rose sixfold.
TL;DR
Deduplication in programmatic SEO is a ranking problem that shows up three months after launch, not a tidy-up before publishing. On dentaltourismalbania.com “can I eat chocolate after tooth extraction” fell from position 1 on 1 July 2026 to position 9 on 22 September because a second page answers the same intent, and 5 more position-1 pages stopped returning a URL. My system now runs three dedupe passes, and the one that matters cannot run until the pages have ranked.
How this was researched
The work behind the numbers, so you can judge them for yourself.
EFFORT
Page by page, two dates
Ahrefs top pages for one client site compared on 1 July and 22 September 2026, every page at position 1 checked for what happened to it.
ORIGINALITY
A loss most reports hide
The pages that disappeared under a growing traffic total. Our own client data, not a case study someone else published.
SKILL
Run by an SEO, not a tool
Found by Roman Sadowski in a routine monthly check, and now built into the three dedupe passes.
ACCURACY
Limits
Visit figures are Ahrefs estimates, not analytics, and it’s one site.
What does cannibalisation look like in the data?

Ahrefs top pages, compared 1 July against 22 September 2026, dentaltourismalbania.com:
- “Can I eat chocolate after tooth extraction”: position 1 to 9, estimated visits 30 to 1
- “When can I eat chocolate after tooth extraction”: position 10 to 11, visits 28 to 4
- “What soft food to eat after tooth extraction”: position 1 held, visits 25 to 3, with two near-duplicate soft-food URLs on the same site
- “When can I eat McDonalds after tooth extraction”: position 12 to 9, visits 89 to 13, which is a volume drop rather than a ranking loss
- Five pages with July traffic of 141, 64, 41, 25 and 23 visits that return no URL in September
| Page (question) | Position 1 Jul | Position 22 Sep | Visits 1 Jul | Visits 22 Sep |
|---|---|---|---|---|
| can I eat chocolate after tooth extraction | 1 | 9 | 30 | 1 |
| when can I eat chocolate after tooth extraction | 10 | 11 | 28 | 4 |
| what soft food to eat after tooth extraction | 1 | 1 | 25 | 3 |
| when can I eat McDonalds after tooth extraction | 12 | 9 | 89 | 13 |
| five pages with no URL in September (141, 64, 41, 25, 23 visits in July) | 1 | none | 294 | 0 |
Table 1. Ahrefs top pages, dentaltourismalbania.com, 1 Jul vs 22 Sep 2026. Visits are Ahrefs estimates.
This data was compiled from Ahrefs Site Explorer top-pages exports taken on 1 July and 22 September 2026, matched by URL.
The two chocolate pages are the clean case. Same site, same question, two URLs with different word order.
In July one ranked first and the other tenth. By September both had lost, because the search engine resolved them as one intent and demoted the pair rather than picking a winner.
That is the pattern to look for: not one page falling, but a cluster of near-identical pages falling together.
Why did the spreadsheet dedupe pass both pages?
The keyword deduplicator in my system was built on 13 March 2026. It compares keyword lists, strips exact duplicates, and surfaces unique terms per list. In May it gained URL-derived comparison: it reads titles, H1s, H2s and meta descriptions from a URL list, derives keywords from them, and compares those against a candidate list.
Both chocolate questions passed every version of it. They are different strings.
“Can I eat” and “when can I eat” share four of six words but are not identical, and no string rule would merge them without also merging questions that deserve separate pages. The clustering layer was supposed to catch it: group keywords the way a reader would, so “dental implant cost” and “how much are dental implants UK” land in one cluster.
It did that at small scale. At 1,600 keywords on 3 June it put baseball terms in a dental cluster.
At 5,000 to 11,000 keywords it failed outright until batching with a cross-batch pass was added on 9 May, and it was still unreliable a month later.
Source: Rankdough content system build record, deduplicator (13 Mar 2026), URL comparison (17 to 18 May), clustering (9 May, 3 Jun)
The lesson is not that the tool was bad. It is that intent overlap is not visible in the keywords.
It is visible in which pages the search engine returns for them. Two keywords belong to one page if their results pages overlap, and that information does not exist until something ranks.
What are the three dedupe passes, in order?
- At intake, on keywords. Exact and near-exact string duplicates removed. Lists compared against each other so a term already covered by an earlier batch is not briefed twice. Cheap, mechanical, catches perhaps a third of eventual duplicates.
- Before briefing, on URLs. Candidate keywords compared against terms derived from the site’s existing titles and headings. Stops the system briefing a page that already exists under a different working title. This is the pass most content operations skip, and it is why migrated sites end up with 2019 and 2024 versions of the same article.
- After ranking, on SERP overlap. Once pages have positions, pull the queries each page ranks for from Search Console. Where two URLs rank for the same queries, or where one URL’s position fell as a sibling’s rose, merge them: keep the stronger URL, fold in any unique content from the weaker, and 301 the weaker. This is the pass that would have caught the chocolate pair, and it cannot run on day one.
Rankdough runs the third pass monthly. It reads two reports: pages whose position fell while a sibling with overlapping queries exists, and pages that held a top-3 position last month and return no URL this month. Both reports came out of the September comparison above.
Why does “publish wide, then merge” beat “dedupe first”?

The conventional advice is to resolve duplicates before publishing. In practice that means guessing which of two questions the search engine will treat as the same intent, and guessing wrong in both directions: merging questions that deserved separate pages, and keeping pairs that did not.
The data from the same site argues for the opposite order. In September 2025 dentaltourismalbania.com had 884 keywords ranking at position 11 or worse and 12 in the top 3.
In September 2026 it had 59 at position 11 or worse and 364 in the top 3, after peaking at 593 in July. The long tail did not get optimised; it got replaced. Publishing wide produced the ranking data; the ranking data showed which pages to keep.
Source: Ahrefs keywords history by position band, monthly; Google documents canonical selection and redirects as the consolidation tools: Google Search Central, Consolidate duplicate URLs
The cost of that order is what the chocolate pages show: a period where duplicates compete and both lose. The cost of the other order is pages never written because a spreadsheet rule thought they were duplicates.
On this site the first cost is a few visits a month for one quarter. The second is unmeasurable, which is the problem with it.
What is the missing-URL report?

Five pages that held position 1 in July and 294 estimated visits a month between them return no URL in Ahrefs in September. Nobody noticed, because the site’s total traffic was up 6.1 times year on year (284 to 1,734 estimated visits). A green top line hides individual losses.
That is the second monthly check: every page that held a top-3 position last month, matched against this month. A page that drops out gets a status check (removed, redirected, noindexed, or deindexed) before anything new is written. Restoring a page-one page is cheaper than earning a new one.

What should you run on your own site this month?
- Search Console, Pages, 28 days, sorted by position change. For every page that fell, search the site for a URL answering the same question. If one exists, you have a pair.
- For each pair, compare the queries both rank for. More than a third shared means merge.
- Pull last month’s top-3 pages from Ahrefs or Search Console. Check each still returns 200 and is indexed.
- Only then brief new pages. Nothing new goes out while a known pair is still competing. That order is how Rankdough runs every programmatic site it manages.
FAQ
Why not merge duplicates before publishing?
Because intent overlap is only visible in the results pages, which do not exist until the pages rank. Pre-publish dedupe catches string duplicates and misses intent duplicates such as “can I” versus “when can I”.
How do I know two pages are cannibalising rather than one just dropping?
Both pages lose position in the same period, and they rank for overlapping queries in Search Console. One page falling alone is a different problem.
Which URL do you keep when merging?
The one with more clicks and backlinks over the last 90 days. Unique sections from the weaker page are folded in, then the weaker URL is 301-redirected.
How often should the SERP-overlap pass run?
Monthly, alongside a check that last month’s top-3 pages still return a URL. Both reports came out of a July-to-September comparison that found two cannibalising pairs and five missing pages.
Does this apply at small scale?
Less often, but yes. Any site with more than one page per topic can have a pair. The test is the same: overlapping queries and a shared fall.
Sources
- Google Search Central, Consolidate duplicate URLs
- Google Search Central, Redirects and Google Search
- Ahrefs Site Explorer, top pages, dentaltourismalbania.com, 1 July 2026 compared with 22 September 2026
