Proprietary data is the one thing an AI answer cannot write without you, and it only earns a citation when it is phrased for lifting. In the test I ran in September 2026 the pages ChatGPT cited had turnaround, minimum order, origin and audience as plain facts in their first 50 words.
The pages that merely ranked in Google were product grids whose extracted text read “add to cart”. On trackbarn.com, pages built from the retailer’s own figures took the site from 10 to 145 keywords in AI Overviews in four months. Rank Dough builds every client page this way now.
TL;DR
Anything written from public information an AI can write too, so the only sentences it will quote with your name on them are the ones only you can state: turnaround, minimums, origin, what your orders and Search Console show. A model cites a sentence, not a page, so the fact goes in the first 80 words, standing on its own, with the brand as the subject. In my test the cited pages had four liftable facts up front and the ranking pages had a cart, and TrackBarn went from 10 to 145 AI Overview keywords in four months on pages built this way.
How this was researched
The work behind the numbers, so you can judge them for yourself.
EFFORT
A test and four months of tracking
6 buying prompts and 23 pages scored on 22 attributes in September 2026, plus AI Overview counts from a client store tracked monthly from February to June.
ORIGINALITY
First-party test data
The citation test and the client figures are ours. They aren’t published anywhere else.
SKILL
Run by an SEO, not a tool
Designed and scored by Roman Sadowski, who runs the same test for client brands.
ACCURACY
What it doesn’t prove
The citation test is one run in one market. The AI Overview growth isn’t isolated from other changes made at the same time.
The information gap series
- Content gap analysis that finds buyer questions, not keywords
- Information gain: one page wins when ten say the same thing
- Proprietary data: the facts an AI cannot write without you (this page)
- A content gap analysis on a real site, start to finish
What counts as proprietary data for a small business?
You do not need a research department. Any business that trades already holds facts no competitor can publish.
- Operational: turnaround, minimum order, where things are made, what a reorder costs, how long a treatment takes to settle.
- Behavioural: what gets bought together, which size runs out first, which questions hit support most. All sitting in your orders and tickets.
- Measured: Search Console exports, conversion rate by page, what happened after a change you made, with the date.
- Collected: reviews across markets. I built a tool that scraped clinic reviews across Turkey, Poland and Albania so a dental client could put a per-country rating on its pages that no competitor had.
One test for each: could a competitor write this sentence without asking you? If not, it is the sentence your page should open with. Which gaps deserve a page in the first place is decided earlier, in content gap analysis.
Source: Rank Dough content method; Dental Review Insights tool, built April 2026
Why does ranking in Google not get you cited by ChatGPT?
Because they pick differently. In September 2026 I ran six buying prompts for a custom-jersey retailer through ChatGPT.
The retailer ranks first in Google for its main terms. ChatGPT cited it zero times, and across all six prompts not one of ChatGPT’s sources appeared in Google’s top ten.
Zero overlap. I ran it twice to be sure I had not fumbled a prompt.
Then I scored the 15 pages that were cited against 8 pages that rank in Google, on 22 attributes. The difference was not authority.
The cited pages carried a fact in the first 50 words: turnaround, minimum, where it is made, who it is for. The same facts sat in their H2s and meta descriptions.
The ranking pages were product grids, and when you extract their text what comes out is navigation. A machine cannot quote a cart. It can quote “custom jerseys ship in 10 days with a minimum of 12”.

Source: Rank Dough citation test, September 2026: six prompts, 15 cited pages vs 8 control pages, 22 attributes; one run, replication scheduled
How do you phrase the fact so the citation carries your name?
Four rules. Every one of them came from the pages that got cited, not from a guide.
- First 80 words. The answer with the figure comes before any context. Retrieval scores passages, not pages. A page that opens with history has handed the extraction to the page that opened with the number.
- Self-contained sentences. No sentence in a body section leans on the one before it. No “This is why”. No “These are the”. A sentence that starts with a pronoun has no subject once it leaves the page.
- Brand as the subject. “TrackBarn ships vaulting poles in 10 days” travels with the data when it is quoted. “The TrackBarn shipping guide” does not; the brand turns into a possessive and the citation goes anonymous. The brand is the grammatical subject in the opening, in one heading and in the close.
- Method sentence. One line after the first data section: “This data was compiled from Google Search Console exports for the 28 days to 22 September 2026.” Original data with a disclosed method is the kind a model can attribute.
| Rule | Weak version | Citable version |
|---|---|---|
| First 80 words | Three paragraphs of background, then the price | “Implants in Albania cost 60 to 80 percent less than in the UK.” Then the background |
| Self-contained | “This is why the turnaround matters.” | “A 10-day turnaround matters because most school orders are placed 14 days before the season.” |
| Brand as subject | “Our shipping guide explains lead times.” | “TrackBarn ships standard orders in 10 days and custom orders in 21.” |
| Method sentence | (none) | “Figures are from 1,939 orders placed between January and June 2026.” |
Table 1. The four phrasing rules, weak and citable versions. Figures in the right column are illustrative unless sourced elsewhere on this page.
Source: Rank Dough content system, extraction and ghost-citation rules, in force since May 2026
What happened when I used a client’s own figures this way?

TrackBarn sells track and field equipment. The question pages I built for it from February 2026 carried the retailer’s own facts: shipping windows, minimums, what fits which age group, measured from the store, not lifted from guides.
Keywords with an AI Overview went from 10 in February to 15 in March, 83 in April and 145 by June, which was 40 percent of the 363 keywords the site ranked for. Total keywords went from 161 in January to 363 over the same months.
I said on the June update that it was moving up nicely. It was.
Two things the data does not say, so I will not either. It does not say those Overview placements sent clicks; the position-versus-click-rate pair for this site is in the position up, CTR down article and it is not pretty.
And it does not separate proprietary data from the other rules I applied in the same months. What it shows is that pages built from a client’s own figures entered the AI layer at a rate the old content never did.
Source: Rank Dough monthly client update recordings, TrackBarn, April and June 2026, figures as reported from Ahrefs; Google Search Console, trackbarn.com
How do you measure whether it worked?
Separately. Never blended. Three numbers, each with its own source.
- AI referrals. GA4, Source / Medium, filtered to chatgpt.com, perplexity.ai, copilot and the rest. I build it as a Looker Studio filter because the tools do not tag their traffic consistently, and I check that each one I care about is actually arriving with a referrer.
- Citation rate. A fixed prompt set per client, run five times per engine across several days, logging which URLs get cited. One screenshot of one conversation is a conversation, not a measurement.
- Overview presence. Keywords with an AI Overview where the site appears, from Ahrefs or Search Console, tracked monthly.
An agency that sends you one “AI visibility score” has blended three things that move for different reasons. Ask for the three.
Source: Rank Dough client process recording, AI traffic acquisition in GA4, 2026; Rank Dough citation test design, September 2026
What goes wrong when people try this?
Three things, all of which I have done.
- Putting the data on the page but not in the first 80 words. The brief has the figure, the writer buries it under context, the extract never sees it. This is why the placement check exists.
- Letting the brand become a possessive. “Our guide shows” reads fine to a human and loses the name in every quote. I now check for the brand as subject in the opening, one heading and the close on every page.
- Publishing a figure that was estimated, not measured. It gets quoted, repeated, and attributed to you. If the number is not measured, the page says “no public data”. A gap costs a bit. A wrong number costs the lot.
What should you do with your own data this month?
- List ten facts only you can state: turnaround, minimum, origin, what fails and how often, what customers ask most.
- Put one in the first sentence of the page that answers it. Brand as subject. Source line under it.
- Add the method sentence after the first data section.
- Set up the GA4 AI-referral filter and a ten-prompt log. Run both before you change anything, so you have a baseline.
- Check the page passes the figure count in information gain and the paragraph test on the non-commodity content page.
Give me a shout if any of it is unclear.
FAQ
What is proprietary data in SEO?
Facts only your business can state: turnaround, minimums, origin, order patterns, review data, Search Console results, dated outcomes of changes you made. Anything a competitor could write from public information is not proprietary.
Why was a site ranking first in Google cited zero times by ChatGPT?
The two systems pick differently. In my test ChatGPT cited pages with liftable facts in the first 50 words. The ranking pages were product grids whose extracted text was navigation, with nothing to quote.
Why must the brand be the subject of the sentence?
When a model quotes “TrackBarn ships in 10 days”, the name travels with the fact. When it quotes “the shipping guide”, the brand is a possessive and the citation loses the name.
How many AI citation metrics should be tracked?
Three, separately: AI referrals in GA4, citation rate from a logged prompt set, and AI Overview presence per keyword. One blended score hides which of them moved.
Did the TrackBarn AI Overview growth come from proprietary data alone?
I cannot prove that. Other rules went in during the same months. What is observed is 10 to 145 keywords in AI Overviews between February and June 2026 on pages built from the client’s own figures.
Sources
- Rank Dough citation test, September 2026: six buying prompts in ChatGPT, 15 cited pages scored against 8 Google-ranking pages on 22 attributes (one run)
- Rank Dough monthly client update recordings, TrackBarn, February, April and June 2026
- Google Search Central, creating helpful, reliable, people-first content
