Skip to content
Launch GuidesUpdated September 30, 2026

What ChatGPT Actually Cites Now, and How to Get Your Product on the Shortlist

On August 8, 2026 ChatGPT changed who it names. Reddit fell out of the top ten, official pages walked in, and the citation moved to a second step most builders never optimise for. The data, the mechanism, the Cloudflare switch that can drop you out of Google, and the exact order to fix your own site.

ChatGPT citations · after August 8
Findable gets you on the list. Cited is the step after.
GEOAEOChatGPTAI SearchSEOCitationsCloudflareIndie HackingDistribution
Outcome
On the shortlist, then cited
Effort
~1 afternoon of plumbing
Read
~10 min
Channels
ChatGPT · Cloudflare · robots.txt
What you will get
  • ✓Who ChatGPT cites after August 8, domain by domain, from the Profound slides and three independent panels
  • ✓The find-then-hone mechanism: fanouts, site: queries, the safe_urls list, and where the citation is actually decided
  • ✓Which page shapes make the shortlist, with the cite rates behind each claim
  • ✓The Cloudflare crawler setting that can pull you out of Google while you block AI training
  • ✓A five-item fix list in order, and a free 14-check test to run on your launch page
Share

Your buyers are asking ChatGPT what to use. On August 8, 2026 it changed who it names, and almost nobody shipping a product noticed, because the change happened in the step after the one everybody optimises for.

This guide is built from one set of slides and four independent measurements. Profound pulled 3,264,282 ChatGPT citations and put the before and after on a slide. Josh Blyskal presented it at the Shenzhen SEO Conference on September 17. Lily Ray photographed three slides. There is no public deck, so every Profound number here is read off a photo; where another panel measured the same thing, I say what it found too.

The second half is what I did about it on my own site, which scored 33 out of 100 on its own AI readiness test. Twice.

The shift

August 8: the day ChatGPT stopped trusting the crowd

Share of all ChatGPT citations, the week of August 1 to 7 against August 14 to 21. Before: reddit.com 5.46%, techradar.com 1.67%, arxiv.org 0.94%, g2.com 0.81%, wsj.com 0.71%. Reddit alone out-cited the next four domains put together. After: g2.com 0.86%, pubmed.ncbi.nlm.nih.gov 0.66%, microsoft.com 0.61%, consumerfinance.gov 0.60%, pmc.ncbi.nlm.nih.gov 0.59%. The IRS sits at number eight. Reddit, TechRadar and arXiv are gone from the top ten.

Reddit before
5.46%
of all citations, Aug 1 to 7
Reddit after
top 10: out
Aug 14 to 21
New #1 (G2)
0.86%
nothing dominates
Sample
3.26M
ChatGPT frontend citations

Three other panels caught the same cliff with their own prompt sets, and the numbers differ because the panels differ. Promptwatch: Reddit from 3.83% to 0.52%, an 86% drop, with the share of fanouts using a site: operator jumping from 0.37% to 16.8% on August 8. Qwairy: 2.04% to 0.10%, with institutional and .gov share rising from 17.1% to 29.6%. Otterly: Reddit citations per day from 497 to 132 across sixteen brand reports, while official brand pages gained in ten of fifteen. Elmo compared the app to the API: the app's Reddit share fell 92%, the API's stayed flat at under 1%.

Why the API comparison matters
Same model, different result. The retrieval layer in the consumer app changed, not what the model knows. That is why this is a distribution problem and not a prompting one, and why it can change again.

One panel disagrees. Ahrefs Brand Radar, published September 2, still lists reddit.com as the most cited domain in ChatGPT at 16.8% and reports no drop, because it measures a search-backed query set built differently. Expect someone to quote it back at you. The honest reading is that Reddit lost most of its weight in the consumer app for product and advice queries, and that official pages took it.

Now look at the size of the winner. The old number one held 5.46% of every citation. The new number one holds 0.86%. The whole thing spread sideways. Official pages grew, and "official" does not automatically mean your product page. It means the page an authority would publish.

The mechanism

Find, then hone: fanouts and the safe_urls list

The second slide indexes activity so the August 1 to 6 average equals 100. From August 8 to September 7: fanout queries 170, safe URLs 102, citations 112. Seventy percent more searching to cite roughly the same number of pages. The filter got pickier.

The third slide is the useful one. For a question like "best running shoes?", ChatGPT runs a fanout, reasons about what came back, then runs another fanout. Everything it finds lands in a list the slide labels safe_urls. Then it reasons over that list and the slide quotes the model: these are the ones worth citing.

Being findable gets your URL into safe_urls. The citation happens at the step after that.
The guide, in one line

Two studies from earlier in the year put numbers on how narrow that second step is. AirOps ran 15,000 prompts and watched 548,534 pages get retrieved; 15% were cited. A fanout happened on 89.6% of searches, and 32.9% of cited pages only appeared through a fanout query, never the original prompt. Ahrefs looked at 1.4 million prompts: 88.46% of citations came through the search index, and the pages that got cited had titles far closer to the fanout query than the pages that did not. The gate is the title, the slug and the snippet, before anyone reads the page.

Retrieved pages cited
15%
AirOps, 548,534 pages
Searches with a fanout
89.6%
AirOps
Cited only via fanout
32.9%
never for the original prompt
Fanouts after Aug 8
170
indexed, Aug 1 to 6 = 100
What a site: fanout means for you
After August 8 a large share of fanouts interrogate a specific official domain. If a buyer asks for a tool in your category, ChatGPT may run a search against G2, against a government page, and against the vendor it already suspects. Your product page gets into that list when your domain is the obvious official source for your own name and category. Being a thread about yourself no longer works.
The pages

Which pages make the shortlist

Across engines, DeltaV Digital's July study of 25,337 citations found articles at 23.7%, listicles 19.6%, product pages 16.3% and homepages 10.8%. Comparison pages had the highest citations per retrieval of any format, 1.87, about 45% above average. In B2B tech, listicles alone took 61%. On Google AI Mode, Pillarbase found 47.7% of citations point at a single paragraph; the median passage was 117 words and 80% of them led with the answer.

So the shortlist is made of pages that look like this:

  • A title that matches the sub-question. Ahrefs: cited pages scored 0.656 on title-to-fanout similarity against 0.484 for retrieved but uncited pages. AirOps: 50% or more title overlap doubled the cite rate, 20.1% against 9.3%.
  • A natural-language slug. 89.78% of cited URLs had one, against 81.11% of the rest. Small, free, and it compounds.
  • The answer in the first paragraph. A 100 to 200 word capsule that a model can lift whole, with the number, the date and the source in it.
  • A comparison table or an attributed statistic. The formats with the best citations-per-retrieval are the ones that save the model a second search.
  • Age and refresh. The median cited page was about 500 days old. You do not need to be new. You need to be maintained, and your commercial pages need a visible update every few weeks.
Tally, as the proof it pays
Tally, an eleven-person form builder at $5M ARR, counted 125,000 ChatGPT visits in one month, 9.6% of all referrals, and attributes a quarter of new signups to AI answers. Ahrefs reports ChatGPT at 0.5% of visits and 12.1% of signups. Buffer measured LLM traffic converting at 20.15% against 7.06% for organic. Small share, high intent, because the person arrives with the answer half made.

Traffic from a citation also changed shape in May, when ChatGPT started linking brand names inline. Similarweb saw referrals to homepages rise 354.7% in a week and pages per visit go from 3.8 to 4.7. The homepage became the landing page again, which makes the first paragraph of your homepage the most read paragraph you own.

The plumbing

The Cloudflare switch and the plumbing underneath

None of the above matters if the crawler cannot reach you, and in 2026 the most common way to block it is by accident. Cloudflare rebuilt its crawler controls on September 15 after reports in August that the old "Block AI Bots" toggle was returning 403s to Googlebot and Bingbot on some sites. Read this before you touch the dashboard.

You now set three things separately:

  • Search: whether you get crawled for a search index. Allow, Block on pages with ads, or Block.
  • Training: whether your pages train a model. The same three, plus a fourth, Disallow AI Training, which publishes a Disallow line in your robots.txt and leaves crawling alone.
  • Agent: whether an assistant fetches you live for a user. Allow, Block on pages with ads, or Block.
Block on Training takes Googlebot with it
Googlebot, Bingbot and Applebot serve search and training both. Cloudflare's own words: on Training, selecting Block stops them entirely. If you want to stop training and keep search, the setting is Disallow AI Training, not Block. Sites that had the old toggle on were migrated to Search Allow, Training Disallow AI Training, Agent Block on pages with ads. If you had already set Search to Block yourself, it is still Block, and that still costs you Google.

Two more notes from the same change. New domains get a preset: no ads on the site and everything stays on Allow; ads on the site and Training goes to Disallow AI Training with Agent blocked on ad pages. And Bing does not read the robots.txt preference yet; Microsoft is targeting early 2027, so until then the Bing opt-out is the NOARCHIVE tag. AI Labyrinth, the honeypot feature, does not touch Googlebot, so it is not the culprit if you see 403s.

Then the part below Cloudflare. I built a free AI readiness test: paste a URL, it fetches the page once, reads the site around it, and runs 14 checks. No account. I ran it on my own launch directories page and got 33. Then on the homepage. Also 33, on a completely different set of failures. On a site I built for builders, it caught all of this:

  • robots.txt names no AI crawlers at all; GPTBot and ClaudeBot fall through to the default rule
  • no Markdown version of the page, so an agent swallows the full HTML
  • the navigation gets read as if it were content
  • a request asking for text/markdown still comes back as HTML
  • no API catalog at the well-known path, while llms.txt names a live MCP server

The last one stung. The surface exists and the file that advertises it to a crawler does not. Three checks passed: llms.txt answers, nothing important is blocked, and the heading outline is a real outline. This is plumbing: robots.txt, a Cloudflare setting, and whether the page is readable. It is also the part that decides whether you ever reach the hone step.

The order

What to fix first

This is the exact order I am fixing mine. Cheapest and most damaging first.

Today
  • Open Cloudflare's crawler settings and read the Training row. If it says Block, change it to Disallow AI Training. Check Search is Allow.
  • Name the AI crawlers in robots.txt: OAI-SearchBot and ChatGPT-User allowed, GPTBot and ClaudeBot decided on purpose, not by the default rule.
  • Run the 14-check readiness test on your launch page and fix whatever it names first.
This week
  • Rewrite the first paragraph of your homepage and your top three pages as the answer: what it is, for whom, the number, the date.
  • Match titles and slugs to the questions buyers actually ask, in their words. One question per page.
  • Ship a Markdown twin of your key pages, or return Markdown when a request asks for text/markdown.
This month
  • Add one comparison table and one attributed statistic to each commercial page, with a visible updated date.
  • Get your entity data consistent on G2, Capterra and LinkedIn, and listed in the comparison hubs ChatGPT now queries directly.
  • Ask ChatGPT the five questions your buyers ask, monthly, and write down what it names. That is your citation report until you pay for one.

The number I still cannot see is how often ChatGPT and Claude cite this site. Cloudflare's AEO tab is built for exactly that: it runs Claude and GPT on the questions buyers in your category ask and shows how often they cite your site, how often they name your brand, and your share against competitors. I have asked for early access to mine. Agent Readiness is live on the same screen and calls this site Almost ready. That gets fixed too.

If you are launching this week
Do the Today list before the launch, not after. A launch sends crawlers to your pages for a few days; that is the window where a readable answer-first page gets into safe_urls and a blocked or HTML-only one does not.
Coding got cheap. Being found did not. Being cited is a third thing, and it is decided after you are found.
Where this leaves you

FAQ

Did ChatGPT really stop citing Reddit?

Mostly, in the consumer app. Profound's sample of 3.26 million citations shows reddit.com going from 5.46% of all citations to out of the top ten. Promptwatch measured an 86% drop, Elmo 92%, Otterly 73%. The API's citation mix barely moved, which is why it reads as a retrieval change, not a new model. One panel, Ahrefs Brand Radar, still shows Reddit at the top because it measures a different query set.

Is this good or bad for a small product?

Both. Official pages gained, so your own site counts for more than a thread about you. But the new number one domain holds under 1% of citations, so nothing dominates, and the filter got pickier: 70% more background searches to cite the same number of pages. You win by being the clean, answer-first source for your category, not by being loud.

What is a fanout?

A background search ChatGPT runs before it answers. AirOps saw one on 89.6% of searches, and a third of cited pages only appeared through a fanout query. After August 8 many fanouts are site: queries against official domains, which is how government, documentation and vendor pages walked into the top ten.

Does blocking AI training on Cloudflare hurt my Google rankings?

It can. Googlebot, Bingbot and Applebot are mixed-use crawlers that serve search and training. Setting Training to Block stops them entirely. Cloudflare's own instruction since September 15, 2026 is to use Disallow AI Training, which publishes a robots.txt preference and leaves search crawling alone.

How do I check how often ChatGPT cites my site?

Server logs show OAI-SearchBot and ChatGPT-User fetches but not citations. Cloudflare's AEO tab, Profound, Promptwatch, Qwairy and Otterly all run prompt panels and report citation share; the free route is to ask ChatGPT the five questions your buyers ask and note what it names. Run the 14-check readiness test first so the plumbing is not the reason you are missing.

Sources

Disclaimer: the AI readiness test and the directory catalog linked above are mine and free to use without an account. Profound's figures are read from photographs of conference slides and should be treated as a signal, not a paper. Panel numbers differ because panels differ.

Share

Get the weekly recap

New launch case studies, top launches, and what is worth a look. One email a week.