Getting Cited by ChatGPT, Gemini, and Perplexity Is Three Different Jobs

Getting Cited by ChatGPT, Gemini, and Perplexity Is Three Different Jobs
Three engines, three indexes: the citation work for ChatGPT, Google, and Perplexity barely overlaps, so the audit has to be run per engine.

ChatGPT, Google AI Overviews, and Perplexity pull answers from three separate indexes, and a Profound analysis of 680 million citations found that only 11% of the domains ChatGPT cites also show up in Perplexity. Getting cited by all three is three different jobs with three different checklists. The fastest win is confirming your robots.txt is not quietly blocking one of the crawlers before you touch any content.

Most of the "AI search optimization" advice floating around right now is the same 2019 SEO checklist with the word "entity" sprinkled on top. Some of it is fine. But the citation data from the last twelve months points at different work for each engine, so this piece goes engine by engine with the one or two things I think matter for each. (The sibling piece on what AI actually costs a marketing team covers the tooling side. This one is purely about being the source the machine quotes.)

Three engines, three indexes, one robots.txt that can block all of them

The plumbing matters more than people expect. ChatGPT search draws from Bing's index plus its own fetching. Gemini and AI Overviews sit on Google's index. Perplexity runs its own crawler and layers Bing on top. So a page that ranks well on Google is not automatically in ChatGPT's candidate pool, and vice versa.

Each engine also runs two different bots, and the difference trips people up. OpenAI documents four crawlers: GPTBot collects training data, OAI-SearchBot indexes for ChatGPT search, and ChatGPT-User fetches pages when a person asks for one live. Their docs say it plainly: a site can allow OAI-SearchBot to appear in search while blocking GPTBot to opt out of training. Perplexity's setup is nearly identical. PerplexityBot builds the index and respects robots.txt, while Perplexity-User handles live requests and, per their own documentation, "generally ignores robots.txt rules."

The failure I keep seeing is a blanket rule someone added in 2023 when the training-data panic was at its peak. A single User-agent: * block on a subfolder, or a WAF rule that returns 403 to anything with "bot" in the user agent string, and you have removed yourself from ChatGPT search entirely without any dashboard telling you so.

The 20-minute action: pull 30 days of server logs and grep for OAI-SearchBot, PerplexityBot, and Google-Extended. If OAI-SearchBot has hit your site fewer than a handful of times in a month, you are not in that pool, and no amount of content work will fix it. Then read your robots.txt line by line, not from memory. I would bet a decent number of sites reading this have a leftover disallow they forgot about.

Google: rank first, then mostly stop worrying about it

Google is the engine where old-fashioned ranking still does most of the work, though less than it did a year ago. Ahrefs pulled 1.9 million AI Overview citations in July 2025 and found 76% came from pages already in the top 10 for the same query. When they reran a bigger version in early 2026, across 863,000 keywords and 4 million cited URLs, that number had fallen to 38%, with the rest split almost evenly between positions 11 to 100 and pages that did not rank at all.

Ahrefs is careful to say their detection improved between the two studies, so the drop is partly measurement. But the direction is backed up elsewhere. Originality.ai's cut of the same problem put the #1 organic result at a 57.91% chance of being cited. A coin flip with a slight edge. And BrightEdge's 16-month tracker shows the overlap varies enormously by vertical: 75% in healthcare, 23% in e-commerce. When the topic carries risk, Google leans on pages it already trusts. When it is a product query, it roams.

The part that changes what you write is the fan-out. Ahrefs attributes a lot of the spread to Google's query fan-out process, where one search gets broken into several sub-questions and the AI cites whichever pages answer those well. So a page that ranks #1 for the head term but skips the three obvious follow-up questions can lose the citation to a page ranking #40 that happens to answer one of them cleanly.

Google's own AI features documentation says there are no extra requirements to appear in AI Overviews or AI Mode, no special file, no special schema. Gary Illyes has said the same thing on stage. I mostly believe them, with one caveat: "no special optimization" means standard SEO plus covering the fan-out. Our pillar on how Google decides what ranks covers the standard part.

The action with a benchmark: for each of your top 10 target queries, type it into AI Mode and note the follow-up questions Google surfaces underneath. Take the three that recur. If your page does not already answer each of them in a self-contained paragraph of 40 to 70 words, add one. Then check citations again in 30 days. From what I have seen in the data above, a page that ranks in the top 10 and covers the fan-out should be cited in AI Overviews well over half the time. If you are top 10 and cited under a third of the time, the fan-out gap is the first place I would look.

ChatGPT: the engine that does not care where you rank on Google

This is the one that breaks people's mental model. Semrush's analysis of AI search found ChatGPT cited pages ranking in traditional position 21 or worse almost 90% of the time. Read that again if you have been treating Google rank as a proxy for AI visibility. For ChatGPT, it barely correlates.

Part of that is the Bing index. Part of it is source preference. In the Profound data, Wikipedia is ChatGPT's single largest source at 7.8% of all citations, followed by Reddit at 1.8% and Forbes at 1.1%. Profound describes the pattern as favoring "authoritative knowledge bases and established media." My read is that ChatGPT wants a source it can defend. It reaches for things that look like reference material, and it reaches for them from an index most marketers have never opened.

Which brings me to the unglamorous fix. Open Bing Webmaster Tools. Most teams I talk to have never verified their site there, or did once and never went back. Check that your key pages are indexed in Bing at all. Submit the sitemap. Turn on IndexNow if your CMS supports it. Bing traffic on its own might be 3% of your organic, which is why nobody bothers, but Bing's index is the front door to ChatGPT search, and that is a different calculation.

One more wrinkle: ChatGPT does not run a single retrieval path. We covered ChatGPT's four hidden search pipelines already, and the short version is that a quick answer and a thinking-mode answer can cite almost entirely different sets of sources. If your tracking tool only samples one mode, you are seeing a fraction of your real exposure.

The content side for ChatGPT, honestly, is about looking like reference material. Specific numbers with the source and date attached. A definition paragraph near the top that could be lifted as-is. Fewer opinions in the first 100 words, more of them after. That last one is slightly painful for a site like ours, but it seems to be what the citation data rewards.

Perplexity: a Reddit engine with a chat box on top

Perplexity's top source in the Profound dataset is Reddit at 6.6% of all citations, then YouTube at 2%, then Gartner. Profound's phrasing is that Perplexity "prioritizes community discussions and peer-to-peer information." I would put it more bluntly: Perplexity trusts people talking to each other more than it trusts brands talking to people.

That is a hard thing for a brand site to optimize for directly, because you cannot publish your way onto Reddit. What you can do is be the source the thread links to, or be the person in the thread. The second one has to be done honestly. Reddit punishes obvious brand accounts, and the threads that get cited are the ones where a practitioner shares a real result with numbers and someone else argues with them. The threads that rank are messy. That is sort of the point.

YouTube is the other lever, and it is bigger than most teams assume. In the 2026 Ahrefs study, YouTube alone was 5.6% of all AI Overview citations and had grown 34% in six months. We wrote about the odd part of this already: AI engines will happily cite a YouTube video with a few hundred views over the brand's own product page. A 6-minute walkthrough of the thing you sell, titled to match the question people ask, is a Perplexity citation asset in a way your beautifully designed feature page is not.

The action here: pick the five questions where you most want to be the answer. For each, search Perplexity and note whether a Reddit thread or a YouTube video is currently cited. If a Reddit thread is there and you have a real result to add, add it as a person. If a YouTube video is there and it is not yours, you know what to make next. The benchmark I would use is modest: one of your five queries citing something you control within 60 days.

What I would skip, at least for now

llms.txt. I know it is everywhere. The idea is a plain-text file at your root that tells language models which pages matter. The problem is that no major engine has confirmed it reads the file, Google has said outright that it does not need one, and I have not seen a study that ties it to a citation lift. It costs nothing to add, so add it if it makes someone feel better. Just do not count it as work done.

Same with the more exotic schema advice. Article and FAQ schema are fine and cheap. Beyond that, the correlation studies I trust do not show structured data moving citations on its own. The Victorious study we covered put the correlation between domain authority and AI citations at r=0.017, which should also make you suspicious of anyone selling "authority building" as the AI fix.

And to be fair, none of this is entirely new. Search engines have always retrieved from an index, scored candidates, and picked a few. The difference is that there are now three or four indexes that matter, and the scoring on each one leans a different direction. It feels less forgiving mainly because you cannot see the ranking anymore.

The three-engine audit I would run this week

Build a spreadsheet. 20 queries down the side, ideally the ones where a customer is close to buying. Four columns across: Google AI Mode, ChatGPT search, Perplexity, and Gemini. For each cell, record every domain cited, not just yours. Do it manually the first time. It takes about two hours, and you will learn more from reading the actual answers than any tool will summarize for you.

What comes out is a map of who owns each question on each engine. Usually it is a surprise. A competitor you dismissed owns ChatGPT because they maintain a reference page that reads like documentation. A Reddit thread from 2024 owns Perplexity. A YouTube creator with 900 subscribers owns Gemini. Once you can see that, the work stops being generic.

Why bother, given AI search sends a fraction of the traffic Google does? Semrush's number is the one I keep coming back to: visitors arriving from AI search converted 4.4 times better than organic visitors in their data. Small volume, unusually warm. Someone who asked a machine a question and then clicked through anyway had a reason.

A prediction, with a number attached: within 12 months, at least one of the three engines will publish an official citation report inside its own webmaster tooling, the way Google eventually gave us Search Console. My money is on Bing surfacing ChatGPT citations first, because it is the quietest way for Microsoft to make Bing Webmaster Tools matter again. Until then the spreadsheet is the console.

Frequently asked, briefly answered

Does ranking #1 on Google get me cited in AI Overviews? More often than not, but not reliably. Originality.ai puts the #1 result at roughly a 58% citation chance, and Ahrefs found only 38% of AI Overview citations now come from top-10 pages. Rank plus fan-out coverage is the combination that seems to work.

Should I block GPTBot? Blocking GPTBot only affects training data. You can block it and still allow OAI-SearchBot for ChatGPT search, and OpenAI's documentation explicitly supports that split. Just be sure the two rules are actually separate in your robots.txt.

Why I would pick one engine and ignore the other two for a quarter

The temptation is to chase all three at once. I would resist it. The 11% overlap number means the work barely transfers, and a team that spreads across three checklists tends to finish none of them. Look at your referral data, find the engine that already sends you something, and spend 90 days getting cited there consistently before touching the next one. Perplexity if you have a community presence. Google if you already rank. ChatGPT if you have reference-grade content and a Bing index nobody has checked. Then widen out. I do not think the sites that win here are the ones with the most tactics deployed. It is probably the ones that figured out which engine was theirs and stopped pretending the other two were the same job.

Notice Me Senpai Editorial