Yext published an analysis of 17.2 million AI citations this month, and the headline is not that AI search is hard. It is that AI search is not one thing.
The four major engines apply fundamentally different retrieval logic. Not different weightings of a shared ranking method. Different logic. Verified, structured, directly distributed data made up 54.53% of distinct citation sources overall, but the mix shifts hard depending on which engine you ask. ChatGPT retrieves largely through Bing's index via OAI-SearchBot and cites seven to eight sources per response. Claude cites user-generated content at two to four times the rate of the other three.
Source: https://www.yext.com/research/ai-citation-behavior-across-models, corroborated by https://cmotech.uk/story/study-finds-ai-models-develop-distinct-citation-habits
That is a large-scale finding, and large-scale findings are easy to nod at and hard to feel. So here is the same thing at a scale of thirty-six.
On 21 August we asked three engines twelve questions about who does answer engine optimisation, in fresh incognito sessions, never naming our own studio. Twelve blind questions, three engines, thirty-six answers. We are a small design studio in Mumbai, so the honest expectation was a low number.
The number was nine of thirty-six. What matters is the shape of it.
ChatGPT named us in nine of twelve. Perplexity named us in zero of twelve. Gemini named us in zero of twelve.
That is not a twenty-five percent problem. It is a one-engine-in-three problem, and those need completely different work. The two engines on zero describe the studio accurately when asked about it directly. They hold the facts. They just never volunteer them.
Watching what Perplexity actually cited explained why. Every source it used in the blind set was somebody else's roundup article, and we were on none of them. It was not evaluating studios and ranking them. It was reading published lists and repeating the names on them. No amount of markup on our own site reaches that, which is worth knowing before anyone sells structure as the fix for it.
ChatGPT, meanwhile, was doing something closer to evaluating. It read what we publish and formed a view.
Two engines, two entirely different mechanisms, one query. That is the Yext finding, reproduced accidentally at a scale small enough to see the individual answers.
The practical consequence for anyone with a website: "improve our AI visibility" is not a task. It is at least two tasks that share almost nothing.
If the engine reads roundups, your work is off-site. Get named in the lists that actually exist, or publish the measurement those lists get written from. Structured data on your own domain does not touch it.
If the engine evaluates, your work is on-site. Say plainly what you do, who for, and what proof exists. Make it liftable. That work is real and it moves that engine, and it will do nothing for the first one.
Doing only the second and calling it AI visibility is the common mistake, partly because it is the half you control and partly because it is the half most vendors sell.
We published our full run, including the questions where all three engines missed us entirely, at https://reidify.design/ledger. The misses are the useful part. A citation rate with no denominator is a horoscope.
One more thing worth saying plainly. Our thirty-six answers are not a study. They are one studio, one day, one set of questions. What makes them worth publishing is that they point the same direction as seventeen million citations do, and that we recorded them before we knew the result.