Skip to content
← Blog

What decides whether ChatGPT names your company

The research is public and predates the question being asked. Statistics, quotations and named sources, in that order.

Start hereSources read 2026-09-095 min read

The traffic bargain that funded the web has changed shape, and Cloudflare published the measurement. In a piece dated 1 July 2025, David Belson and Sam Rhea reported crawl-to-refer ratios for the week of 19 to 26 June 2025: the number of HTML page requests a platform’s crawler makes for every visitor it sends back through a referral. The range ran from about 70,900 to 1 for Anthropic’s crawler at one end down to 0.1 to 1 for Mistral, which sent ten referrals for every page it requested. The spread matters more than any single figure. Being crawled and being sent readers have come apart, and for most platforms the first now happens tens of thousands of times for each instance of the second.

What should you optimise for when the reader never arrives?

Which changes what there is to optimise. If a model answers in place and the reader never arrives, then the question is not how to rank but whether the answer names you. That is a different objective with different mechanics, and it has been studied since before most people started asking about it.

What does the research say gets a page cited?

The paper to read is GEO: Generative Engine Optimization, by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande, published at KDD 2024 and on arXiv as 2311.09735. They built a benchmark of queries with the web sources that answer them, then tested what happens to a page’s visibility inside a generated answer when the page is changed in specific ways. The headline is a visibility improvement of up to 40%. The useful part is which changes produced it: adding relevant statistics, adding quotations, and citing sources. Statistics helped most on questions of law, government and opinion, and quotations on explanatory and historical ones. Keyword stuffing, the tactic that transferred most obviously from the old game, was not among the things that worked.

That result reads as mechanical once stated plainly. A model composing an answer needs material it can attribute. A page carrying a specific number with a named source and a date hands it a sentence it can lift and stand behind. A page carrying an argument, however good, hands it nothing quotable, and the model paraphrases somebody else who was easier to quote. Fluency helped in the same study, which is the same point from the other side: text a model can parse cleanly is text it can reuse.

How do you write a page so a model can quote it?

So the practical form of writing for extraction is narrow. Answer one question per block in one sentence, then support it. Attach every figure to a named source and the date you read it, in the sentence, not in a footnote, because a model lifting the sentence does not lift the footnote. Say the same facts in the same words everywhere they appear, since a model resolving what your company is will encounter your pricing page, your comparison page and your documentation, and disagreement between them reads as uncertainty about the entity. Give the page a date it was last checked, and mean it.

Two things to avoid, one obvious and one not. The obvious one is inventing figures, which works until somebody checks and then costs the whole document its credibility, and models are increasingly the thing doing the checking. The other is serving different content to crawlers than to people. Microsoft has been explicit that cloaking is a violation, and the practice fails on its own terms anyway: the crawler that reads the special version is not the surface a buyer eventually lands on.

Why does a model need to know what your company is?

The structural point about being an entity at all is worth separating from the writing. A model has to be able to say what your company is in one clause before it can recommend it for anything, and that clause is assembled from whatever is consistent across the pages that mention you. A category sentence you repeat, a page that compares you against the alternatives a buyer is actually weighing, and a set of answers to the questions buyers ask do more for this than any volume of publishing, because they give the model the shape of the answer rather than more prose to compress.

How do you measure whether models mention you?

What none of this buys is a guarantee, and the measurement problem is real. There is no rank to check. An answer varies between users, between sessions and between models, and a platform that cites you in July may not in August because its retrieval changed rather than because your page did. The honest way to track it is to ask the models the questions your buyers ask, on a schedule, and record what comes back, which is a sampling exercise rather than a metric. Our own tool at instinctgtm.com/tools/ai-visibility does exactly that and no more than that, and it is free without an account.

The summary is short and slightly deflating. The way to be named by a model is to publish things that are true, specific, sourced and dated, in the same words each time, on a page a crawler can read. The tactics that worked in the research are the tactics of somebody being careful about facts. This is the first time in twenty years that the incentives of the search game and the incentives of writing honestly have pointed the same direction, and it is worth taking while it lasts.

Questions this answers

How do you get cited by ChatGPT?
Publish pages carrying specific claims with a named source and a date in the sentence itself, answer one question per section in one sentence, and keep the same facts in the same words across every page that mentions you. The GEO paper (Aggarwal et al., KDD 2024, arXiv 2311.09735) measured visibility gains of up to 40% from adding statistics, quotations and source citations, and no gain from keyword stuffing.
What is answer engine optimization?
Optimising to be named inside a generated answer rather than to rank in a list of links. It is a different objective from search ranking because the reader may never visit the page, so the unit of success is the citation rather than the click.
Does AI search send any traffic?
Far less than crawling would suggest. Cloudflare reported on 1 July 2025 that for the week of 19 to 26 June 2025, crawl-to-refer ratios ranged from about 70,900 HTML requests per referral for Anthropic’s crawler down to 0.1 for Mistral, which sent ten referrals per request.
Does keyword stuffing work for AI search?
No. It was among the tactics tested in the GEO paper at KDD 2024 and was not one of the ones that improved visibility in generated answers, unlike adding statistics, quotations and cited sources.
Should you serve different content to AI crawlers?
No. Microsoft states plainly that cloaking is a violation of its guidelines, and the practice defeats itself, because the version a crawler reads is not the page a buyer lands on.
How do you measure whether AI models mention your company?
By sampling rather than by rank, since there is no ranking to read and answers vary between users, sessions and models. Ask the models the questions your buyers ask on a schedule and record what comes back.

More of this kind