AI Visibility

LLM Visibility Is the Metric Your Analytics Cannot See

LLM visibility is how often, and how accurately, an AI assistant names your brand when someone asks it a question you should be the answer to. You measure it by putting real buyer questions to the models, counting whether you appear at all, and recording where in the answer you land and what gets said about you.

I checked mine. A mention search across ChatGPT responses returned nothing for my domain, nothing for my name, and nothing for my product. Not low. Nothing.

That is an uncomfortable number to publish. It is also the reason I can write this honestly, because I am not selling you a recovery from a position I never held.

How is LLM visibility different from a search ranking?

A ranking is a position in a list of links. LLM visibility is whether you exist inside the written answer that increasingly sits above that list, or replaces it.

The practical difference is that you cannot see one of them. Google’s own documentation confirms that sites appearing in AI features are included in the overall search traffic in Search Console, inside the ordinary Web search type. There is no separate report. So the surface that is taking your clicks is folded into the same row as the surface that used to give them to you, and your dashboard shows you one number where there are now two behaviours.

This is why LLM visibility has to be measured deliberately. Nothing you already have will surface it by accident.

What is AI visibility, and is it the same as LLM visibility?

AI visibility is how often your brand turns up in the answers AI systems write, and yes, it means the same thing as LLM visibility. The two phrases describe one behaviour, and measuring AI visibility uses the same method either way.

They arrived from different directions, which is why both are still in use. “LLM visibility” names the thing doing the answering, a large language model. “AI visibility” is the wider word, and it stretches to cover surfaces that are not a chat window at all, such as the AI Overview sitting above your ordinary search results. Vendors pick whichever phrase their buyers type. The wider acronym tangle, AEO and GEO against plain SEO, is a separate question and I have taken it apart in what AEO and GEO actually mean.

I use them interchangeably here and I would suggest you do the same, with one caution. When you compare two products, check which engines each one actually queries rather than which word is on the pricing page. A tracker branded for AI visibility that reads one assistant is narrower than one branded for LLM visibility that reads five.

What metrics actually measure LLM visibility?

There are four that matter, and they answer different questions: whether you appear, where you appear, whether you are used as a source, and how much of the category you hold. Any single one of them on its own will mislead you.

Profound, which publishes its definitions openly, describes citation share as a measure of “how often a brand’s domain is cited as a source”, and share of voice as “the AI equivalent of market share in search”. Those two are not the same thing, and the gap between them is usually where the real problem sits.

MetricWhat it countsWhat it misses
Visibility scoreHow often you appear at all across tracked questionsSays nothing about position or about what was said
Mention positionWhere in the answer you land, first or fifthA first mention inside a rarely asked question is still rare
Citation shareHow often your domain is used as a sourceYou can be named without being cited, and cited without being named
Share of voiceYour share of the category’s total visibilityMoves when a competitor moves, even if you did nothing

Sentiment is a fifth metric and most tools offer it. I would treat it as a late-stage concern. Sentiment on a sample of two mentions is not a finding, it is noise with a label on it.

What is an AI visibility score, and should you build one?

An AI visibility score is a single number that compresses the metrics above into one figure, usually the share of tracked questions you appeared in, sometimes weighted by position or by sentiment. Almost every tool puts one on the front of the dashboard, because a dashboard needs a number.

I would not build a weighted one on your first pass, and I say the same in the step by step audit. A composite feels rigorous and it hides the thing you most need to see. Blend position and sentiment into a low appearance rate and the result is a middling score that reads like a tuning problem, when what you actually have is absence.

Use the plain version instead. Answers you appeared in, divided by questions you asked. It is crude, it only moves when something real moves, and you can explain it to somebody in one sentence, which is more than most composites manage.

How do you measure LLM visibility?

You write down the questions your buyers actually ask, put them to the models, and count the results. Every tool in this category is a convenience layer on top of that one loop.

The version that costs nothing takes an afternoon, and I have written it up step by step as how to run an AI visibility audit. Write twenty to forty real questions, the ones a buyer types when they do not yet know your name. Not “what is Vantage”, which only a person who already knows you would ask. Closer to “how do I compare my site against a competitor without buying an enterprise tool”. Run each one. For each answer, record four things: whether you appeared, where in the answer you appeared, what was said about you, and which sources were cited instead of you.

That last column is the one people skip and it is the most useful. The domains cited in place of you are the ones the model currently trusts on your topic. That is a target list, not a curiosity. It is the same logic I use when reading a competitive report, which I have written about in what competitor intelligence means for a growing brand: the evidence is only worth having if it tells you what to do next.

What is query fan-out, and why does it change what you measure?

Query fan-out is the technique where one question from a person becomes many searches behind the scenes. Google describes both AI Overviews and AI Mode as “issuing multiple related searches across subtopics and data sources” in order to build a single response.

This breaks the habit of tracking a keyword list. You are no longer competing for the phrase the person typed. You are competing across a spread of subquestions you never see, generated on the fly, any one of which can be the one that pulls in a source. A page that answers the headline question and nothing around it can lose to a page that answers the six questions underneath it.

So measure at the level of the question, not the keyword. Ten well-chosen buyer questions will tell you more than two hundred tracked phrases.

What really drives LLM visibility?

Being genuinely useful on the topic, in prose a machine can lift without ambiguity. There is no technical trick underneath it, and Google has said so in unusually plain language.

Its guidance on generative AI features states that “Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add”, that “There’s no requirement to break your content into tiny pieces for AI to better understand it”, and that creating content people find unique, compelling and useful “will likely influence your website’s presence in generative AI search in the long run more than any of the other suggestions”.

I have watched this pattern arrive more than once. I started in SEO in November 2010, on brands like Citibank, Taj Hotels and Tata Motors, and added performance marketing from 2014. Fifteen years on, I have seen new search and ad surfaces turn up on roughly the same schedule, and every one of them landed the same way. First a technical checklist sold as the secret, then a slow realisation that the checklist was table stakes and the content was the actual variable. Answer engines are on the same curve. I wrote about where that leaves ordinary SEO practice in how to win in SEO in the age of AI.

The one thing I would add to Google’s framing is structural rather than technical. Answer the question in the first two sentences under the heading that asks it. Not because a model rewards the format, but because a passage that answers a question completely and on its own is a passage that can be quoted without the surrounding paragraph. Make it easy to lift and it gets lifted.

Do you need an LLM visibility tool?

Not to start, and probably not to find out whether you have a problem. A tool buys you scale, history and a schedule. It does not buy you the answer, and it cannot tell you which questions matter in your market.

Worth knowing before you shop: in my own keyword research this month, the phrase “llm visibility tool” carried a cost per click of just over eighty dollars in the United States, against roughly twenty-eight for “llm visibility” itself. That is the market telling you how much a vendor will pay to reach someone at the buying end of this question, and it explains why almost every page answering it is published by a company selling one. Read those pages with that in mind.

Start manual. Move to a tool when the counting becomes the bottleneck rather than the thinking, and when you get there, the three categories these products fall into and what each actually measures are in LLM SEO tools and what they measure.

How often should you re-measure?

Monthly, with the same question set, on roughly the same day. Model answers vary between runs even when nothing about you has changed, so a single reading is a data point and not a trend.

That repetition is all AI visibility tracking really is. Same questions, same interval, written down, which is why a spreadsheet does the job as well as a subscription until the volume gets away from you.

Change the questions and you have thrown away your baseline. Add new ones by all means, but keep the original set intact alongside them, or you will never be able to tell improvement from a change in what you asked.

What do you do when the number comes back zero?

You write it down, dated, and you treat it as the starting line rather than a verdict. A zero is not a failure of the work you have done, it is a measurement of a surface you have not competed on yet.

Mine is zero as I publish this, and I have put the full baseline and its caveats on the record rather than describing it in the abstract. What I am doing about it is the ordinary version: publishing genuinely useful things on the specific questions I want to own, making the answers easy to lift, and re-measuring monthly against the same list. I will publish the second reading whether or not it moved, because a public baseline that only gets an update when the news is good is marketing, not measurement.

If you would rather not assemble the question set and the counting by hand, Vantage runs the comparison for you against up to three competitors, including how your brand shows up inside AI answers, and returns a graded report in about thirty minutes. The findings are AI created, so read them with your own judgement before you act on them. Your first analysis is free, which is enough to see your own number without deciding anything.

Either way, get the number. It is the only part of this that cannot be argued with.

Written by
Vineeth Nair

Fifteen years in growth: VP Digital Lead at Vodafone Idea, digital lead for Nestlé Indonesia at dentsu, and nine years at Performics on brands like Citibank and Taj Hotels. Now co-founder at ShopLoco and Adroit Digital. He writes here and builds tools in the Lab.