AI Visibility

AI Visibility Tools: Why Two of Them Give You Different Numbers

Point two AI visibility tools at the same brand in the same week and they will give you different numbers. That is not a bug in one of them, and it is not a reason to distrust the category. It is what happens when you measure something that is not stable, and understanding why is the difference between buying one of these tools well and buying one badly.

I have not run a controlled comparison of these products. There are no scores here and no winner. What follows is what the category costs, taken from the vendors’ own published pricing, the mechanisms that make two honest tools disagree, and one measurement of my own that shows how large the gap can get.

What is an AI visibility tool?

An AI visibility tool asks AI assistants a list of questions on your behalf, on a schedule, and records whether your brand appears in the answers. Everything else in the product is reporting built on top of that loop.

The loop is simple. The hard part is that the thing being counted moves. I have written separately about the three kinds of tool in this market and what each one measures, because the category names are slippery and the differences between a tracker, a suite module and a content optimiser matter more than the feature lists suggest. This piece assumes you have picked a category and are now trying to work out what the number it hands you means.

What do AI visibility tools cost?

Entry pricing starts around 29 dollars a month and enterprise plans start around 1,000. The variable that moves the price is almost never engine coverage. It is the number of prompts you are allowed to track.

Otterly.ai publishes its full price list: Lite at 29 dollars a month for 15 search prompts, Standard at 189 for 100, Premium at 489 for 400, and Enterprise from 1,000. Extra prompts are sold at 99 dollars per 100 per month. Its base plans cover ChatGPT, Google AI Overviews, Perplexity and Microsoft Copilot, with Claude, Google AI Mode and Gemini sold as add-ons.

What you are buyingWhat actually variesWhat it means for the number
Prompt allowance15 prompts at the entry tier, hundreds higher upA small allowance measures a narrow slice of your category, so the score is precise about very little
Engine coverageFour engines as standard at Otterly, with three more as paid add-onsTwo tools covering different engines are not measuring the same population
Refresh frequencyDaily, weekly or on demand, depending on planA weekly reading and a daily reading of a moving number are different instruments
Location and languageOften a single market unless you pay for moreA United States reading and an India reading of the same brand are different results

Profound sits at the other end of the coverage question and lists Perplexity, ChatGPT, Claude, Gemini, Grok, Microsoft Copilot, DeepSeek and Google AI Overviews. Set that against Otterly’s four included engines and the disagreement between two tools stops being mysterious. They are counting across different populations before either one runs a single query.

Why do two tools report different numbers for the same brand?

Four things differ between any two tools, and each one moves the number on its own: which engines they query, which prompts they ask, when they ask, and where they ask from. Stack all four and two honest products can report results that look nothing alike.

The engine question is the largest of the four, and I have a measurement for it. On 8 August 2026 I ran the same set of eleven category queries across two engine groups on the same day, through my own research engine, and looked at who was being cited. On ChatGPT, seo.ai came back with 134 mentions. On Google’s AI surfaces, the same domain on the same queries on the same day came back with 3,686. That is a factor of twenty seven, from one variable, with everything else held still. A tool covering only ChatGPT and a tool covering only Google would have told me two completely different stories about the same competitor.

The timing question is not a matter of tool quality either, because the engines themselves do not promise a stable answer. OpenAI’s own API documentation offers a seed parameter so developers can receive “(mostly) deterministic outputs across API calls”, and notes that even then “determinism may be impacted due to necessary changes OpenAI makes to model configurations”. The word in the brackets is doing a lot of work. If the vendor of the model will only claim mostly, no tool sitting on top of it can claim more.

Google says the same thing in plainer terms. Its Search Central documentation states that “AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary”, and that AI Overviews “are only shown when our systems determine that it is additive to classic Search, and as such, often don’t trigger”. A tool measuring a surface that often does not appear will record a zero that means “no AI Overview ran today”, not “you were left out of one”.

What is an AI visibility score actually counting?

Most visibility scores are counting brand string matches inside AI answers, then dividing by something. The counting step is where they break, because a brand name is rarely unique and a string match cannot tell two owners of the same name apart.

I found this in my own data rather than reading about it. When I measured the name “Vineeth Nair” as an entity, the search returned 84 mentions. Not one of them was me. Every single one belonged to a Malayalam film namesake, and a tool reporting on that string would have handed me a healthy looking score built entirely on somebody else’s career. The number was not wrong. It was answering a different question from the one I thought I had asked.

That is worth checking before you trust any share of voice figure. Ask what the tool does with a brand name that is also a common word, a person, or another company in another industry. Ask whether a mention without a link counts the same as a citation with one. Two tools can be counting honestly and still disagree by a wide margin because one counts mentions and the other counts linked citations, and nothing on either dashboard says so.

Can you measure AI visibility without paying for a tool?

Yes, and I would do it before spending anything. Write down the ten questions a buyer in your category would actually type, ask them in the assistants your buyers use, and record what comes back and who gets cited.

That is exactly how I produced my own baseline, and I published the result when it came back at zero. Doing it by hand is slow and it does not scale past a few dozen prompts, but it teaches you something a dashboard cannot: what the answers in your category actually look like, who keeps appearing in them, and which of your questions the engines refuse to answer at all. I have written the longer version of the method in how to run an AI visibility audit.

What you buy when you start paying is frequency and memory. A tool runs the same prompts every week without you remembering to, and keeps the history so you can see a trend rather than a snapshot. That is genuinely worth money once you have something to trend. It is worth very little on day one, when you have no history and no idea which prompts matter.

What should you ask before you buy an AI visibility platform?

Ask which engines are included rather than merely supported, and ask what a zero means. Those two questions separate the products faster than any feature comparison.

The full list I would take into a demo:

The last one matters more than it looks. If you can export the answers, you can check the tool’s arithmetic and you can leave without losing your history. If you cannot, the score is unfalsifiable and you are trusting a vendor’s summary of a number that vendor also chose how to calculate.

Do you need one at all?

If you have never measured your AI visibility, you do not need a tool yet, you need a baseline. Run the questions by hand once, write the number down with the date beside it, and you will know within an hour whether there is a trend worth paying to track.

Buy the tool when the manual run stops being enough: when you have more prompts than you can sit through, when someone else needs to see the number without asking you, or when you need to prove a change over months rather than describe one. Until then the honest answer is that the most expensive part of AI visibility is not the software. It is deciding which questions are the ones you need to win, and no dashboard will do that for you.

If you would rather have the competitive half of that work done for you, Vantage runs the category analysis and hands back what your competitors are doing and where the gaps are. The reports are AI-created, and they are built to be checked: every finding carries its evidence so you can apply your own judgement rather than take the output on trust.

Written by
Vineeth Nair

Fifteen years in growth: VP Digital Lead at Vodafone Idea, digital lead for Nestlé Indonesia at dentsu, and nine years at Performics on brands like Citibank and Taj Hotels. Now co-founder at ShopLoco and Adroit Digital. He writes here and builds tools in the Lab.