Key takeaways

An AI visibility score is not a temperature reading, it is a proprietary formula applied to a sample of questions the vendor chose. Change tools and you change formula and sample. The number moves, your visibility does not.

You test two tools in parallel on the same brand, in the same week. One shows 62 out of 100, the other 18. You conclude one of them is lying.

Neither probably is. I co-founded Cockpyt AI, so I am a participant in this market, and I am going to describe the mechanism that produces these gaps, including my own score in the demonstration.

Does ChatGPT recommend your brand?

Measure your presence and find out which brands get cited instead of you. No credit card.

Start a 14-day free trial →

Why do two AI visibility tools never give the same score?

Because a score out of 100 is not a physical measurement. It is a proprietary formula applied to a sample of questions the vendor selected. Two vendors share neither the formula, nor the sample, nor the engines, nor the frequency.

Five causes stack up, listed here by weight.

Cause 1: they do not measure the same questions

This is the heaviest cause, and the least visible.

Some tools start from an aggregated database. Ahrefs builds its index by extracting queries from its keyword database, turning them into questions through People Also Ask and semantic expansion, then running those prompts through each platform. Semrush works similarly, from a database of more than 300 million prompts drawn from clickstream data and its Google keyword set.

Others start from your questions. You declare the queries your prospects actually type, and the tool measures your presence on those.

Both approaches are legitimate and serve different needs. But they cannot produce the same number. A consulting firm well described in the generic questions of its sector will score well on an aggregated database, and close to zero on its ten real commercial queries.

Ahrefs documents this honestly: its index works best for established brands with meaningful search demand, and coverage may be limited for brands that are rarely searched. For a small business, that changes everything.

Cause 2: they do not weight the same things

Two scores out of 100 can add up entirely different dimensions. Three public examples show it.

Tool Score composition Consequence
HubSpot AI Search Grader Sentiment 40 points, presence quality 20, brand recognition 20, exposure 10, market competition 10 More than half the total covers how people talk about you
Semrush AI Visibility Score Mention frequency against the median of automatically identified competitors Relative score: it moves when competitors move, even if you change nothing
Cockpyt Score Presence in the answer, position within it, whether a link appears. Zero without a citation Binary at its base: you are cited or you are not

Take a respected but rarely recommended brand. At HubSpot it scores on sentiment and recognition, up to 60 potential points. On my side it drops to zero on the questions where it does not appear. Both figures correctly describe a different reality: one measures perception, the other presence.

The absolute-versus-relative distinction deserves particular attention. A score relative to a competitor median can rise because your competitors fell back. That is useful information, but it is not the same as “my visibility is improving”.

Cause 3: they query neither the same engines nor the same versions

Scopes differ far more than people assume. Otterly includes ChatGPT, AI Overviews, Perplexity and Copilot in its base plans, and sells Gemini as an add-on. Other tools include Gemini but not Google’s surfaces. A score computed on four engines and one computed on three different engines do not describe the same object.

The model version matters just as much, and almost nobody publishes it. HubSpot documents its own: GPT-5.2 for OpenAI, the Sonar model for Perplexity, and Gemini 2.0 Flash for Google, a lightweight version built for speed. A reading taken on a light model does not exactly describe what a user of the main app sees.

In my view, this transparency about model versions should be a market standard and it is not. When a vendor does not publish it, you cannot know whether its number describes your customers’ experience or that of a cheaper substitute model.

Cause 4: they do not run the same number of passes

This is the most technical cause, and it is often confused with frequency. They are distinct.

  • Frequency determines how often a new measurement is produced. Daily for some, weekly for others, monthly for the Ahrefs index, re-tested once a month on a 90-day window.
  • Number of passes determines how many queries are averaged to produce a single measurement. Meteoria advertises 15 to 30 passes per prompt. Cockpyt AI runs one per question per scan. Tools billed per check generally run one per update.

A daily single-pass tool produces 30 fragile measurements a month. A weekly ten-pass tool produces 4 solid ones. Both can display very different scores for the same brand, and the second will be more stable without being more frequent.

Cause 5: the model itself does not answer the same way twice

This cause stacks on top of all the others, and it is independent of the tool.

A study published by SparkToro in January 2026 ran 12 identical prompts almost 3,000 times through ChatGPT, Claude and Google’s AI, using 600 volunteers. The odds of getting the same brand list twice came out below 1 in 100, and around 1 in 1,000 for the same list in the same order.

A study from the Association for Computational Linguistics published in December 2025 confirms it on different ground: five models configured to be deterministic, tested across eight tasks and ten runs, show performance gaps of up to 15 %, with a maximum spread of 70 % between best and worst result.

In other words, even if two tools used the same questions, the same weighting, the same engines and the same number of passes, they still would not produce exactly the same result.

How many repetitions would it take?

Work published in September 2025 on evaluation reliability provides a rarely cited figure: two repetitions remove roughly 83 % of the ranking inversions caused by a single run, and going from one to three repetitions cuts the standard error by only about 5 %.

That result suits nobody in this market. It shows a single pass is fragile as soon as you want to rank, and that the race to dozens of daily passes buys precision with sharply diminishing returns. These studies cover model evaluation, not brand citation: the transposition is reasonable, it is not demonstrated.

How to compare two tools honestly

Since the scores are not comparable, compare what is.

  • Use a common denominator. Ten identical questions, declared in both tools, on the same engines. Everything else is methodological noise.
  • Count raw observations, not marks. Across those ten questions, in how many answers does your brand appear? That is the only data two tools can produce comparably.
  • Record the competitors cited. If both tools return the same competitors on the same questions, their readings converge, even if their scores diverge.
  • Give them three weeks. A gap between two isolated readings proves nothing. A gap in trend across six readings says more.
  • Ask how the score is composed. If a vendor publishes neither its weighting nor its model versions, you cannot interpret its number. That alone is reason enough to rule it out.

In my view, the most useful rule fits in one sentence: a score is only worth something compared to itself. Following one score across twelve weeks teaches you something. Comparing two scores from two tools only teaches you that the formulas differ.

What to stop expecting from a score

That it be stable day to day. It cannot be, whatever the quality of the tool.

That it gives a reliable ranking position. Rand Fishkin concludes his study bluntly: any tool reporting a ranking position in AI is talking nonsense. What stays measurable is an appearance frequency aggregated across many runs. That caveat applies to every vendor, including the position component of the Cockpyt Score.

That it compares across tools, or across time after switching tools. If you migrate, export your data, keep it as an archive, and start from zero while annotating the switch date in your reports.

That it replaces a decision. A score says where you stand. It does not say what to fix. A diagnosis and an action plan do that.

Frequently asked questions

Which score should I believe if two tools give me 62 and 18?

Both, for what they measure. Look at how each score is composed before deciding. If one weights sentiment heavily and the other requires a citation, a 44-point gap is normal for a brand that is well described but rarely recommended.

Is a relative score better than an absolute one?

Neither, they answer different questions. A score relative to your competitors places you in your market. An absolute score measures your presence independently of others. A relative score can rise without you having improved anything.

Why does my score move when I changed nothing?

Because models do not answer the same way twice, your competitors publish, and engines get updated. A few points of movement week to week is variance. Read the trend over 3 to 8 weeks.

How many times should an AI be queried for a reliable result?

No consensus exists. Academic work on model evaluation suggests two repetitions remove most of the inversions caused by chance, with returns dropping fast after that. The main point is that a single pass is never enough to conclude on an individual question.

Can you add or average the scores of two tools?

No. That would average two different formulas applied to two different samples. The result would mean nothing.

What is a free score worth against a paid one?

It does not depend on price but on method. A well-documented free diagnostic beats a paid score whose weighting you do not know. Ask first how the number is calculated.

Should you test several tools before choosing?

Yes, but on an identical question set and comparing raw observations rather than marks. It is the only protocol that teaches you something about the tools rather than about their formulas.

Sources

  • Rand Fishkin and Patrick O’Donnell, New Research: AIs are Highly Inconsistent When Recommending Brands or Products, SparkToro, January 2026 — sparktoro.com
  • Berk Atıl et al., Non-Determinism of “Deterministic” LLM System Settings in Hosted Environments, Proceedings of the 5th Workshop on Evaluation and Comparison of NLP Systems, Association for Computational Linguistics, December 2025 — aclanthology.org/2025.eval4nlp-1.12
  • Miguel Angel Alvarado Gonzalez et al., Do Repetitions Matter? Strengthening Reliability in LLM Evaluations, arXiv:2509.24086, September 2025 — arxiv.org/abs/2509.24086
  • HubSpot, AI Search Grader page, score dimensions and models queried, accessed 2 September 2026 — hubspot.fr/ai-search-grader
  • Semrush, knowledge base, AI Visibility Toolkit: score calculation and data sources, accessed 2 September 2026 — semrush.com/kb/1493
  • Ahrefs, Brand Radar page, AI Visibility Index methodology and refresh frequency, accessed 2 September 2026 — ahrefs.com/brand-radar
Florian Zorgnotti

I’m Florian Zorgnotti, an SEO consultant based in Nice since 2016. I’ve led 300+ projects, specializing in WordPress, Shopify, and Generative Engine Optimization (GEO) to help brands grow their visibility in search and AI platforms. Linkedin