• An AI visibility tracking tool is judged on 10 criteria, six of which measure the reliability of the data produced rather than the features advertised.
  • The single most discriminating criterion is model access mode: a measurement taken through an API with no web search enabled does not describe what your prospect actually sees in ChatGPT.
  • LLMs are non-deterministic, so a tool that queries a prompt once a week delivers a snapshot, not a measurement.
  • The 20-minute reproducibility test, doable during any free trial, reveals whether a tool genuinely samples responses or simply caches them.

Twenty platforms have been telling the same story for eighteen months: your customers ask ChatGPT questions, and your brand is nowhere in the answer. The diagnosis is correct. The trouble starts when you reach for your card, because every available comparison hands you the same four selection criteria: how many engines, how many prompts, which metrics, what price.

Those four criteria are easy to check. They are also the least useful. Two tools with identical spec sheets can produce scores forty points apart on the same brand and the same prompt, because they do not query the models the same way.

I sell an AI visibility tracking tool. Let me get that out of the way: I co-founded Cockpyt AI, so I have an obvious interest in you picking mine. The grid below is built to be used against Cockpyt as readily as in its favour, and I will show you at the end of this article which criteria my own tool fails.

Does ChatGPT recommend your brand?

Measure your presence and find out which brands are named instead of yours. No credit card.

Start a 14-day free trial →

How do you choose an AI visibility tracking tool?

You choose an AI visibility tracking tool by scoring ten weighted criteria, six covering measurement reliability and four covering functional scope. A tool scoring below 60 out of 100 produces numbers you will not be able to defend in front of a client or a board.

Here is the grid. It is deliberately scored so you can compare two quotes side by side instead of comparing two sales pitches.

Criterion What you check Weight
Model access mode Raw API or browser simulation, publicly documented 14
Web search enabled Answer grounded in live web results or generated from training data 12
Repetitions per prompt Number of actual calls behind one tracked prompt 12
Metric granularity Mention, citation and recommendation kept separate, not merged 12
Data replayability Raw response viewable, exportable, timestamped 10
Link to real traffic Google Analytics 4 and Search Console connection 10
Engines covered ChatGPT, Gemini, Perplexity, Mistral, AI Overviews 10
Actual response volume Prompts multiplied by variations multiplied by engines 8
Geolocation Configurable and documented 6
Actionability Prioritised action plan, not just a dashboard 6

The four criteria every comparison publishes sit at the bottom of the table. They total 30 points. The six reliability criteria, which almost nobody asks a vendor about before signing, total 70.

The 4 visible criteria, and why they fall short

Public criteria are elimination filters, not decision tools. They tell you whether a tool is out of the running. They never tell you whether it is good.

Engines covered. Counting engines is the most expensive false debate on this market. A tool tracking nine models when 85 % of your audience uses ChatGPT bills you for eight decorative dashboards. The question that matters is whether Mistral is covered for a French-speaking B2B market, and whether Google AI Overviews are included, live in France since July 2026.

Prompt volume. The number on the pricing page means nothing on its own. A 50-prompt plan testing each prompt once a week on one engine yields 200 responses a month. A 15-prompt plan tested five times a week across three engines yields 900. Always multiply prompts, variations and engines before comparing two prices.

Metrics. Almost every tool advertises mentions, citations, share of voice and sentiment. The word that counts is often missing: recommendation. Being named in a list of ten players is not the same as being singled out as the right answer. A tool that merges those two states into a single score leaves you flying blind.

Price. Entry tickets start around 29 dollars a month for lightweight trackers and climb to 295 dollars a month for premium platforms. In its June 2026 comparison, Justa notes that more than twenty platforms are competing around the same narrative. Price alone says nothing. Price divided by responses generated says a lot.

The 6 reliability criteria vendors do not publish

The reliability of an AI visibility tracking tool rests on its model querying protocol, and that protocol is rarely documented. L’Agence WAM made the point back in January 2026: most solutions reveal only part of their methodology, without specifying which prompts, which models or which weighting. Here are the six questions to put to a vendor in writing before you sign.

1. Raw API or browser simulation?

A tool querying the OpenAI API measures a model’s response under laboratory conditions. A tool simulating a browser measures what a real user sees. Both measurements are legitimate. They are not comparable. LLM-GEO.fr documented this in April 2026, pointing out that most tools never specify which of the two they perform.

In my view, this criterion should be disqualifying. A vendor who will not answer it in writing is selling a score, not a measurement.

2. Is web search enabled?

An API call without a search tool enabled returns what the model retained from training, several months out of date. The same prompt with search enabled returns an answer built from pages crawled that day. A tool measuring the first case will tell you your new brand does not exist, while ChatGPT has been naming it correctly for three weeks.

3. How many real calls sit behind one tracked prompt?

Language models are non-deterministic. The same question asked twice produces two different answers. A tool querying a prompt once per cycle is not measuring, it is photographing. Ask for the number of calls per prompt per cycle. Below three, the data will not support client reporting.

4. Are mention, citation and recommendation kept separate?

Three very different states hide behind the word visibility:

  • Mention: your name appears in the answer, with no link and no favourable context.
  • Citation: your URL is referenced as the source of the information.
  • Recommendation: the model names you as the answer to the stated need.

A tool that adds these three states into a single percentage produces a flattering, unusable indicator. You will know you moved from 30 to 45 %, without knowing whether you gained recommendations or plain mentions inside a list.

5. Is the data replayable?

A score with no raw response behind it cannot be verified. You need to open any data point, read the full text generated by the model, and see the date, time and exact prompt sent. Without that trace, you can neither challenge a number nor defend it in a meeting. The Media Leader framed this well in April 2026: media audience studies also rest on samples, but those samples are standardised and audited, which this market has not achieved yet.

6. Is real traffic tied to the measurement?

Knowing you appear in 40 % of answers says nothing about visits generated. A tool connecting to neither Google Analytics 4 nor Search Console leaves you with an awareness metric disconnected from the business. This is the point that sinks most GEO results presentations at board level.

The 20-minute reproducibility test

You can verify a tool’s reliability yourself during its free trial, with no technical skill required. The protocol runs to five steps and takes twenty minutes.

Step 1 – Pick a commercial prompt

Choose a purchase-intent query from your market, along the lines of “best invoicing software for small businesses”. Avoid branded queries, which are too stable to reveal anything.

Step 2 – Run three on-demand scans

Trigger the same prompt three times, at three different points in the day. If the tool offers no manual scan, log that as a replayability failure.

Step 3 – Record presence and rank

For each run, note whether your brand appears and where it sits in the answer. Record the first three competitors named as well.

Step 4 – Calculate the spread

Subtract the lowest score from the highest across your three runs.

Step 5 – Interpret

A zero spread across three runs signals caching: the tool is serving you the same answer. A spread above 30 points signals sampling too thin to produce a trend. A spread between 5 and 20 points matches the normal behaviour of a model queried properly several times.

Finish by asking the same prompt manually inside the engine’s own interface. If the manual result bears no resemblance to the tool’s, you have just discovered that the vendor measures an API with no web search.

The 4 tool families and who they suit

The market splits into four families, and choosing a tool starts with choosing a family. This segmentation, proposed by the agency Justa in June 2026, remains the most workable to date.

Full-stack platforms cover monitoring, analytics and content production in one interface. Broad coverage, average depth on each component, entry tickets between 99 and 300 dollars a month. Suited to marketing teams with no existing stack.

Pure trackers handle monitoring and citations well, with no optimisation layer. Good cost per response, but you see the problem without knowing what to fix. Suited to teams that already run SEO and content in-house.

Modules bolted onto an SEO suite spare you an extra subscription. Clean integration with your existing tooling, prompts often fixed, limited model coverage. Suited to heavy users of a suite they already pay for.

Repositioned content generators produce LLM-optimised text without measuring anything. A useful complement, never a substitute for a tracker.

In my view, the one structural mistake is buying a full-stack platform without someone dedicating at least two days a week to working the data. A dashboard with no operator is a fixed cost line, not a management tool.

The 5 mistakes that lead to the wrong purchase

  • Buying before writing your prompts. The relevance of your tracked queries determines the value of the measurement. With no list of thirty strategic prompts, you are not ready to subscribe, you are ready to run a workshop.
  • Counting engines instead of counting responses. Nine models queried once are worth less than three models queried five times.
  • Confusing presence with recommendation. The first metric climbs easily. The second is the one that generates leads.
  • Accepting a proprietary score without its formula. Ask what the score aggregates and with what weighting. A refusal is an answer.
  • Ignoring the state of your site. A measurement tool pointed at a poorly structured site with no authority produces data, not results.

Full disclosure: where Cockpyt AI lands on this grid

Cockpyt AI is the tool I co-founded with Laurent Séjourné. Here is its honest run through the grid above.

What it clears. Three engines tracked (ChatGPT, Gemini, Perplexity), each prompt tested from several angles every week to smooth out variability, mention and position kept separate inside the Cockpyt Score, raw responses viewable, Google Analytics 4 and Search Console connections tying visibility to traffic, a Query Fan Out page covering reformulations, a Sources page and a Competitors page for share of voice, and a monthly technical and editorial GEO audit with a prioritised action plan. Pricing starts at 29 euros excluding VAT per month on the annual plan, with a 14-day trial and no credit card.

What it fails. Mistral is not tracked today, a genuine limitation on a French-speaking B2B market. Google AI Overviews and AI Mode are not integrated yet. And the prompt volume on entry plans is calibrated for small to medium organisations: a brand needing to cover several hundred queries across several markets will be better served by a premium platform.

Apply the grid yourself, to Cockpyt as much as to anyone else. That is exactly what the scoring tool published alongside this article is for.

Frequently asked questions

Does an AI visibility tracking tool replace an SEO tool?

No. The two measure different surfaces: your position in a results page on one side, your presence inside generated text on the other. A GEO tool reports no search volumes, no backlinks and no technical indexing errors. Budget for both, not one.

How many prompts do you need for a reliable measurement?

Count on 20 to 30 prompts for a single-product business, and 80 to 150 for a catalogue or a multi-segment market. Beyond that, the difficulty is no longer volume but quarterly review: your prospects’ queries evolve, and a frozen list loses relevance within six months.

Why do two tools give different scores for the same brand?

Because they do not query the models the same way. Access mode, web search activation, model version, number of repetitions and call geolocation all produce diverging results on an identical prompt. No measurement standard exists on this market yet.

Can you measure AI visibility for free?

Yes, at small scale. Manually querying ten prompts each month across ChatGPT, Gemini and Perplexity, then logging results in a spreadsheet, costs two hours a month and gives a usable trend. Automation pays off beyond thirty tracked prompts.

Do GEO tools cover Mistral?

Rarely. The model most relevant to part of the French-speaking B2B market is absent from the majority of mainstream platforms. If your audience is French and technical, ask before subscribing. The question will eliminate several candidates.

How often should you measure AI visibility?

Weekly measurement is enough to steer, with occasional manual scans after each publication or optimisation. Daily frequency generates noise from model variability, not signal.

Does a small local business need a tool?

Only if you have identified local commercial-intent queries where your competitors appear. Start by manually testing five queries along the lines of “best [trade] in [city]”. If you are absent from all three, a tool becomes useful for tracking the fix.

Sources

  • Justa — “Les meilleurs outils GEO et AI visibility en 2026 : comparatif et framework de décision”, Benoît Eveillard, June 2026 — justa.fr
  • LLM-GEO.fr — “Outils de tracking GEO : Peec.ai, Profound, Otterly et les vraies questions sur la précision”, April 2026 — llm-geo.fr
  • L’Agence WAM — “Outils GEO : promesses, biais et méthode pour mesurer votre visibilité”, January 2026 — agence-wam.fr
  • The Media Leader FR — “Comment les outils de mesure GEO accèdent aux réponses des LLM (et leurs limites)”, April 2026 — fr.themedialeader.com
Florian Zorgnotti

I’m Florian Zorgnotti, an SEO consultant based in Nice since 2016. I’ve led 300+ projects, specializing in WordPress, Shopify, and Generative Engine Optimization (GEO) to help brands grow their visibility in search and AI platforms. Linkedin