Key takeaways

  • A single manual check proves nothing: the odds of getting the same brand list twice from ChatGPT are below 1 in 100.
  • A serious manual protocol takes around 1 hour 30 per round for 10 questions across 3 AI engines, which is 18 hours a year at a monthly rhythm.
  • The tipping point sits around 10 questions checked every month. Below that, stay manual.
  • No tool gives a reliable ranking question by question. What can be measured is an aggregated appearance frequency over time.

You open ChatGPT, you type “best SEO consultant in Nice“, your brand does not show up. You try again the next day and there it is. Neither answer taught you anything, and that is the core problem with manual checking.

A study published by SparkToro in January 2026 ran 12 identical prompts almost 3,000 times through ChatGPT, Claude and Google’s AI, using 600 volunteers. The odds of getting the same brand list twice came out below 1 in 100, and around 1 in 1,000 for the same list in the same order.

A paid tool remains pointless for plenty of people. I am going to give you the numeric threshold at which one pays for itself, and the cases where manual testing is perfectly sufficient.

Does ChatGPT recommend your brand?

Measure your presence and find out which brands get cited instead of you. No credit card.

Start a 14-day free trial →

Do you need a tool to check whether ChatGPT talks about your brand?

No, if you track fewer than 10 questions and check once or twice a year. Yes, if you track 10 questions or more every month, or if you have to report to a client or a management team.

The calculation fits in one line. A proper manual protocol takes about 3 minutes per question and per engine, recording time included. Ten questions across three AI engines means 1 hour 30 per round. At a monthly rhythm, 18 hours a year. At 60 EUR an hour, that is 1,080 EUR of billable time against 348 EUR excl. VAT for an entry-level annual subscription.

The threshold does not depend on your company size but on three variables: the number of questions, how often you check, and what your hour is worth. The calculator further down combines them.

The manual protocol that produces a usable result

Typing your brand name into ChatGPT measures nothing. You are testing recall, meaning the model’s ability to repeat what it has read about you. What matters is discovery: does your brand come up when nobody names it.

The minimum protocol has five steps.

  • List 10 buyer questions phrased the way your prospects phrase them, never mentioning your brand. “Which tool tracks AI visibility” rather than “what do you think of Cockpyt AI”.
  • Open a fresh session for each question, in private browsing with no account signed in, to rule out personalisation from your conversation history.
  • Record three things: whether your brand appears, its rank among the brands listed, and which competitors come up instead of you.
  • Repeat on every AI engine your market actually uses, at minimum ChatGPT, Gemini and Perplexity.
  • Date and archive in a spreadsheet. A reading that compares to nothing is worth nothing.

Allow 90 minutes for 10 questions across 3 engines. The time sink is not the querying, it is opening a clean session every time and filling in a structured record.

Three reasons your manual test is not reproducible

Manual testing is not just imprecise. It produces a result you cannot compare with next month’s, for three distinct reasons.

A single pass draws a random card

This is the point the SparkToro study documents best. Each prompt was run 60 to 100 times per platform, and almost every response was unique in three ways: the list of brands, their order, and the number of items returned. An isolated query gives you a draw, not a measurement.

This behaviour is not a defect. Models are probabilistic by design, and variation is expected operation.

Your history and your account distort the answer

A signed-in account carries your conversation history, your saved preferences and your context. If you have already discussed your company with ChatGPT, you are no longer measuring your visibility, you are measuring its memory of your exchanges. Private browsing solves part of the problem, not all of it.

Your test conditions change without your knowing

Between two checks a month apart, the model may have been updated, web search may have triggered in one case and not the other, the interface may have shipped a new version. You think you are comparing two measurements. You are comparing two protocols.

A study published by the Association for Computational Linguistics in December 2025 tested five models configured to be deterministic across eight tasks and ten runs. It reports performance gaps of up to 15 % between runs. Even with settings locked, stability does not exist.

The tipping point: when a tool starts paying for itself

The threshold is calculated, not guessed. Here are the orders of magnitude for a manual protocol at 3 minutes per question per engine, across 3 AI engines.

Questions tracked Frequency Hours per year Time cost at 60 EUR/h Verdict
5 Twice a year 1.5 h 90 EUR Stay manual
5 Quarterly 3 h 180 EUR Stay manual
10 Quarterly 6 h 360 EUR Borderline, your call
10 Monthly 18 h 1,080 EUR A tool pays for itself
15 Monthly 27 h 1,620 EUR A tool pays for itself
15 Weekly 117 h 7,020 EUR Manual is unsustainable

The line to remember: at 10 questions checked monthly, you pass the cost of an entry-level annual subscription by the third month. And this calculation is generous towards manual work, since it counts a single pass per question. Approaching a reliable measurement would mean repeating each question several times, multiplying the time accordingly.

In my view, most people underestimate this figure because they think about query time and forget recording time. Asking the question takes 20 seconds. Opening a clean session, reading the full answer, spotting the competitors cited and filling in a spreadsheet row takes ten times longer.

When manual testing is perfectly sufficient

Three situations where I would advise you not to pay for anything.

You are starting out and looking for a baseline. Five questions, a clean session, an hour of your time. If your brand appears nowhere, you do not have a measurement problem, you have a visibility problem, and no dashboard will fix it. Work on content and third-party citations first.

You are preparing a redesign or a one-off audit. A thorough manual reading paired with a consultant’s judgement produces a better diagnosis than a subscription taken out for three months. Continuous tracking exists to catch drops over time, not to establish an initial diagnosis.

You track one brand in a very narrow market. If three competitors share your category and answers stay stable from one test to the next, a quarterly check is enough to spot a change.

What a tool gives you that manual work never will

Four things, and only one of them truly justifies the spend.

A dated history. This is the decisive argument. An isolated manual reading compares to nothing. Twelve weeks of history show whether your curve is rising, flat or dropping, and let you connect a movement to something you did.

Volume. Querying 15 questions across 3 AI engines every week means 45 observations per round, or 2,340 a year. No reasonable person does that by hand.

Share of voice. Knowing who gets cited instead of you, and how often. Manual work shows it occasionally, a tool quantifies it.

Catching what you were not looking for. A hallucination about your pricing, a third-party source describing you badly, a competitor appearing out of nowhere. You will not find these by testing the questions you already thought of.

Be wary of tools selling you a ranking position

Rand Fishkin closes his study bluntly: any tool that reports a ranking position in AI is talking nonsense. On a single pass, the order of cited brands is close to random, with roughly a 1 in 1,000 chance of seeing the same order twice.

What the same study does validate is appearance frequency measured across a large number of runs. In tight categories, leading brands come out with a presence rate clearly above the rest.

In my view, this distinction should be your first filter when comparing tools, mine included. An aggregated presence metric over time holds up. A position reported for a given query should be read with a lot of caution, whoever sells it to you. One honesty caveat on this study: it is co-signed by the head of an AI tracking product, and its conclusion also serves that positioning. The methodology remains the most solid published on the subject so far.

Frequently asked questions

Is private browsing enough for a neutral test?

It rules out your conversation history and personalisation, which is necessary. It does not neutralise your location, the model version served that day, or whether web search triggers. The test gets cleaner, not reproducible.

What is the difference between a mention and a citation?

A mention means your brand name appears in the text of the answer. A citation means one of your pages is shown as a source, with a link. You can be mentioned without being cited, and cited without being mentioned. A mention influences the decision, a citation brings traffic.

How many times should you repeat a question for a reliable result?

No figure has consensus, and the SparkToro study explicitly leaves the question open. Academic work on model evaluation suggests two repetitions already remove most of the inversions caused by chance, with returns dropping fast after that. The main point is that a single pass is never enough.

Can free tools handle this tracking?

Partly. Google Search Console and Bing Webmaster Tools show traffic arriving from AI engines, meaning the people who clicked. They do not show the answers where you are cited without a click, which is most of the subject. Useful complements, not substitutes.

Should you test on mobile and desktop?

Display gaps exist, especially on enriched formats and carousels. For manual work, pick one device and stick with it. Switching between readings adds a variable you do not need.

How often should you check if tracking is manual?

Once a quarter for most businesses. Your visibility moves at the pace of the content you publish and the citations you earn elsewhere, meaning over 3 to 8 weeks. Checking weekly by hand will make you react to variation that means nothing.

Does a tool guarantee you will be cited by ChatGPT?

No, and no vendor can promise it. A tool measures and prioritises. It does not decide what the model answers. Be wary of any guaranteed citation promise.

Sources

  • Rand Fishkin and Patrick O’Donnell, New Research: AIs are Highly Inconsistent When Recommending Brands or Products, SparkToro, January 2026 — sparktoro.com
  • Berk Atıl et al., Non-Determinism of “Deterministic” LLM System Settings in Hosted Environments, Proceedings of the 5th Workshop on Evaluation and Comparison of NLP Systems, Association for Computational Linguistics, December 2025 — aclanthology.org/2025.eval4nlp-1.12
  • iAdvize and Ifop, study on AI in the online purchase journey, surveyed late January 2026 among 1,051 representative French respondents — reported by Ecommerce Nation, March 2026, ecommerce-nation.fr
  • Miguel Angel Alvarado Gonzalez et al., Do Repetitions Matter? Strengthening Reliability in LLM Evaluations, arXiv:2509.24086, September 2025 — arxiv.org/abs/2509.24086
Florian Zorgnotti

I’m Florian Zorgnotti, an SEO consultant based in Nice since 2016. I’ve led 300+ projects, specializing in WordPress, Shopify, and Generative Engine Optimization (GEO) to help brands grow their visibility in search and AI platforms. Linkedin