An AI visibility score can be useful as a compact summary of a fixed test, but it is not a universal grade for how visible a business is in ChatGPT or other AI tools. The owner should trust it only as far as they can inspect the questions, products, locations, collection dates, repetitions, counting rules, and raw answers behind it.

That distinction matters when a neat percentage arrives in an agency report. A score may summarize careful work. It may also hide a weak question set, a convenient denominator, or a mixture of tests that have little to do with the jobs your business wants. The number alone cannot tell you which.

Start with what the score counts

There is no standard unit called AI visibility. One platform can count whether a business appears anywhere in an answer. Another can count citations, positions, estimated exposure, or a share of all brand mentions in its own prompt set. Those measures answer different questions.

Semrush defines its 0–100 AI Visibility score using topic coverage and mention consistency. Ahrefs reports several measures separately, including mentions, citations, share of voice, and estimated impressions. Ahrefs also explains that estimated impressions use Google search demand to model potential exposure. They are not a count of people who saw an AI answer.

Neither method is inherently useless. The problem begins when a report drops the definition and presents the result as though 62% meant the same thing everywhere. It does not. Before comparing two tools, two agencies, or two months, check that the numerator, denominator, prompt set, and collection method still match.

One small dataset can produce four honest percentages

AnswerPrism ran four fresh ChatGPT conversations that each asked for three emergency plumbers in Austin. ABC Home & Commercial appeared in three of the four answers, and that business occupied three of the 12 recommendation slots.

That means that the evidence supports a 75% rate of a prompt containing that business name, and a 25% share of recommendation slots. Both describe the same four answers. One divides appearances by conversations; the other divides occupied positions by all positions.

The test also contained two runs that used identical wording. If a report showed only that narrower pair, ABC appeared in both conversations and held two of the six available slots. The resulting figures would be 100% prompt presence and 33% slot share.

So the same recorded evidence can support 25%, 33%, 75%, or 100%. The arithmetic is not the disagreement. The definition and sample selection are. None of those figures means ABC owns that share of ChatGPT, local plumbing demand, or customer attention. This was one bounded test, not a market benchmark.

The example is why a percentage should always travel with a plain-language label. “Present in three of four tested answers” is harder to misunderstand than “75% AI visibility.”

The questions choose the score before the math does

A report can look rigorous while testing the wrong work. A roofer may care about questions from homeowners comparing replacement options, looking for storm-damage help, or asking about a flat commercial roof. Fifty broad prompts about roofing history and materials can make the sample larger without making it more useful.

Branded and discovery questions also belong in separate groups. “What does Smith Heating charge?” already gives the system the company name. “Who repairs a heat pump near me today?” asks the system to discover a suitable business. Success on the first question says something about how the company is described. It does not show that the company would be found for the second. A local business owner should insist the report split them.

Location and wording can change the answer too. OpenAI uses approximate or precise location, rewrites a question into targeted searches, and uses relevant memories as set out in ChatGPT Search input conditions; any change in these conditions can affect the answer, even for the same business.

Recent research gives that caution some scale. A study of commercial recommendations found that natural paraphrases of the same buying intent produced only 14% to 29% overlap in recommended brands, while exact reruns overlapped more often. A separate paper on uncertainty in AI visibility measurement argues that a single run can look much more precise than repeated sampling supports. Those studies are not local-business benchmarks, but they show why one favorite prompt and one answer are too fragile to carry a monthly score.

Ask to see what sits underneath the number

A useful report should let you reconstruct the result without trusting a mystery formula. Ask for enough detail to answer these questions:

Keep the raw answers so you can inspect the underlying evidence, including what might be missing. Compare like-for-like collection windows, and do not react to one-day movement. If the provider changes its prompt set or formula, mark the break instead of drawing a smooth trend line across two different tests.

The same caution applies to platform comparisons. Semrush says it keeps platform results separate because combining them can hide useful differences. Ahrefs describes its figures as directional modeled signals, not audience measurement or business performance. A ChatGPT score and a Google AI Overview score should not be blended unless the report also shows the separate results.

  • What does the number count: mentions, citations, recommendations, position, sentiment, or a mixture?
  • Which exact customer questions were tested, and which were branded?
  • Which AI product and mode answered them?
  • What location did the test state or assume?
  • When were the answers collected, and how many times was each question run?
  • Which competitors were included in any share-of-voice calculation?
  • Can you inspect the raw answers and cited sources?

Make the score point to an inspectable problem

A lower number is not a diagnosis. Open the answers and determine what actually happened. The business may be:

Each of those situations points toward a different type of work, and the owner needs to know which situation applies. An incorrect address calls for a fact correction. A weak service description may call for a clearer page or profile. A missing citation may justify inspecting the sources the answer did use. A report full of irrelevant questions needs a better test, not more content.

Measure mentions, citations, recommendations, and traffic separately. Then use repeat runs to see whether a recommendation is stable or just one possible answer. Calls and qualified leads remain business outcomes; an AI score is only supporting evidence.

A useful score leads to a measurement decision: a page, a profile, a source, or a collection approach. If the next move stays vague, ask for the raw answers before renewing the report. A single number that cannot help make a decision about where to put the checkbook is not yet a useful visibility score.

  • absent from a question it genuinely fits;
  • present but described with an old address, service, or hours;
  • mentioned without a useful supporting source;
  • recommended for a broad topic but not for the work it wants; or
  • visible in answers that send no qualified visits, calls, or leads.

Sources

  1. OpenAI: Searching the web with ChatGPT
  2. Semrush: AI Visibility data and methodology
  3. Semrush: AI Visibility Index methodology
  4. Ahrefs: Brand Radar methodology
  5. Ahrefs: AI visibility metrics
  6. Quantifying Uncertainty in AI Visibility
  7. Paraphrase Brittleness in AI Recommendations