AI answers move around. A useful measurement system expects that. It repeats a small set of buying questions under recorded conditions, keeps the full evidence, and reports whether the business was mentioned, supported, recommended, and described correctly. The purpose is to make a better decision, not produce one impressive percentage.
Start with a business decision
Decide what the benchmark should help you decide. A clinic may want to know whether it enters the consideration set for a specific treatment in Fredericton. A roofing company may care about commercial flat-roof work, not every mention of roof repair. That boundary keeps the test tied to revenue and service fit.
A question such as ‘What is our company?’ can confirm name recognition, but it says little about a buying decision. Build the set around the services, markets, customers, and constraints that matter. Keep the wording stable after the baseline. Explore new questions separately so a changing test does not masquerade as improving performance.
- Discovery: which providers solve this problem?
- Specific need: who handles this type of job or customer?
- Comparison: how do the credible options differ?
- Trust: what proof matters before choosing?
- Fit: which provider suits the location and constraints?
Save enough context to repeat the test
An answer without its conditions is hard to compare. Record the exact question, product and mode, date, market or location, and whether web search was active or visible. Save the whole answer and every source URL the interface exposes.
Keep products separate. A ChatGPT Search result and a Google AI Mode result are two observations, not two votes in a universal ranking. If a product changes its interface, model, or search behaviour, note the change. The measuring instrument moved, even if the business did not.
- Exact question and its category
- Product and mode when visible
- Date, language, location, and signed-in state
- Businesses named and how each one is presented
- Source URLs and a complete capture of the answer
Keep five outcomes separate
A citation is not a recommendation, and a mention can be unhelpful. Imagine an HVAC company cited for an explanation of heat pumps while three competitors receive the actual recommendations. The page has source visibility, but the business has not entered the list of options. Or the company may be recommended with the wrong service area, which is visible but inaccurate.
AnswerPrism keeps the prompt-level evidence and records each outcome separately. A summary can show direction, but anyone reading it should be able to return to the answer and see what the number means.
- Retrieved: the page appears to have been surfaced or used
- Cited: a visible source link points to the page
- Mentioned: the business name appears
- Recommended: the business is offered as a suitable option
- Accurate: important details match verified business information
Compare collection windows, not favourite answers
Generated answers can vary between runs. Repeating a question until the preferred answer appears is selection, not measurement. Decide in advance how many times and over what period you will run the set. Preserve every valid result selected by that rule.
Compare a later window with the starting snapshot. Report which outcomes improved, stayed flat, or became worse. If one question improves in one product, keep the finding at that scale. A broader claim needs a broader pattern.
For a small local business, monthly or campaign-based collection is often more useful than daily checking. The right cadence depends on how quickly the underlying pages and profiles change and how much business value the question carries.
Let each analytics tool answer its own question
Google announced a dedicated Generative AI performance report in Search Console in June 2026 and is rolling it out to a subset of sites. Where available, it can show visibility in Google's generative search features by pages, countries, devices, and dates. It says nothing about your position in ChatGPT or Claude.
Site analytics answers another question: who arrived and what they did. Track useful visits and enquiries from known AI referrers while respecting consent choices. Search Console measures activity within Google Search; site analytics measures visits to your site. Their totals will differ because they observe different events.
Neither tool can tell you why an AI product recommended a competitor. Pair traffic data with the saved answers when you need to understand the story customers saw before the click.
Treat crawler access as a pass or fail check
OpenAI documents OAI-SearchBot for ChatGPT search inclusion, and Perplexity documents its own search crawler. Allowing a relevant crawler removes a barrier. It does not guarantee visibility or a place in the answer.
Check robots rules, canonical URLs, indexing, internal links, and visible page text before blaming content or reputation. Once those basics pass, do not count them as monthly visibility gains. A door being unlocked is not the same as a customer walking through it.
Report the change without inventing a cause
A later improvement may follow a technical repair, clearer service information, a corrected profile, new independent coverage, a platform update, or several changes at once. Unless the test isolates one factor, say that the result changed after the work. Do not claim that a single tactic caused it.
Record the unknowns too. Private conversations, undisclosed training data, hidden selection systems, personalization, and sources the interface does not show are outside the audit. This honesty does not weaken the report. It tells the owner which decisions the evidence can safely support.
End with an action. Continue when a useful pattern is improving. Change direction when the same supported gap remains. Stop when the answer is accurate, the business is a fair fit, and further work would only chase fluctuations.

Written by Tristan Michel