ChatGPT can give two people different local recommendations without either person doing anything wrong. Their questions may look similar while the location, wording, conversation context, search results or details of the job point toward a different recommendation list.
That variation matters when you measure whether your business appears. One result can show you what a customer might see. It cannot establish a permanent rank or prove that a change to your website worked.
The same question may not produce the same search
ChatGPT Search does more than pass every sentence unchanged to a search engine. OpenAI says it may rewrite a question into one or more targeted searches. It may also use an approximate location based on the user's IP address, an optional precise device location and relevant saved memories when forming a search.
Sometimes the wording you enter is rephrased before search-style matching happens.
That means two customers who both want a plumber may still be asking different questions. One types "plumber near me." Another asks for an emergency plumber near South Congress who can reach a burst pipe tonight. The second question adds a neighborhood, a service, urgency and availability. A different answer is reasonable if those details change which businesses fit.
The account can matter too. OpenAI says memory can personalize ChatGPT when it is enabled. Temporary Chat does not use or create memories. A test run in a normal signed-in conversation and another in Temporary Chat are therefore not the same measurement conditions, even if the visible prompt matches.
Four searches produced seven businesses
AnswerPrism ran a small check using Austin emergency plumbers. We asked four fresh ChatGPT conversations for three businesses that could handle a burst pipe that night. Two conversations received the exact same question. A third used a shorter version, and a fourth added the South Congress neighborhood.
Across the twelve recommendation slots, ChatGPT named seven distinct businesses. The two identical prompts shared two of three names, but one name changed. The South Congress version shared only one name with the first run and introduced two businesses that had not appeared there.
What's interesting about the answer given below is that different phrasing doesn't change just the results but also the "problem statement" that is being solved.
The test case implies that recommendations can vary even when prompts are very similar. It does not show how often this happens in every city or trade, and we did not call the businesses to confirm that a crew was actually available. The result is useful for measurement: a three-name answer should be treated as a sample, not the local market carved in stone.
A larger preregistered study reached a compatible conclusion in a different setting. Researchers collected 2,208 local food-and-drink responses from four AI systems over seven days. Their two-week comparison found that answer similarity across periods was comparable to same-day reruns. The study covered venues in two Bali markets rather than American service businesses, so its figures should not be applied to plumbers. It does provide stronger evidence that recommendation changes can be ordinary behavior instead of proof that the market changed overnight.
Location and context change the job
OpenAI recommends including a city, neighborhood or postal code when local results show the wrong area. That advice is useful for a customer, but an audit has to record those details because each one can alter the businesses eligible for the job.
Four kinds of context deserve separate treatment:
Live information adds another layer. A search may find changed hours, a recently updated service page, a closed business or a different set of available sources. In addition to the wording and location signals, timing and availability are also a separate, verifiable source of changes.
Not every movement deserves an explanation. If two suitable businesses exchange first and third place while both remain present, the owner has learned less than if the business is repeatedly absent from questions it is demonstrably qualified to answer.
- Place: A city, neighborhood, postal code, approximate IP location or precise device location can narrow the working area.
- Need: Emergency repair, routine maintenance and a major replacement are different jobs even when the trade is the same.
- Constraint: Timing, budget, property type, accessibility and required credentials can remove otherwise plausible choices.
- Conversation: Earlier messages and saved memories can add preferences that are absent from the final sentence.
Measure presence before position
A stable test starts with a fixed set of realistic customer questions. Keep the main questions unchanged during a comparison period. Put experiments with new neighborhoods, jobs or constraints into a separate exploratory set so they do not quietly change the benchmark.
When running your tests, you need to track your results systematically and accurately. Record:
Once you've picked out your set of test questions, run each one several times on different days or even multiple times in a single session. This will let you track whether ChatGPT continues to see your business as a genuine recommendation in response to the question being asked, or whether it's disappeared off the list.
Make sure to save all the evidence you can get, including:
Those notes let you report several observations separately: whether the business appeared, whether it was genuinely offered as a fit, whether the description was accurate and which visible pages supported the answer. The broader measurement guide explains why mentions, citations and recommendations need separate fields.
There is no honest universal number of repetitions. A high-value question in a volatile market may justify more checking than a minor exploratory question. What matters is deciding the collection rule before seeing the result and keeping every valid run selected by that rule.
- The exact wording of each question
- The city/neighborhood each time
- Things like constraints (24/7) or product/model requested
- Whether or not the user is an account holder
- What time you run each test and on what day
- The visible source that ChatGPT calls (e.g., "from Yelp" or "from Wikipedia")
- The reasons it gives for that answer (e.g., "mentions emergency service," "one of the most popular plumbers in the whole state," etc.)
- As well as noting any conflicts or omissions that concern your business.
Fix errors, not ordinary movement
A changed recommendation list becomes actionable when it exposes a factual error or a recurring evidence gap. Check the service, territory, hours, qualifications and customer fit stated in the answer. Then open the visible sources rather than assuming the model invented every detail itself.
If ChatGPT gets a factually incorrect detail about you, then you need to look at the underlying source it uses for that detail and see if that's because you haven't published that information accurately on the relevant source page.
Finding such an error, you need to fix it. It's not enough to change ChatGPT's opinion. To get it to change its mind about you on the public question, you must correct the underlying source. That means changing what's visible there, whether it's Google Business Profile, Yelp, your website, or wherever you hold the information. Then, re-run some of your queries to see if ChatGPT's answers change accordingly.
The source may already be correct. In that case, preserve the answer and check whether the same error recurs before rebuilding a page around it. ChatGPT can make a mistake even when the public record is accurate. The business cannot edit an individual customer's answer or guarantee the next recommendation list.
A persistent absence deserves a different investigation. Compare the businesses repeatedly recommended for the same job. Can their service area, response window, credentials or relevant work be verified more easily? If your company genuinely fits but its public information leaves those questions unanswered, improve the page or profile carrying the gap. This competitor guide shows how to compare the stated reasons without copying another company's claims.
Ignore tiny movements in where you fall in the list. Focus on whether you're present or absent in the list the customer sees. A measurement system becomes useful when it can distinguish a normal shuffle from a repeated, supportable reason that a qualified business is being left out.

Written by Tristan Michel