AI Discoverability

How to measure AI discoverability without fooling yourself

One answer in one app is a sample of one, and a sample of one can say whatever you hoped it would.

Say an estate planning attorney in Thousand Oaks asks ChatGPT for the best estate planning attorney in town, sees her firm in the answer, and relaxes. Her office manager asks the same thing from home that evening and the firm is not there. Neither has measured anything. Each drew one card from a shuffled deck.

This page is about turning those single draws into a number you can trust, so that when you change something to improve your AI discoverability you can tell whether it worked.

Why the same question gets different answers

The variation is not a glitch, it is how these systems are built. OpenAI's help page on searching the web with ChatGPT explains that ChatGPT search "typically rewrites your query into one or more targeted queries" for its search providers, that it may use an approximate location from your IP address, and that saved memories may influence the rewritten query. It adds that a VPN or network location "may affect the approximate location." Google says its AI Overviews and AI Mode "may use different models and techniques" and may fan a question out into multiple related searches (Google Search Central; see query fan-out).

We see the same thing in our own sampling. When we asked one assistant the same buying question eight times in each of four local categories, 37% of the businesses it named appeared in only one of the eight answers, as published in our share of voice sample dated 2026-09-12.

What to count

For each answer, record four things and nothing fancier:

  • Named or not. A plain yes or no. Position in a short list moves too much to mean anything.
  • Which sources were cited. The domains in the citations or source panel. This tells you why you were or were not named, and our page on the sites AI reads about you explains what to do with that list.
  • Whether the facts were right. Being named with the wrong phone number or a service you stopped offering is a finding, not a win.
  • The date and the conditions. Which assistant, signed in or not, web search on or off, which city in the question.

Your headline number is simple: answers that named you, divided by answers collected, for each question.

How to ask

  • Fix the questions. Write five to ten in a customer's words and never edit them. "Best estate planning attorney in Thousand Oaks" and "estate lawyer Thousand Oaks" are different questions and cannot be compared.
  • Put the city in the question. That removes most of the guesswork about where the assistant thinks you are.
  • Keep your own name out. "Is Smith Law good?" tests whether the assistant knows you exist. "Who should I call?" tests whether it recommends you. Only the second is the job.
  • Hold conditions steady. Same assistant, same account state, web search on, new conversation each time.

How many repeats before a change means anything

Here the arithmetic is unforgiving. Suppose an assistant truly names your firm in half of its answers. If you ask ten times, plain coin flip odds say you will see anywhere from three to seven mentions about 89% of the time. So a score of four out of ten this month and six out of ten next month can be the same firm with nothing changed.

A working rule

Ask each question at least ten times per round. Run a baseline round before you change anything. Make the change, wait long enough for the pages involved to be fetched again, then run a second round the same way. Only call it progress if the improvement shows up across most of your questions and holds in a third round a few weeks later. One question moving by one or two mentions is weather, not climate.

This is slow, and that is the point: decisions made from screenshots credit the wrong change or abandon the right one.

What a measurement cannot tell you

A share of answers is one view of AI discoverability: it describes what one assistant did, under your conditions, in one period. It does not predict what a customer in another neighborhood sees, and it cannot prove why a change happened. For the on-site half, which does not vary between runs, run the free AIOInsights check: it reads the same public signals the same way every time, so its result is a fixed baseline rather than a sample. The first week plan shows where measurement fits alongside the fixes.

Questions

Why does ChatGPT give a different list of businesses each time I ask?

ChatGPT gives different lists of businesses between runs because each answer is assembled fresh. OpenAI says ChatGPT search typically rewrites a question into one or more targeted queries sent to search providers, may use an approximate location from the IP address, and may draw on saved memories. Different queries, places and retrieved pages produce different names, so a single answer is one sample, not a ranking.

How many times should I ask an AI assistant before trusting the result?

To measure whether an AI assistant names a business, ask the same fixed question at least ten times under the same conditions and record the share of answers that name it. Even then, a business named in half of all answers will often score anywhere from three to seven out of ten by chance alone, so treat small moves as noise and look for a gap that holds across two separate rounds.