Here is an experience almost everyone tracking AI visibility has had. On Monday, you ask ChatGPT for the best tools in your category and your brand is second on the list. On Tuesday, you ask the exact same question and it is gone.
Nothing changed on your site. No competitor launched anything. The answer simply moved.
This variability is not a bug in the tools or a sign that tracking is pointless. It is a basic property of how AI assistants generate answers. Understanding it is the difference between reacting to noise and reading real trends.
There are several independent sources of variation, and they stack on top of each other. It helps to separate them, because each one calls for a different response.
Language models generate text one piece at a time, choosing each next word from a range of likely options. Consumer assistants usually allow some randomness in that choice so answers feel natural rather than robotic.
For a recommendation question with several reasonable answers, that randomness can change which brands are named and in what order. The brands at the edge of the model's confidence are the ones that come and go.
Many assistants decide on the fly whether a question needs a live search. OpenAI describes how ChatGPT chooses when to search in its ChatGPT search documentation. An answer built from live results can look very different from one built from training data alone, even for the same prompt.
We explained how retrieval shapes citations in the citation funnel. The short version: when search is involved, what ranks today matters.
For answers that do use search, the underlying results shift constantly. A new roundup gets published, an old one drops in the rankings, a review site reorganises its category pages. Each change can alter which sources the assistant reads, and therefore which brands it names.
AI companies update their models and assistant behaviour frequently, sometimes without announcement. A new model version can have different knowledge, different preferences for sources, and different habits in how many brands it lists.
These updates tend to cause step changes: visibility jumps or drops across many prompts at once, and stays at the new level.
Some assistants use memory, location or earlier messages in a conversation to shape answers. Two people asking the same question can get different results. Clean, repeatable tracking removes as much of this as possible by asking prompts in fresh sessions without personal context.
Not every assistant moves in the same way. As a rule of thumb, how much an assistant moves follows how it sources its answers.
| Assistant | Main driver of change | What it looks like |
|---|---|---|
| Perplexity | Live search results on nearly every query | Frequent small movements as the web changes |
| ChatGPT | Mix of training data and on-demand search | Noticeable swings depending on whether search was used |
| Gemini | Google's index and knowledge graph | Tends to track changes in Google visibility |
| Claude | Training data by default | More stable day to day, with shifts at model updates |
| Copilot | Bing search results | Follows Bing rankings and freshness |
This is also why Google's AI Overviews and standalone chat assistants can disagree about the same brand, a topic we covered in AI Overviews vs AI chat. Google publishes guidance on how its own AI features work with sites in its documentation for site owners.
If single answers are unreliable, the unit of measurement has to change. Stop asking "were we in the answer?" and start asking "how often are we in the answer?"
A mention rate is the share of runs in which your brand appears. Run the same prompt set on a regular schedule and your rate becomes a stable number, even when individual answers bounce around.
A brand mentioned in 70 percent of runs for a prompt is a default answer. One mentioned in 15 percent is on the edge. Both would look identical in a single lucky screenshot.
Frequent runs give you enough data points to separate signal from noise. Reporting over a longer window then smooths out the daily jitter. This combination is the reason we built Bold GEO around daily refreshes rather than occasional spot checks.
Variation in answers is hard enough to reason about without adding variation in questions. Freeze a core set of prompts and change it only deliberately. Our guide to building a prompt set covers how to structure a stable core with a small rotating exploratory set.
Presence alone is a blunt signal. Where you appear in the list and how much of the total conversation you own tell you more. See tracking share of voice in AI answers for how to calculate both.
When your numbers move, ask three questions before you react.
Keep a simple change log alongside your tracking: your own content releases, notable competitor moves, and known model updates. Patterns become obvious much faster when you can line them up against the data.
Imagine two brands competing for the same prompt, "what is the best invoicing tool for freelancers". Over a month of daily runs, Brand A appears in 21 of 30 answers, a 70 percent mention rate. Brand B appears in 6 of 30, a 20 percent mention rate.
Now imagine each brand's marketer checks the prompt once, by hand, on a random day. There is a 30 percent chance that Brand A's marketer sees no mention and concludes something is badly wrong. There is a 20 percent chance that Brand B's marketer sees their brand and concludes they are doing fine.
Both conclusions would be wrong, and both would be perfectly reasonable readings of a single answer. This is why spot checks produce so much anxiety and so many bad decisions. The underlying positions are very different; the individual answers are not reliable enough to show it.
The same logic applies to position. A brand that averages second place will sometimes appear fifth and sometimes first. Only the average over many runs tells you where it really stands.
Sooner or later, a founder or executive will ask ChatGPT a question on their phone and forward a screenshot that contradicts your report. How you respond shapes whether they trust the data.
Setting this expectation early saves many uncomfortable conversations later.
There is a useful pattern hiding in all this variation. Some brands appear consistently, run after run, across assistants. Others flicker in and out.
The consistent brands are rarely lucky. They are the ones models have the most evidence for: clear positioning, consistent descriptions across the web, and plenty of independent sources recommending them for specific use cases. When many paths through an answer lead to the same brand, randomness has less room to leave it out.
That gives you a practical goal. You cannot remove the randomness, but you can move your brand from the edge of the model's confidence toward the centre. The work that does that, strengthening third-party evidence and entity clarity, is the same work that grows visibility in the first place.
Variability also changes how you should judge individual pieces of work. If you publish a new comparison page and your brand appears in the next answer you check, that is not proof the page worked. If it does not appear, that is not proof it failed.
Give changes time, compare mention rates before and after over several weeks, and look at whether the new page starts appearing in the cited sources. That last signal is the most direct evidence you will get that a specific piece of content is doing its job.
Variability is the part of AI visibility clients find most confusing. A screenshot from a client's own search that contradicts your report can undermine trust quickly.
Explain variability up front, before the first report. Show mention rates and trends rather than individual answers, and include a short note on what drives day-to-day movement. We cover reporting formats in more depth in GEO for agencies.
AI answers change because of sampling randomness, on-demand search, a shifting web, model updates and personalisation. That makes single answers unreliable.
Measure mention rates over many runs, keep your prompts fixed, look at position and share of voice, and investigate only sustained, broad or citation-linked changes. Do that and the noise turns into a trend you can act on.
Not by itself. Variation is normal. What matters is your mention rate across many runs and whether it is rising or falling over time. A brand named in 60 percent of runs is in a much stronger position than one named in 10 percent, even though neither appears every time.
There is no single magic number, but one run is never enough. Running the same prompt set on a regular schedule, ideally daily, and looking at rates over a week or a month gives you a far more reliable picture than any individual answer.
You cannot control the randomness in the models, but you can make your brand a more obvious answer. Brands with clear positioning, consistent descriptions and strong third-party evidence tend to appear more consistently, because more of the paths through an answer lead to them.
Bold GEO monitors how your brand is cited across ChatGPT, Perplexity, Gemini, Claude, and Copilot on a daily refresh. 7-day free trial, no credit card.