Why AI answers change between runs (and why one check misleads)
AI answers change between runs because large language models sample each response from a probability distribution instead of looking up a stored record, so the same prompt can legitimately return a different brand list on consecutive tries. Retrieval, index freshness, and provider-side updates add more variance on top of the sampling itself. That makes any single check one draw from a distribution: an anecdote, not a measurement, and a shaky basis for decisions.
Why does the same prompt return different answers?
When you ask ChatGPT™, Claude®, Gemini™, or Perplexity™ to recommend vendors in your category, nothing in the system fetches a fixed answer. The model writes the response token by token, and several independent mechanisms add randomness along the way:
- Sampling temperature. At each step the model holds a ranked set of plausible next words and, at any temperature above zero, does not always pick the top one. Two runs can diverge at a single early token ("Top options include..." versus "Many teams start with...") and never reconverge. Brands sitting near the model's confidence threshold blink in and out between runs.
- Retrieval variance. Search-grounded assistants run a web search before answering, and that search is not deterministic either. The assistant may rewrite your prompt into different queries, pull a different mix of pages, or run a different number of searches per attempt. Different sources in, different brands out.
- Index freshness. Each provider works from its own crawl and cache of the web, refreshed on its own schedule. A page published last week may be visible to one provider and invisible to another, or present in the morning run and absent by evening as caches roll over.
- Grounding differences. Some answers come mostly from the model's trained knowledge, others mostly from retrieved pages, and the blend shifts by provider, by prompt, and by run. A brand that is strong in training data but weak in live search results (or the reverse) will look inconsistent for that reason alone.
On top of all this, providers ship model updates and quietly test variants, so the system you query on Monday may not be exactly the system you query on Friday.
How much do answers actually vary?
Enough to change conclusions. Across a sample of our own scans covering 58 brands between April and August 2026 (33,000+ responses from the web-search-enabled modes of the four major AI search providers, each scan re-asking its own prompts with identical settings), the churn is easy to state: in scans that ran each category prompt three times, only 45.7% of the brands named for a given prompt and provider were named in all three runs, and only 21.6% of prompt-and-provider pairs returned the identical brand set three times. A single run surfaced just 78.5% of the brands that three runs surfaced. One check misses roughly a fifth of what the model is willing to say about your category.
Those are corpus-wide rates, and the spread around them is wide. In the steadiest tenth of scans, about seven in ten prompt-and-provider pairs held an identical brand set across runs; in the most volatile tenth, about one in ten did. Our own earliest published test (406 responses for one small brand in a crowded generic category) sat at 6%, at the volatile extreme. The smaller the brand and the more crowded the category, the harder the answers churn.
Web grounding is where most of the movement comes from. As a baseline we compared a legacy model endpoint with no web search: it returned identical brand sets on 79% of repeat runs, while the search-grounded providers ranged from 37% to 53%. Retrieval reshuffles the evidence underneath the answer on every attempt, and the brand list moves with it.
| Signal | What it measures | What we observed across runs |
|---|---|---|
| Mention rate | How often the brand appears in answers at all | Most stable: rates settle after a few repeated runs |
| List position | Where the brand ranks when an answer lists options | Noisiest: the order reshuffles even when the same brands appear |
The practical reading: whether you get mentioned is fairly reproducible once you average a few runs, but where you rank inside a list is not. If a dashboard shows your average position moving week over week with no spread attached, some of that movement is sampling noise. The noise is not small, either: in our corpus, one scan-provider measurement in ten swung by 14 percentage points or more of mention rate between same-day runs of the identical prompts.
Why is a single check misleading?
Because one response is a sample of one. If your brand appears third in today's ChatGPT™ answer and is missing from tomorrow's, neither response is the truth. The truth is a rate: how often you are mentioned across repeated runs and providers, the same mention-rate math that underpins AI share of voice.
Single-check thinking produces predictable failure modes:
- False alarms. A brand "disappears" in one run, someone escalates, and it is back in the next run. Nothing changed except the dice.
- False comfort. One flattering screenshot circulates internally as proof the brand owns a prompt it only sometimes wins.
- Phantom trends. Comparing one run this month against one run last month reads sampling noise as gain or loss, and teams end up optimizing against movement that was never real.
How do you measure AI visibility honestly?
Four practices separate measurement from anecdote:
- Run every prompt multiple times per provider. Repetition converts "what it said once" into "how often it says it." Three runs per prompt is a practical floor, and the corpus above shows why the second and third runs do real work: one run surfaced 78.5% of the brands three runs found, and a second run lifted that to 92.1%.
- Report the spread, not just the mean. The same headline score with a tight run-to-run spread and with a wide one are two different facts. Without the spread you cannot tell signal from noise.
- Watch run-to-run agreement. If providers mention you consistently across runs, trends in that metric are trustworthy. If agreement is low, treat any movement as directional until more runs accumulate.
- Handle failures cleanly. Timeouts, refusals, and errors happen. A failed response should be excluded from both the numerator and the denominator of any rate; counting it as "not mentioned" quietly deflates your numbers.
And when you trend over time, hold the prompt set fixed so each period measures the same thing. See how to track the prompts that surface your brand.
What should you ask an AI visibility vendor?
If data accuracy and prompt coverage matter most to you, spec sheets will not tell you what you need. These questions will:
- How many runs back each number on this dashboard? If the answer is one, every number is a single sample.
- Can I see the spread or variance behind each headline metric, per prompt and per provider?
- Are failed responses excluded from both the numerator and the denominator of every rate?
- How many prompts does a scan cover, and can I control which ones, so coverage matches how my buyers actually ask?
- Are identical prompts re-asked on a schedule, so trend lines compare like with like?
The category has several established tools that monitor prompts over time, including the Semrush AI Visibility Toolkit, Profound, and Otterly.AI, among others. They differ most in exactly the areas above, how many repetitions back each number and how openly variance is shown, so put these questions to any vendor on your shortlist. For our part, Gen3 AI Visibility runs every scan multiple times, scores each run, and reports the per-run spread alongside every headline number, with a consistency view showing how much providers agreed across runs. That is not a claim of less variance (the variance belongs to the AI systems); it is a commitment to showing it.
Measure it
Want to see how your own brand holds up across repeated runs, not just one lucky answer? Run a free Pulse visibility check from our homepage for a first read, or explore the rest of the guides at /learn/ to go deeper on methodology.
Frequently asked questions
- Why does ChatGPT™ give different answers to the same question?
- ChatGPT™ generates each answer token by token by sampling from a probability distribution, so identical prompts can produce different wording and different brand lists. Web-grounded answers add retrieval variance on top, since each attempt may search differently and pull different sources. This is normal behavior, not a bug.
- Is there a way to monitor changes in AI answers for the same prompt over time?
- Yes. AI visibility platforms re-ask a fixed set of prompts on a schedule across providers like ChatGPT™, Claude®, Gemini™, and Perplexity™, then track mention rates and positions over time. Because single answers vary, look for a tool that runs each prompt multiple times and reports the spread, so real changes stand out from sampling noise.
- How many times should you run a prompt to measure AI visibility?
- More than once, because a single response is one sample from a random process. Three runs per prompt per provider is a practical floor: in our own multi-scan measurements, a single run surfaced only about four fifths of the brands three runs surfaced. Stable signals like mention rate settle quickly, while ranked position needs more repetition before you can trust movement in it.
- Which AI visibility metric is most stable between runs?
- Mention rate, meaning how often a brand appears in answers at all, is the most stable signal between runs. Position in ranked lists is the noisiest: the same brand can lead one run's list and sit mid-pack the next, so trends in ranked position need more repetition before they can be trusted.
- Why is checking an AI answer once misleading?
- One answer is a single sample from a random process, so it can show your brand missing on a day you are usually mentioned, or ranked first on a prompt you rarely win. Reliable measurement comes from repeated runs of the same prompt, reported as a rate with a spread, not from a one-off screenshot.