Why Your AI Visibility Score Swings Every Week: Citation Volatility, Sampling Error and How Many Runs You Actually Need

By Aditya JhaSeptember 27, 202610 min read

Why Your AI Visibility Score Swings Every Week: Citation Volatility, Sampling Error and How Many Runs You Actually Need

A VP of Marketing at a US SaaS company opens the Monday GEO report: AI share of voice is up 11 points. The CEO forwards it to the board. The following Monday it is down 9, nothing on the site changed, and the agency's explanation is 'the models updated.' Both reports were accurate, and both were meaningless. Each was built from one answer per prompt, and AI answers are not a ranking you read off a page. They are a sample drawn from a probability distribution, and a single sample tells you almost nothing about the distribution.

How much do AI answers actually change between runs?

A lot more than most dashboards admit. Parse's July 2026 study analyzed 693,509 answers across 16,143 ChatGPT prompts and 15,805 Google AI Overviews prompts between 26 March and 25 April 2026. Repeat ChatGPT answers to the same prompt shared only 21.2% of their cited domains; AI Overviews shared 31.5%. Put the other way, 78.8% and 68.5% of sources churned from one answer to the next.

The pool behind each prompt is wide. A single ChatGPT prompt drew on an average of 81.3 distinct domains across repeat answers, yet only about 1.4 domains per prompt appeared in 80% or more of them. Parse calls those anchor sources, and they carried just 22.1% of ChatGPT's citation volume (33.2% on AI Overviews). Most of what you see in any single answer is the rotating cast, not the anchors.

Source: Parse, "AI citation volatility by industry" (July 2026), 693,509 answers, 26 Mar to 25 Apr 2026.
Source: Parse, "AI citation volatility by industry" (July 2026), 693,509 answers, 26 Mar to 25 Apr 2026.

Is it the same for brand mentions, not just cited URLs?

Brand mentions are steadier than URLs, but still far from fixed. Glen Allsopp's 28-day study at Detailed tracked more than 70,000 responses across 1,300+ commercial prompts daily. Leading brands appeared on 82% of ChatGPT days and 89% of AI Overviews days, but two random days produced an identical list of brands or cited domains less than 3% of the time. Only 25% of ChatGPT's cited URLs repeated the next day, and the top brand was named first only half the time on ChatGPT.

The practical split is the same one Parse found: a small core that is stable (Allsopp's core brands changed only 13% day to day on ChatGPT) and a long periphery that churns (78%). Your goal in GEO is to move from the periphery into the core, and your measurement has to be able to see that move.

Why are AI answers probabilistic in the first place?

Two layers of randomness stack. First, generation: the model samples tokens from a probability distribution, so wording, ordering and which brands get named vary run to run, as explained in why an LLM gives different answers to the same question. Second, retrieval: an answer engine rewrites your prompt into several sub-queries, pulls a fresh candidate set from its index, and selects a handful to ground on. Small differences in the rewritten queries or the candidate pool change which pages get cited.

Google's AI Mode adds a third layer: personalization, which we covered in how Personal Intelligence breaks rank tracking. The upshot is structural. There is no single 'position' to track. There is an inclusion probability per prompt, per platform, per week, and you can only estimate it by sampling.

How many runs do you need before a number means anything?

Treat each run as a coin flip: your brand is either cited or not. If the true inclusion rate is p and you take n independent runs, the standard error of your estimate is √(p(1−p)/n), and a 95% confidence interval is roughly ±1.96 times that. The arithmetic is unforgiving:

  • **1 run per prompt (p = 0.3):** the result is 0 or 1. It cannot distinguish a 10% brand from a 60% brand.
  • **20 runs of one prompt:** standard error ≈ 0.10, so the interval is about ±20 points. A reading of 30% is compatible with anything from 10% to 50%.
  • **100 runs of one prompt:** ±9 points. Useful for a flagship query, expensive to do for many.
  • **A 50-prompt panel × 10 runs (500 observations):** about ±4 points on the panel-level share, before correcting for the fact that runs of the same prompt are correlated. With that clustering, the effective sample is smaller, so budget for ±6 to ±8 points in practice.

How should a GEO measurement panel be designed?

Build it like a survey, not like a rank tracker:

  • **Prompt panel from buyer language.** 40 to 100 prompts sampled from real sales-call questions, support tickets and comparison queries, stratified by funnel stage and market (US, UK, UAE phrasing differs). Freeze the panel for a quarter so trends are comparable.
  • **Repeated runs, fresh sessions.** 5 to 10 runs per prompt per platform per period, logged out or in clean sessions, spread across days so you average over index refreshes rather than one moment.
  • **Three metrics, not one.** Inclusion rate (share of runs naming or citing you), anchor status (share of prompts where you appear in 80%+ of runs), and share of voice against named competitors. Anchor count is the number that compounds.
  • **Significance before celebration.** Compare periods with a two-proportion test or bootstrap the panel. If the week-over-week change sits inside the interval, the report should say 'no detectable change', not 'up 11 points'.
  • **Tie it to money.** Pair the panel with AI-referral sessions and pipeline in analytics, the method in how to track AI citations. Share of voice alone doesn't pay a retainer.

What actually moves a brand from rotating source to anchor?

Anchors are the pages the retrieval layer keeps choosing because they answer the sub-query cleanly and are corroborated elsewhere. That points to three levers: answer-first pages that match the exact sub-questions buyers ask, third-party corroboration (reviews, analyst mentions, earned media) so the engine sees consensus, and freshness on the pages that matter, covered in how often to update content to stay cited. Parse also found churn highest in commerce, software, AI and IT (about 81% of sources changing), so SaaS brands should expect noisier numbers and need larger samples than a travel or finance brand.

How AIBOOTSTRAPPER helps

Our GEO reporting is built on repeated-run prompt panels with confidence intervals, because a board deck built on one-shot screenshots eventually embarrasses someone. The outcomes we point to are commercial ones: one B2B SaaS founder we work with, Rahul Mehta, told us that within 90 days they were being cited by ChatGPT for their category and inbound doubled. For PropLock, the UK PropTech platform, the GEO-optimized site we built reached 12k organic visitors a month within 90 days alongside a 47% lift in qualified viewings (see case studies).

If your current AI visibility report can't tell you its margin of error, book a call and we'll audit the methodology, or see our AI marketing and GEO services.

Want this done for you?

Book a free strategy call and we'll show you how to build and market your business with AI.

FAQ

Questions, answered

Everything you might want to know before we hop on a call.

Because AI answers are sampled, not ranked. Parse found repeat ChatGPT answers share only about 21% of cited sources, so a report built from one answer per prompt mostly measures random churn rather than real change.

At least 5 to 10 runs per prompt per platform per period, across a panel of 40 to 100 prompts. One prompt run 20 times still carries a margin of error of roughly ±20 points; a pooled 500-observation panel gets closer to ±4 to ±8 points.

A domain that appears in at least 80% of repeat answers to the same prompt. Parse found only about 1.4 anchors per ChatGPT prompt, and becoming one is the most durable GEO outcome.

Somewhat. In Parse's study, repeat AI Overviews answers shared 31.5% of sources versus 21.2% for ChatGPT, but both churn heavily enough that single checks are unreliable.

Keep reading

Let's talk

Ready to build and sell with AI?

Book a free 30 minute strategy call. We'll map the highest ROI AI move for your business, no pitch, just value.