GEO Knowledge Hub
Why AI Visibility Reports Fluctuate: Signal vs Sampling Noise
Coverage was 85% last time and 62% now — a real decline or normal variance? This post explains the sampling behaviour of AI answers, how to read three metrics together, and how to separate signal from noise.
Anson Ng
6 min read
"Coverage was 85% last time and 62% now — did we get worse?" We get this question almost every month. The usual answer: not necessarily — you are looking at a sampling metric that fluctuates.
How far apart can two reports be?
Take New Win Tec (LEEO): the 07.09.2026 report showed 85.7% coverage, AI mention rank #2.5 and SOV 24.5%. By 15.09.2026 coverage had moved to 62.1% — but rank improved to #1.8 and SOV rose to 36.7% (full report and screenshots).
Read coverage alone and it looks like a decline; read rank and SOV together and the picture changes: the brand appears earlier and takes a bigger share when mentioned — a quality gain.
Why AI answers are unstable
- Generative models are stochastic: the same question can produce different answers each run
- Sampling conditions differ: question set, aliases, engine and scan time all change results
- Competitor sets change: when AI starts comparing more brands, the list inside the answer gets longer
- Engines update: model refreshes change citation behaviour
Read the three metrics together
| Metric | What it measures | How to read it |
|---|---|---|
| Coverage | Probability of being mentioned (breadth) | Fluctuates short-term — track the trend |
| AI mention rank | Where you appear when mentioned (position) | Lower is better |
| SOV | Your share of all brand mentions (concentration) | Shows how much AI focuses on you |
Four ways to separate signal from noise
- Read the trend, not a single day: 3–4 consecutive reports are far more reliable than any snapshot
- Check whether rank/SOV move with it: coverage down while rank and SOV rise usually means better-quality mentions
- Look at individual questions: did everything drop, or only one or two comparison-type questions? The latter is usually an answer-structure issue, not a broad decline
- Watch competitor lines: if rivals rise at the same time, the market changed rather than your performance
Five questions to ask before comparing reports
- Do both reports use the same scope (website, aliases, keywords)?
- How many questions were tested and how many samples were taken?
- Which engine and which surface (AI Overview / AI Mode / app)?
- Are there dated screenshots and reproducible queries?
- Did the competitor set change?
If a provider cannot answer these five, the numbers are hard to trust no matter how good they look.
How we do it
Every GNS-GEO report states its scope, question set, scan date and screenshots; three AI agents (Insight, Strategy, Execution) run in a human-approved loop with weekly monitoring, so you are reading a trend rather than an isolated number.
Want to know where your brand really stands today? Start with the free AI visibility report, then use the five questions above to audit any GEO report you are shown.
FAQ
Why do the numbers differ when I run the scan again a few days later?
AI answers are generated with randomness; coverage is a sampling probability, not a fixed score. The same question can mention you once and skip you the next time, so short-term movement is normal and does not mean your optimization stopped working.
Does lower coverage mean we got worse?
Not necessarily — read rank and SOV alongside it. Real example: for New Win Tec, coverage moved from 85.7% to 62.1% between 07.09 and 15.09, while AI mention rank improved from #2.5 to #1.8 and SOV rose from 24.5% to 36.7%. The quality of mentions improved; breadth was affected by sampling and comparison-type questions.
How should two reports be compared?
Compare like with like: same website, same alias set, same keyword set and same engine. Look at a 3–4 period trend rather than a single day; most day-to-day differences are noise.
How can I tell whether an AI visibility report is trustworthy?
Check whether it lists the question set, sampling method, scan date, screenshots and reproducible queries. A single number with no method, no dates and no screenshots deserves caution — these metrics fluctuate, so an unsupported number says very little.