By Jarno Saarimies, Founder of AEOvara. This is Part 3 of the series How to Build a Repeatable AI Search Visibility Benchmark.
Ask ChatGPT the same commercial question twice and you may get two different shortlists. Change “best” to “most reliable,” and the shortlist may change again. Run the prompt in a continuing conversation, in another country, or after a model update, and you may no longer be measuring the same system.
That creates an awkward problem for AI visibility reporting: a dashboard can show a precise score even when the experiment behind it is not precise.
This does not make AI visibility measurement useless. It means the score must be treated as an estimate produced by a documented test—not as an objective ranking retrieved from a fixed index.
In Parts 1 and 2 of this series, I explained how to build a prompt corpus and track mentions, citations, competitors, and answer quality. This article addresses the harder question:
When is a change in AI visibility evidence, and when is it only measurement noise?
The answer requires more than repeating prompts. You need to control the test environment, preserve the raw evidence, quantify uncertainty, and separate effects that dashboards often collapse into one number.
First, identify what actually changed
Suppose your mention rate rises from 30% to 40%. At least three different things could have produced that result.
1. The brand changed
You published new evidence, earned independent coverage, corrected an entity description, improved a relevant page, or became more widely discussed. This is the effect an AEO or GEO programme is trying to measure.
2. The engine changed
The provider released a new model snapshot, changed retrieval behaviour, refreshed an index, altered citation rules, or modified the product. OpenAI explicitly states that model outputs are variable and that prompting behaviour can change between model snapshots. For API applications, it recommends pinned model versions and evals when consistency matters.
Consumer AI products rarely give researchers complete control over those variables. A visible product name such as “ChatGPT” is not a permanent experimental condition.
3. The sample changed
The wording, session history, location, language, login state, retrieval mode, time, or number of repetitions changed. Even if the website and model stayed untouched, the measured score could move.
If your log cannot distinguish these explanations, it cannot support a causal claim such as “our optimization increased AI visibility by 10%.” It can only say that the observed answers changed.
That distinction sounds cautious. It is also what makes the result credible.
A prompt is not a keyword
Traditional rank tracking usually observes a defined query, location, device, and search engine. AI prompts carry more semantic degrees of freedom.
Consider these questions:
- What are the best technical SEO consultancies for a Finnish SME?
- Which technical SEO consultancy is most reliable for a small Finnish company?
- Who can help a Finnish business become visible in AI-generated answers?
- Recommend an affordable Finnish specialist in SEO and AI search optimization.
They occupy the same commercial neighbourhood, but they are not interchangeable samples. “Best,” “reliable,” “affordable,” and “AI search optimization” introduce different selection criteria.
A 2026 controlled study by Jan Ehrlinspiel, Malte Landwehr, and Tomek Rudzki examined 37,804 responses across five AI answer engines. It found that prompt format and archetype can shift visibility baselines, with the clearest sensitivity appearing in unbranded, middle-funnel commercial prompts. The authors are affiliated with Peec AI, so that commercial relationship should be disclosed, but the methodological lesson is useful: one wording cannot represent an entire buyer intent.
Freezing a prompt list is therefore necessary, but not sufficient. A frozen list gives you repeatability for those exact prompts. It does not prove that the list represents the full market.
Use prompt families, not a bag of convenient questions
A stronger benchmark separates two layers:
- Core prompts remain unchanged so trends can be compared over time.
- Prompt variants test whether the result survives realistic changes in wording.
For each buyer intent, write one core prompt and two or three variants before checking which version favours your brand. This prevents an easy form of unconscious cherry-picking: keeping the wording that produces the most flattering result.
| Buyer intent | Core prompt | Controlled variants |
| Category discovery | What are the best AEO agencies in Finland? | most reliable; established; suitable for SMEs |
| Problem solving | How can a Finnish SME improve visibility in AI answers? | get cited by ChatGPT; appear in AI search; improve generative search visibility |
| Supplier selection | Who should I hire for an AI visibility audit in Finland? | independent specialist; consultancy; local expert |
Do not average branded questions such as “What is AEOvara?” into the same score as unbranded discovery questions. A branded prompt measures recognition and description accuracy. An unbranded prompt measures discovery. Combining them can make a weak discovery result look healthy.
Record the test environment like an experiment
Every observation should include enough metadata for another person to understand what was tested:
- exact prompt text;
- platform and visible product mode;
- model identifier, when exposed;
- web search or grounding status, when exposed;
- date, time, language, and country;
- signed-in or signed-out state;
- clean session or continuing conversation;
- personalization or memory state, when known;
- repetition number;
- full answer and cited URLs;
- scorer name or scoring-rule version.
“ChatGPT, September” is not adequate provenance. Neither is a cropped screenshot that hides the prompt, date, citations, and product mode.
API-based monitoring improves control, but it creates a different test. An API model with a pinned snapshot and explicit search configuration is not necessarily equivalent to the consumer interface used by customers. Report API observations as API observations. Do not quietly label them “what users see in ChatGPT.”
Measure a rate—and show its uncertainty
If a brand appears in 3 of 5 runs, the observed mention rate is 60%. That does not mean the true probability of a mention is known to be 60%.
With only five binary observations, a 95% Wilson confidence interval is approximately 23% to 88%. The interval is wide because the sample is small.
If the brand appears in 30 of 50 comparable runs, the observed rate is still 60%, but the corresponding interval narrows to approximately 46% to 72%. More observations have not changed the headline percentage; they have changed how much confidence we should place in it.
For a binary mention metric:
mention rate = mentions / comparable runs
Report the numerator and denominator—not just the percentage:
ChatGPT web search: 12/20 mentions, 60%
95% Wilson interval: approximately 39%–78%
The interval is not decoration. It helps prevent claims that the evidence cannot support.
If period A is 6/20 and period B is 8/20, the dashboard has moved from 30% to 40%. But with samples that small, the difference may be ordinary variation. Call it an observed increase, not a confirmed improvement.
Do not use one noise budget for every metric
Mention, citation, prominence, sentiment, and factual accuracy behave differently.
- Mention is binary and relatively easy to score.
- Citation requires a link or source reference and should be stored at URL level.
- Prominence depends on an explicit rule: first named, top three, recommended, or merely included.
- Sentiment is a classification judgement and is usually noisier.
- Accuracy must be checked against a dated source of truth.
A 2026 production-data paper on GEO measurement reported that sentiment was 6.7 times noisier than mention in its dataset. That is one vendor’s observational result, not a universal constant, but it illustrates why a single blended “AI visibility score” can hide more than it reveals.
Publish the components separately. A brand can have a rising mention rate while its citation rate falls or its description becomes inaccurate. Those are materially different outcomes.
Establish a noise floor before claiming a win
Before changing the site, run the unchanged benchmark over several measurement windows. This creates a baseline for natural movement.
Then document the intervention:
- page or entity changed;
- exact publication time;
- URLs affected;
- intended prompt family;
- no other known campaign, migration, or major PR event;
- first date the changed page was observed or cited.
Continue measuring with the same core prompts and conditions. Compare post-change movement with pre-change variation.
This is still not a perfect causal experiment. AI engines can update during the same period, and third-party sources can change independently. But it is much stronger than editing ten pages and attributing the next dashboard increase to whichever tactic sounds most impressive.
| Evidence | Defensible wording |
| One answer changed | “The brand appeared in this observation.” |
| Repeated rate increased | “Observed mention frequency increased in this prompt set.” |
| Increase exceeds baseline variation | “The change is larger than the noise previously observed in this test.” |
| Controlled intervention with comparison group | “The intervention is a plausible cause of the measured difference.” |
| Repeated controlled result across engines and periods | “The evidence supports a robust effect under the tested conditions.” |
Avoid jumping from the first row to the last.
Build a holdout set to catch overfitting
If you optimize pages against the exact prompts used for reporting, you may improve performance on the test without improving broader buyer discovery.
Reserve 20–30% of your prompt families as a holdout set. Do not use those prompts to decide what to write or edit. Run them only at defined validation points.
If the core tracking prompts improve but the holdout prompts do not, you may have optimized to the benchmark rather than the market. This is the AI-search equivalent of teaching to the test.
The holdout set should remain commercially relevant. Random prompts do not create scientific rigour; they only create irrelevant data.
Publish the method, including its limitations
E-E-A-T is not a numeric score that can be added with an author box or a block of schema. For readers—and for systems trying to evaluate a source—trust is easier to establish when claims have visible provenance.
A credible public benchmark should state:
- who ran the test;
- what was measured;
- the exact period and platforms;
- how prompts were selected;
- how many observations were collected;
- how answers were scored;
- whether the evaluator had a commercial interest;
- what changed since the previous version;
- what the data cannot prove.
That is the approach used in the continuously updated AEOvara AI Visibility Benchmark Suomi 2026: the platform-level results, prompt scope, cadence, and limitations are part of the evidence, not hidden behind a proprietary score.
Publishing an unfavourable or inconclusive result can strengthen the work. A benchmark that only produces success stories looks like marketing. A benchmark that preserves negative findings begins to look like a useful dataset.
A minimum viable reproducibility checklist
Before publishing the next percentage, ask:
- Are branded and unbranded prompts reported separately?
- Were prompts written before seeing which ones favoured the brand?
- Are the exact wording and realistic variants preserved?
- Were runs made in clean, comparable sessions?
- Are platform, model, retrieval mode, location, and date recorded?
- Does every percentage show its numerator and denominator?
- Is uncertainty visible when the sample is small?
- Are mentions, citations, prominence, sentiment, and accuracy separate?
- Was a pre-change noise floor established?
- Is there a holdout prompt set?
- Are raw answers and cited URLs archived?
- Are limitations and commercial relationships disclosed?
If several answers are “no,” the result may still be useful as an observation. It should not be presented as a stable market measurement.
The bottom line
AI visibility is measurable, but not as a single universal rank.
What you can measure is a defined brand’s behaviour across a documented sample of prompts, engines, configurations, places, and times. Repeat that protocol and you can build evidence. Hide those conditions behind a polished score and you create false precision.
The strongest AI-search report is not the one with the most decimal places. It is the one another practitioner could challenge, reproduce, and learn from.
That is also the kind of source worth citing.
FAQ
How many times should I run each AI-search prompt?
There is no universal number. Three to five repetitions can expose obvious variability and provide a directional baseline, but a small sample does not produce a precise probability. Report the raw count and an uncertainty interval, then increase repetitions for important or volatile prompt families.
Should I freeze my prompts or update them?
Do both in separate layers. Keep a stable core set for trend comparison, and maintain a versioned discovery set that reflects new buyer language, products, and competitors. Never silently replace old prompts and continue the same trend line.
Can I compare an API result with the consumer AI product?
Only with an explicit caveat. The API may use a different model snapshot, system configuration, search policy, location input, or citation interface. Treat it as a separate measurement surface unless equivalence has been demonstrated.
Is a higher mention rate proof that my content change worked?
No. It is evidence that the observed rate changed. A causal claim requires stronger controls: stable test conditions, a pre-change noise floor, documented timing, enough observations, and preferably a comparison or holdout group.
Does a citation mean the AI answer used my page correctly?
Not necessarily. Store and inspect the cited URL, the claim it appears to support, and whether the answer accurately represents the source. Citation presence and citation correctness are separate metrics.
Sources and further reading
About the author: Jarno Saarimies is the founder of AEOvara, a Finnish consultancy focused on technical SEO, Answer Engine Optimization, Generative Engine Optimization, and reproducible AI-search visibility measurement. His work emphasizes documented methods, source verification, and claims that remain useful after the hype cycle moves on.