How to Measure Brand Visibility in AI Answers: an AI Visibility Methodology Without Self-Deception

ai marketing seo brand measurement

Not long ago, a marketer had a fairly clear picture of what "brand visibility in search" meant.

There is a search results page. There is a website position. There are impressions, clicks, CTR, share of search traffic.

With generative AI, that logic stops working.

A user asks:

"Which bank is best for travel?"

And gets not ten links but a ready answer. It may name three banks: one recommended first, another mentioned in passing, a third given as an alternative. Sometimes the AI shows sources. Sometimes it doesn’t.

Ask the same question again, and the answer may change.

Phrase the question slightly differently, and the set of banks changes.

Ask it a different model from the same vendor, and the result changes again.

And if the model fundamentally doesn’t know the Belarusian banking market well, an even more fundamental question arises: what exactly are we measuring? The visibility of a specific bank, or the model’s ignorance of the market?

So AI Visibility cannot be reduced to the question:

"How many times did ChatGPT mention our brand?"

A proper measurement methodology is needed.

1. The main mistake: treating an AI answer as a search results page

One of the most dangerous approaches looks roughly like this.

Take 100 questions. Ask each question to the AI once. Count how many times our brand and competitors are mentioned.

You get:

Bank A: 67 mentions Bank B: 51 mentions Bank C: 39 mentions

And conclude: "Bank A’s AI Visibility is 67%, and it’s the category leader."

The math here can be perfectly correct. The problem is the experimental design.

The answer of a generative system is not a fixed SERP. So a single AI answer should be treated as one observation, not as an established "brand position."

2. AI Visibility should be thought of as a probability

Suppose we independently ask the AI the same question five times:

"Which bank should I choose for travel?"

Run Bank A Bank B Bank C
1 Yes Yes No
2 Yes No Yes
3 Yes Yes No
4 No Yes Yes
5 Yes No No

Now we no longer say: "Bank A is present in the answer."

We say:

Under the specified experimental conditions, bank A appeared in 80% of observations.

AI visibility is a distribution of outcomes, not a single screenshot.

3. What counts as the unit of observation

The unit of observation for AI Visibility = one answer from a specific AI system to a specific prompt under fixed experimental conditions.

The conditions include at minimum:

  • the platform;
  • the model or available mode;
  • the version, if known;
  • the language;
  • the date of the study;
  • whether external search was enabled, if that mode is controlled;
  • the conversation context;
  • personalization;
  • the wording of the question.

From this follows a crucial principle: the name of an AI product by itself is not a precise enough description of the measurement environment.

4. Start not with the brand but with market knowledge

This matters especially for small markets.

If a language model has received relatively little information about a local banking system, it may answer general finance questions well while knowing very little about the specific banks of a country.

So the first metric of the study should be KPI 0. Market Knowledge.

5. KPI 0: Market Knowledge

Before analyzing brands, a special set of non-branded questions is asked:

  • Which banks operate in Belarus?
  • Which banks offer cards to individuals?
  • Which banks are worth considering for business banking?
  • Which banks offer mobile banking apps?

A reference list of existing banks is prepared in advance.

Market Recall

Market Recall = correctly named banks / number of banks in the reference set × 100%

If the benchmark has 20 banks and the AI correctly named 15:

Market Recall = 75%

Market Precision

If the AI produced ten names, of which seven are correct:

Market Precision = 7 / 10 = 70%

Recall shows how completely the system covers the market. Precision shows how correctly it describes it.

6. Why Market Knowledge changes the interpretation

AI Market Recall Mention Rate of bank A
AI 1 92% 48%
AI 2 87% 31%
AI 3 18% 4%

For AI 3, a Mention Rate can be computed mathematically. But analytically, the 4% figure is far less meaningful: the model barely covers the category under study.

Low brand visibility should not automatically be read as a brand problem if the model itself barely knows the market.

The sufficient Market Knowledge threshold should be set before the study, and the chosen value should not be presented as a universal industry standard.

7. Which questions to ask

Another common mistake: measuring brand visibility only with branded questions.

For example:

"Which is better: bank A or bank B?"

That question is interesting, but it already forces the AI to reason about two named banks.

Far more valuable is:

"Which bank should I choose for everyday purchases?"

Here the bank must get into the AI’s consideration set on its own.

So prompts should be split into at least two categories:

Prompted

The brand is present in the question itself.

Unprompted

The question describes a user need but contains no brands.

And these metrics must not simply be mixed together.

8. Research intents, not a random list of questions

For banks these might be:

  • a bank for salary;
  • a card for travel;
  • favorable transfers;
  • a deposit;
  • a loan;
  • a mortgage;
  • a mobile app;
  • premium banking;
  • a bank for business;
  • international payments;
  • security;
  • family banking;
  • the youth segment;
  • everyday purchases.

An intent is a more fundamental unit than a specific phrase.

9. Ask one intent several ways

Intent: choosing a bank for travel

Prompt A: "Which bank should I choose for travel?"

Prompt B: "Recommend a bank whose card is convenient to use abroad."

Prompt C: "Which bank is best for a card if I travel often?"

If the results shift sharply, we have found Prompt Sensitivity.

10. Repeat every query

For example:

100 intents × 3 paraphrases × 5 repeats = 1,500 answers for one AI system.

Five repeats here is not a universal standard. It is an element of a specific study’s design.

11. Never run 100 questions in one conversation

The scheme "question #1 → question #2 → ... → question #100" in a single conversation introduces an extra experimental variable: prior context.

The main benchmark must be built on standardized independent contexts:

Clean context → Prompt → Response → record the result → end the session

Then a new independent observation.

12. Personalization is a separate study

You need to separate:

Clean AI Visibility

What a user gets under the standardized research context.

Personalized AI Visibility

How the result changes when a specific user context is present.

They must not be silently mixed in one sample.

Personalization Lift

Personalization Lift = Personalized Visibility − Clean Visibility

13. KPI 1: Mention Rate

Mention Rate = number of valid answers mentioning the brand / total number of valid answers × 100%

If bank A appeared in 365 of 500 observations:

Mention Rate = 73%

The correct interpretation: under the given experimental conditions, the brand was present in 73% of observations.

14. KPI 2: AI Share of Voice

Mention Rate answers the question: how often does the bank appear at all?

Share of Voice answers: what share of the competitive presence does the bank get?

For research purposes it is better to count brand presence events, not the number of times the name repeats inside a text.

If:

  • bank A is present in 400 answers;
  • bank B in 300;
  • bank C in 200;
  • bank D in 100;

the sum of presence events = 1,000.

Then:

  • SOV of bank A = 40%
  • SOV of bank B = 30%
  • SOV of bank C = 20%
  • SOV of bank D = 10%

15. Why Mention Rate and SOV are not enough

If each of 100 answers contains both bank A and bank B:

  • Mention Rate A = 100%
  • Mention Rate B = 100%
  • SOV A = 50%
  • SOV B = 50%

But the texts may consistently say: "The best option is bank A. Bank B can also be considered."

Formally both are mentioned, but their influence on the user’s decision differs.

16. KPI 3: Recommendation Rate

Define classes in advance:

0 — not recommended The brand is simply mentioned.

1 — consideration The brand is included in the set of options being considered.

2 — explicit recommendation The AI explicitly proposes choosing the brand.

The primary metric:

Recommendation Rate = number of answers with an explicit recommendation / number of relevant answers × 100%

17. KPI 4: Prominence

You can build a transparent weighting system, for example:

  • 1st position = 1.00
  • 2nd = 0.75
  • 3rd = 0.50
  • other significant mention = 0.25
  • no mention = 0

These weights are an analytical construct of the study, not a universal standard. They must be defined in advance and applied identically to all brands.

18. KPI 5: Portrayal

It is important to analyze not only how often the brand is present but how the AI describes it.

A basic classification:

  • positive;
  • neutral;
  • negative;
  • factual accuracy;
  • framing.

Additionally, brand attributes can be coded: innovative, reliable, expensive, mass-market, premium, convenient, complex, technological, outdated.

This is how an AI Brand Perception Map is built.

19. KPI 6: Citation / Source Visibility

You need to separate:

Brand Visibility

The AI mentions the company.

Owned Source Visibility

The AI uses or cites the company’s own resources.

Earned Source Visibility

The AI relies on external materials where the company is present.

For example:

  • Mention Rate = 68%
  • Recommendation Rate = 42%
  • Owned Citation Rate = 7%

That is, the brand is well represented in AI, but its own website almost never appears among the observed sources.

20. Source Attribution as a separate layer

For answers where the system shows sources, it is useful to record:

  • the domain;
  • the URL;
  • the source type;
  • owned / earned / third-party;
  • which brand is associated with the source;
  • whether the citation appeared next to the recommendation.

But it is important not to confuse correlation with causation. The presence of a source next to a recommendation does not yet prove that the source caused the recommendation.

21. KPI 7: Stability

Two concepts must be separated.

Run Stability

What happens when the identical wording is repeated.

Prompt Robustness

What happens across different wordings of the same intent.

These are two different characteristics of AI Visibility quality.

22. Visibility can depend on language

For a multilingual market, the same intent should be tested separately in each relevant language.

For example:

Language Bank A Mention Rate
RU 64%
BE 36%
EN 47%

This yields the diagnostic metric Language Visibility Gap.

23. Platform and model are not the same thing

The name of an AI product must not be treated as the only level of analysis.

You need to distinguish:

Platform Visibility

What an ordinary user of a specific AI product sees.

Model Visibility

What a specific model or mode shows.

Ecosystem Visibility

What happens across the set of relevant AI platforms.

24. The model version belongs in the study passport

For every observation it is desirable to store:

  • Platform
  • Model
  • Version, if available
  • Mode
  • Search ON/OFF, if controlled
  • Language
  • Geo, if relevant and controlled
  • Clean/Personalized
  • Date
  • Prompt ID
  • Run ID

25. Why you cannot compare two waves without controlling conditions

If Mention Rate was 62% in the first wave and 44% in the second, that does not yet prove the brand’s visibility fell.

Between the waves the model, retrieval, prompts, sources, or the product itself may have changed.

Longitudinal AI Visibility requires controlled experimental conditions and explicit marking of methodological breaks.

26. The measurement passport is mandatory

A good report should disclose:

  • the study period;
  • the market;
  • the category;
  • the platforms;
  • the models and modes;
  • the languages;
  • the number of intents;
  • the number of paraphrases;
  • the number of repeats;
  • the context type;
  • the search mode;
  • Market Knowledge;
  • the classification method;
  • the extraction date.

Without this, an AI Visibility Score is poorly reproducible.

27. The absurdity of "let’s check every neural network"

Suppose 11 AI ecosystems are studied:

  1. Alice AI / Yandex
  2. GigaChat
  3. ChatGPT
  4. DeepSeek
  5. Gemini
  6. Perplexity
  7. Qwen
  8. Claude
  9. Grok
  10. GLM / Z.ai
  11. Mistral

For demonstration, take a working factorial design:

11 ecosystems × 3 model/mode × 2 search modes × 3 languages × 2 context types × 100 intents × 3 paraphrases × 5 repeats

That gives:

594,000 AI answers

This is not market statistics. It is the result of one specific demonstration design, showing the scale of the measurement space.

And by the time such a study finishes, part of the technology environment may already have changed.

So it is pointless to strive for completeness of the technical universe. It is more useful to strive for representativeness of the user’s AI universe.

28. Market-Relevant AI Universe

Every serious study should start not with the question "Which neural networks exist?" but with:

Which AI products and usage scenarios matter to our target audience?

That is the Market-Relevant AI Universe.

If there are no reliable data on the audience’s platform usage, do not invent artificial weights. It is better to show each platform’s results separately.

29. A three-tier research architecture

Tier 1. Consumer Benchmark

The core user experience of the AI products the market has chosen.

For example:

11 platforms × 100 intents × 3 paraphrases × 5 repeats = 16,500 observations.

Tier 2. Model Benchmark

We study the effect of specific models where material differences were found.

Tier 3. Deep Diagnostic

For anomalies, we expand to the full factorial:

Model × Search × Language × Context × Paraphrases × Repeated runs

This way the budget is spent only on combinations that yield additional decision-relevant information.

30. A study is not the same as parsing

You can collect a million AI answers and still not conduct a proper study.

The methodology must answer these questions:

  1. What is the population?
  2. What is the unit of observation?
  3. How was the panel of intents formed?
  4. How was the wording controlled?
  5. How many repeats were made?
  6. How was context controlled?
  7. How was personalization accounted for?
  8. How was language accounted for?
  9. Which model was used?
  10. Does the model know the market under study?
  11. How were recommendations classified?
  12. How was prominence computed?
  13. Which answers were considered invalid?
  14. How stable is the result?
  15. Can the measurement be reproduced?

Without these answers, a huge volume of data creates only an illusion of precision.

31. Do not build a "magic AI Visibility number"

Imagine:

Brand A

Mention Rate: 80% Recommendation Rate: 15%

Brand B

Mention Rate: 35% Recommendation Rate: 30%

A has enormous presence inside AI answers but a weak ability to earn recommendations. B appears rarely, but whenever it does, the AI often includes it in the decision.

Collapsing both profiles into one number loses the most interesting part.

32. When a single index is acceptable

For a management dashboard you can build a composite index, for example:

  • 30% Mention Rate
  • 25% Share of Voice
  • 20% Recommendation
  • 15% Prominence
  • 10% Citation

But you must state outright: the weights are an analytical model of this particular study, not a universal industry standard.

The underlying metrics must always sit alongside it.

33. The final KPI system

KPI 0. Market Knowledge

  • Market Recall
  • Market Precision

KPI 1. Mention Rate

The observed probability of the brand appearing.

KPI 2. AI Share of Voice

The brand’s relative share of presence among competitors.

KPI 3. Recommendation Rate

How often the system explicitly recommends the brand.

KPI 4. Prominence

How prominent a place the brand gets.

KPI 5. Portrayal

How the AI characterizes the brand, and how correctly.

KPI 6. Citation / Source Visibility

How the brand and its sources are represented in the observable source system.

KPI 7. Stability

How reproducible the result is across repeats.

KPI 8. Prompt Robustness

How resistant the result is to changes of wording.

Additional diagnostic metrics

  • Personalization Lift
  • Language Visibility Gap
  • Model Sensitivity
  • Platform Variance
  • Recommendation Conversion = Recommendation Rate / Mention Rate

The latter metrics are proposed analytical tools and must be explicitly described in the study methodology.

34. The AI Visibility funnel

Market Knowledge Does the AI know the market at all?

↓

Brand Recall Does the AI recall the brand on its own?

↓

Presence Does the brand appear in the answer?

↓

Prominence Does it get a prominent place?

↓

Consideration Does the brand make the shortlist?

↓

Recommendation Does the AI propose choosing it?

↓

Portrayal How is the brand characterized?

↓

Source Visibility Which observable sources appear?

↓

Stability How reproducible is all of this?

35. Wrong and right approaches

Mistake 1: asking the question once

Right: run a series of independent repeats.

Mistake 2: 100 questions in one chat

Right: standardized independent contexts.

Mistake 3: counting only mentions

Right: Mention + SOV + Recommendation + Prominence + Portrayal + Citation.

Mistake 4: asking only "bank A or bank B?"

Right: add unprompted category/intent questions.

Mistake 5: not checking the model’s knowledge of the market

Right: start with Market Knowledge.

Mistake 6: mixing languages

Right: treat language as a separate experimental cut.

Mistake 7: mixing different models from one vendor

Right: separate Platform Visibility and Model Visibility.

Mistake 8: mixing clean and personalized answers

Right: study them as separate cohorts.

Mistake 9: taking citation for causation

Right: record observable sources without attributing causal effect without separate proof.

Mistake 10: collecting "every AI that exists"

Right: define the Market-Relevant AI Universe.

Mistake 11: showing AI Visibility = 73% without methodology

Right: attach the measurement passport.

Mistake 12: comparing waves under changed conditions

Right: control the changes and mark methodological breaks separately.

36. What a good AI Visibility report looks like

A good report starts from an understanding of the measurement environment:

  1. Which AI systems are relevant to the audience?
  2. How well does each system know the market under study?
  3. How was the panel of user intents formed?
  4. How is variability controlled?
  5. How does the brand look relative to competitors?
  6. How often does a mention turn into a recommendation?
  7. How does the AI characterize the brand?
  8. Which observable sources accompany the answers?
  9. How stable are the conclusions?

Only then does the dashboard appear.

37. An example of the final dashboard

Suppose bank A is analyzed:

  • Market Recall of the studied AI: 89%
  • Mention Rate: 61%
  • AI Share of Voice: 28%
  • Recommendation Rate: 34%
  • Prominence Score: 0.54
  • Owned Citation Rate: 11%
  • Positive Portrayal: 72%
  • Run Stability: high
  • Prompt Robustness: medium
  • Language Gap RU → BE: −17 pp.

Instead of a single number we get a diagnostic AI Visibility profile.

38. AI Visibility and GEO are not the same thing

AI Visibility

This is the measurement of an outcome: how represented the brand is in the AI environment.

GEO

This is an attempt to change that outcome: which actions raise the probability of the brand’s presence, citation, or information being used by generative engines.

You cannot claim GEO worked if measurement before and after is not comparable.

39. Correlation after optimization does not mean causation

If Mention Rate grew from 35% to 47% after a website change, that does not yet prove the website changes caused the growth.

In the meantime, the model, retrieval, the indexable sources, the competitive information field, or the AI product itself may have changed.

The strict formulation:

After the changes, AI Visibility grew by 12 pp. in the comparable part of the measurement panel.

Claiming a causal effect requires a proper causal study design.

40. Why a business should measure all this

In the end, a business needs answers to five questions.

1. Discoverability

Does our brand appear in AI at all?

2. Competitive position

Whom does the AI show instead of us?

3. Recommendation

When the AI helps a person choose, do we make the shortlist?

4. Brand perception

In which words does the AI describe us and our competitors?

5. Source influence

Which observable sources are present in the information environment of the answers?

This is where AI Visibility turns from a pretty digital metric into a marketing strategy tool.

Conclusion

AI has no single stable "results page."

There is:

platform × model × version × retrieval/search × language × context × personalization × prompt wording × generation variability × time of measurement.

So professionalism in AI Visibility lies not in the ability to collect the maximum number of answers.

It lies in the ability to correctly limit the experimental space, standardize observations, and not pass noise off as pattern.

10 principles of quality AI Visibility research

  1. Don’t measure an answer. Measure the distribution of answers.
  2. Check the model’s knowledge of the market first, brand visibility second.
  3. Separate prompted and unprompted visibility.
  4. Study user intents, not a random list of questions.
  5. Use paraphrases and independent repeated runs.
  6. Don’t mix clean and personalized context.
  7. Separate platform, model, and ecosystem visibility.
  8. Don’t collapse Presence, Prominence, Portrayal, and Recommendation into one number without disclosing the methodology.
  9. Don’t try to measure "all the AI." Measure the Market-Relevant AI Universe.
  10. Every AI Visibility Score must have a measurement passport.

A good AI Visibility study answers not the question "what did the AI once say?" but "with what probability, under what conditions, and in what manner does the brand appear in AI answers relative to competitors?"

It is the move from one-off queries and screenshots to a reproducible experiment that turns neural-network monitoring into real marketing research.


Sources and further reading

  • IAB. Measuring Visibility in the AI Era (2026).
  • Aggarwal P., Murahari V., Rajpurohit T. et al. GEO: Generative Engine Optimization (KDD 2024).
  • Martinez O. Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026) (2026).