How to Measure Brand Visibility in AI Answers (and What No Tool Can Tell You)
There is no ranking in ChatGPT and no single metric for AI visibility. There are three layers of evidence: platform data, answer sampling and business signals. How to combine them, and which limits to disclose to leadership.
Table of contents
- First, Define What Counts as “Visibility”
- Layer 1: What the Platforms Measure (and What They Don’t)
- ChatGPT and GA4: What You See and What You Lose
- Layer 2: Sampling Answers Without Chasing a Fake Ranking
- An AI’s answer changes almost every time
- What this means for your method
- A note on monitoring tools
- Layer 3: Business Signals You Can Actually Observe
- A Five-Step Measurement System
- Common Mistakes When Measuring AI Visibility
- Next Step
If someone offers you a “ranking for your brand in ChatGPT,” ask what exactly is being measured. The honest answer today is that there is no single, official, stable metric for visibility in AI answers. What exists are fragments: data some platforms publish about their own surfaces, answer sampling anyone can design, and indirect business signals. Each fragment measures something different, and none of them alone answers the question leadership actually has: “is this helping us sell?”
This article proposes measuring in three layers and disclosing the limits of each: platform data (Search Console, Bing Webmaster Tools, GA4), controlled sampling of answers, and business signals. It doesn’t define GEO or compare disciplines; for that, see GEO vs. SEO and SEO, AEO, and GEO. Nor does it explain what to do to appear; that’s covered in SEO in the AI Era: What to Do for Google and ChatGPT. The question here is different: how do you know whether it’s working, without fooling yourself?
First, Define What Counts as “Visibility”
Before opening any tool, separate four different outcomes. They’re often blurred together, and each is measured differently:
- Mention: the answer names your brand in the text, with no link.
- Citation: the answer shows your site or one of your pages as a source.
- Visit: someone clicks through and lands on your website.
- Accuracy and tone: what the answer says about you is correct, current and favorable.
The fourth is the one most often forgotten, and it may matter most to the business. Appearing with an outdated price, a service you no longer offer or a wrong description is worse than not appearing. A good working definition for your report: visibility = the share of relevant answers, over a fixed set of customer questions, in which your brand is mentioned or cited with an accurate description. That’s our operating definition, not an industry standard.
Layer 1: What the Platforms Measure (and What They Don’t)
Four sources offer first-party data. The table compares what each measures according to its documentation and, more importantly, what it does not measure:
| Source | What it measures | What it does NOT measure |
|---|---|---|
| Search Console, Performance report (Web type) | Clicks, impressions and position for your site in Google Search. Includes AI Overviews and AI Mode: a click on an external link in the response counts as a click, and all links in an AI Overview share a single position. | The documentation describes no click breakdown by AI feature within this report, and it does not include ChatGPT or other assistants. |
| Search Console, generative AI performance report | Impressions of your links in AI Overviews and AI Mode, by page, country, device and date. Available to all sites since August 31, 2026 (if there are enough impressions). | The documentation lists no clicks or queries. The impressions are already inside the main report. Search Labs experiments are excluded. An "impression" means the link was shown, not that anyone read it. |
| Bing Webmaster Tools, AI Performance (public preview, February 2026; expanded June 2026) | Total citations as a source in AI answers, average cited pages per day, page-level activity, trends and a sample of grounding queries. Covers Copilot, AI summaries in Bing and select partner integrations. Since June 16, 2026 (preview) it adds intents, topics, citation share (the percentage of citations shown for a grounding query that go to your site) and period comparison. | Microsoft clarifies it does not indicate ranking, authority or a page's role in an answer; grounding queries are a sample. The announcement does not list other assistants, such as ChatGPT, among the surfaces covered. |
| GA4, "AI Assistant" channel | Sessions arriving from sources such as ChatGPT, Gemini, Deepseek, Copilot or Grok, when the referrer matches a list of AI assistants. | Google states it excludes AI Overviews and AI Mode. It does not capture mentions without a click. Visits with no referrer can't be assigned to this channel. |
Sources: Google Search Console Help (performance and generative AI reports); Bing Webmaster Blog (February 10, 2026); Google Analytics Help (default channel group). Consulted October 9, 2026. Products change: confirm current status.
Three readings follow from the table, and all three are worth passing on to leadership:
1. No single source covers every surface. Google measures Google; Microsoft measures Microsoft; GA4 sees only visits that arrive with a click. By construction, mentions without a click in ChatGPT or Gemini fall outside the first-party data we reviewed. How to plan visibility when the user does not click is covered in building visibility in zero-click search.
2. “Impression” and “influence” are not the same thing. Search Console’s generative AI report tells you how many times your links were shown. It doesn’t tell you how many people saw them, read them or remembered your brand afterward. It’s an exposure figure, useful for tracking trends by page, not a measure of impact.
3. Don’t add what’s already been added. Google states that the generative AI report’s data is part of the main performance report, and that the chart total aggregates by property: if two results from your site appear in the same response, they count as one impression. Add both reports together and you double count.
ChatGPT and GA4: What You See and What You Lose
For traffic arriving from ChatGPT, OpenAI’s publisher FAQ says that if you allow OAI-SearchBot, you can track referral traffic in your analytics tools, and that ChatGPT includes the utm_source=chatgpt.com parameter in referral URLs. Trade press reported in 2025 that the parameter was already on citations and, from June 13, 2025, was added to links in the “More” section as well (PPC Land, 2025).
What this lets you do: segment GA4 sessions sourced from chatgpt.com, see which landing pages they hit and what they do next (inquiries, forms, time on site). What it doesn’t let you do:
- Measure mentions without a click. If a user reads your name in an answer and then searches for you on Google or types your domain, that visit will show up as branded search or direct, not as AI.
- Guarantee clean attribution. Classification depends on a referrer or parameter being present. According to PPC Land (2025), before June 2025 links in ChatGPT’s “More” section carried no parameters and appeared as direct traffic. Check your own data; don’t assume the channel captures everything.
- Compare channels with the same yardstick. An AI session and an organic search session don’t share the same cost or intent. Measure conversion, not just sessions. For attribution generally, see Marketing Attribution: Which Model to Choose.
One practical complement you do control: server logs. They show which crawlers visit your site (OpenAI publishes IP ranges for OAI-SearchBot, GPTBot and ChatGPT-User so you can validate them). They show that a system is crawling you, which is a precondition for being used, but they do not prove it cites you. Treat them as an access diagnostic, not a visibility metric.
Layer 2: Sampling Answers Without Chasing a Fake Ranking
Platform data doesn’t cover ChatGPT, Claude or most conversational answers. That’s what sampling is for: ask the systems what a customer would ask and record what happens. Monitoring tools do this automatically, and you can do it manually at small scale. The problem is how to interpret the result.
An AI’s answer changes almost every time
SparkToro and Gumshoe published a study on this in January 2026. With 600 volunteers, 12 brand-recommendation prompts and 2,961 runs across ChatGPT, Claude and Google’s AI (AI Overviews, and AI Mode when no Overview appeared), they found less than a 1-in-100 chance of getting the same list of brands in two responses, and getting the same order as well was even less likely. Their stated conclusions: visibility percentage across dozens to hundreds of runs is a reasonable metric, and a tool that offers an “AI ranking position” is unreliable.
The limitations the authors themselves disclose matter as much as the finding: the work isn’t peer reviewed, the data comes from November and December 2025 (models have changed since), and it leaves unresolved how many runs are needed for statistical soundness and whether API calls behave like manual ones.
What this means for your method
With that in mind, a defensible sampling approach follows six rules:
- A fixed set of questions, phrased the way real customers would say them (discovery, comparison, decision), not the way your marketing team would write them.
- Multiple runs per question and platform. A single response is not a data point.
- Record the percentage of appearances, not position. If you appear in 7 of 20 runs, your visibility on that question is 35%, with the uncertainty that implies. Report it as a range or a trend.
- Distinguish mention, citation and accuracy in each appearance (the definition in the first section).
- Measure your competitors on the same questions. A 25% figure on its own is a number; 25% against a competitor’s 45% is a strategic conversation.
- Repeat on a fixed cycle (monthly or quarterly) and compare trends, not isolated readings. When the model or interface changes, note it in the log.
An illustrative, hypothetical example (not real data from any client) shows how to read the results. Imagine a services company testing 20 customer questions, each run 5 times on one platform, for 100 responses:
| Indicator | Your brand | Competitor A | How to read it |
|---|---|---|---|
| Responses that mention the brand | 22 of 100 | 41 of 100 | Presence gap across the chosen question set |
| Responses citing your site as a source | 9 of 100 | 18 of 100 | Mentions without links: an opportunity for citable content |
| Mentions with an accurate description | 15 of 22 | Not assessed | 7 mentions with wrong details: a correction priority |
Invented, illustrative figures used to explain the method. In a real measurement, repeat the cycle and compare trends before drawing conclusions.
In this example, the actionable finding isn’t “we’re at 22%.” It’s that nearly a third of the mentions are inaccurate, and that there’s a pattern of questions where the competitor appears and you don’t. That directs content and corrections. An isolated number doesn’t.
A note on monitoring tools
A category of tools automates this sampling. We don’t evaluate specific products here. To choose one, ask the vendor: how many runs it does per question; whether it reports percentages or “positions”; how it handles variability between runs; from which locations and languages it queries; whether it uses the user interface or the API (SparkToro’s researchers leave the difference between the two open); and how it exports data so you can audit it. If they can’t answer, don’t buy a number. Buy a diagnosis.
Layer 3: Business Signals You Can Actually Observe
Visibility in an answer is worth what it causes afterward. Three signals, none conclusive on its own, help connect exposure to business:
Brand demand. Search Console includes a branded and non-branded query filter in the Performance report (announced in November 2025 and extended to all eligible sites on March 11, 2026). Google warns that classification is done by an AI-assisted system and may misidentify some queries. It isn’t available for URL-path or subdomain properties or for sites with low impressions. A sustained upward trend in branded queries, alongside improvements in sampled visibility, is a coherent indication, not proof. Always read brand demand alongside direct traffic: the Similarweb study cited below suggests part of the influence of an AI recommendation shows up as a direct visit rather than a search.
Direct and AI referral traffic. The trend in direct visits and in GA4’s “AI Assistant” channel sessions, read alongside the other two layers. To carry those sessions through to a return calculation, see the digital marketing ROI calculation.
Self-reported source. A “how did you hear about us?” field on forms and in the first sales call, logged in the CRM with options that include “an AI (ChatGPT, Gemini, etc.),” captures exactly what analytics can’t see: mentions without a click. It’s subjective and biased, and it’s often the only direct evidence of the effect of an answer nobody clicked on. Its value increases when you cross it with opportunity quality. For the system that supports this, see how to build a marketing measurement system.
On the link between mentions and visits, there’s first evidence with clear limits. A Similarweb study, with contributions from Rand Fishkin (SparkToro) and discussed by him on June 29, 2026, analyzed three consumer industries (finance, travel and beauty) with panel browsing data. It observed that on devices exposed to an AI recommendation of a brand, direct visits to that brand increased while visits arriving through traditional search dipped slightly: AI influence seems to change the path to the brand more than it adds visits. The authors list the open questions themselves: whether those users would have found the brand anyway, whether the result holds for small brands, and how AI’s influence compares with other channels. For a mid-sized B2B company, the study suggests a reasonable hypothesis to measure. It doesn’t prove the effect in your market.
A Five-Step Measurement System
Editorial framework · Maccam Network
-
Define the business question and the customer question set
Choose 20 to 40 real customer questions, grouped by stage (discovery, comparison, decision) and priority service. Source them from sales calls, forms and site search, not a brainstorm.
-
Turn on and verify first-party data
Search Console with the generative AI report (if available), Bing Webmaster Tools with AI Performance, and a GA4 view that segments the "AI Assistant" channel and sessions by source. Document what each one doesn't measure.
-
Sample methodically and repeat
Run the question set several times per platform, record mentions, citations and accuracy, and calculate percentages. Include two or three competitors. Repeat on fixed cycles.
-
Connect to business signals
Add a self-reported source field with an AI option to the CRM, track branded queries and direct traffic, and review the quality of incoming opportunities. Avoid assigning causality from a single data series.
-
Report with declared limits and decide what to fix
Present each figure with its source, period and limitation. Prioritize actions from specific findings: inaccurate mentions, questions with no presence, pages with no citations. Review quarterly.
This is a working method, not an industry standard. Its value is consistency: same questions, same procedure, same reports, so that differences between cycles actually mean something.
Common Mistakes When Measuring AI Visibility
- Buying an “AI ranking.” As the SparkToro and Gumshoe research shows, brand lists vary in almost every response. A single position has no statistical basis.
- Measuring once. A screenshot of one ChatGPT answer is an anecdote. A series of samples is a trend.
- Adding overlapping reports. Generative AI impressions plus general performance double counts data.
- Forgetting accuracy. Increasing mentions without watching what they say can backfire.
- Attributing every change in direct traffic to AI. Direct traffic has many causes. Cross-check series and ask sales.
- Reporting exposure metrics to leadership as if they were results. Visibility is not a lead. The leadership conversation starts with the marketing metrics that matter to the CEO.
- Measuring before fixing the basics. If your key pages aren’t eligible or your information is inconsistent, measuring will only document the problem. Start with the operating plan in SEO in the AI Era.
Next Step
If you need to build this system, from choosing the questions to configuring Search Console, GA4 and the CRM so they tell a consistent story, the work combines analytics and strategy: deciding what matters to measure before deciding which tool to use. It’s part of our marketing analytics service and builds on Maccam Network’s SEO methodology. If you’d like to talk it through, get in touch.
Sources
- Google Search Console Help. Generative AI performance report (Search). support.google.com/…/16984139
- Google Search Central. (2026, June 3). Introducing Search Generative AI performance reports in Search Console. developers.google.com/…/gen-ai-performance-reports
- Google Search Console Help. What are impressions, position, and clicks? (sections on AI Mode and AI Overviews). support.google.com/…/7042828
- Google Search Central. AI features and your website. developers.google.com/…/ai-features
- Google Search Central. (2025, November 20; updated March 11, 2026). Introducing the branded queries filter in Search Console. developers.google.com/…/search-console-branded-filter
- Google Analytics Help. Default channel group. support.google.com/…/9756891
- Microsoft, Bing Webmaster Blog. (2026, February 10). Introducing AI Performance in Bing Webmaster Tools Public Preview. blogs.bing.com/…/Introducing-AI-Performance-i…
- Microsoft, Bing Blog. (2026, June 16). New AI Visibility Insights in Bing Webmaster Tools: Intents, Topics, Citation Share, Compare. blogs.bing.com/…/New-AI-Visibility-Insights…
- OpenAI. Overview of OpenAI crawlers. developers.openai.com/…/bots
- OpenAI Help Center. Publishers and Developers FAQ. help.openai.com/…/12627856-publishers-and-deve…
- PPC Land. (2025, June 16). ChatGPT adds UTM parameters to more links for better analytics tracking. ppc.land/chatgpt-adds-utm-parameters-…
- Fishkin, R. and O’Donnell, P. / SparkToro and Gumshoe. (2026, January 28). NEW Research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility. sparktoro.com/…/new-research-ais-are-highly-…
- Fishkin, R. / SparkToro. (2026, June 29). New Research from Similarweb: How AI Brand Mentions Influence Direct Visits & Traditional Search Queries. sparktoro.com/…/new-research-from-similarweb…
Sources verified as of October 9, 2026.
Preguntas frecuentes
Not reliably. A SparkToro study with Gumshoe (January 2026), covering 2,961 responses from ChatGPT, Claude and Google's AI to 12 prompts, found less than a 1-in-100 chance of getting the same list of brands in two responses, and an even lower chance of the same order. The authors conclude that visibility percentage across many runs is a reasonable metric, and that any tool offering an 'AI ranking position' is unreliable. It is not peer reviewed, and its data was collected in November and December 2025.
According to Google's documentation, it measures impressions of links to your site shown in generative AI features in Search (AI Overviews and AI Mode), grouped by page, country, device and date. It excludes Search Labs experiments and does not appear if the site hasn't received enough impressions. Those impressions are already part of the main performance report, so they should not be added to it.
Partly. Analytics has an 'AI Assistant' channel in the default channel group for visits arriving from sources such as ChatGPT, Gemini, Deepseek, Copilot or Grok. It is assigned when the referrer matches a list of AI assistants, and it excludes Google's AI Overviews and AI Mode. Visits that arrive without a referrer can't be classified this way, and the channel measures visits, not mentions in answers that didn't produce a click.
In February 2026, Microsoft introduced the AI Performance report in public preview. It shows total citations of your site as a source in AI answers, average cited pages per day, a sample of grounding queries and page-level activity, with a timeline. It covers Microsoft Copilot, AI-generated summaries in Bing and select partner integrations. Since June 16, 2026 it adds, also in preview, views for intents, topics, citation share and period comparison. Microsoft clarifies that the figures do not indicate ranking, authority or a page's role within an answer.
There is no validated number. SparkToro's own researchers list how many times a prompt must be run for statistically sound results as an open question. As a practical method: use a small, fixed set of real customer questions, repeat each one several times per platform, report the share of appearances as a range, and compare trends across several cycles rather than relying on one measurement.
Be honest about causality. A Similarweb study discussed by SparkToro (June 2026) across three consumer industries observed that devices exposed to an AI recommendation of a brand showed more direct visits to that brand and slightly fewer visits arriving through traditional search, and its authors note open questions about causation and about smaller brands. For your own case, combine brand demand signals, a self-reported source question on forms, and inquiry quality in the CRM, and present the relationship as an indication, not proof.
New ideas, analysis and research — directly to your inbox.
Subscribe to receive new Insights publications and other selected content from Maccam Network. No spam. Unsubscribe at any time.
Shall we talk about your business?
Let's talk about what your business needs.
A 30-minute conversation is enough to understand the context, identify the problem and see if we are the right team to help you.