Skip to content

AI search

How Do You Measure AI Search Visibility?

Key takeaways

  • Answers change run to run, so measure how often you appear across many prompts and repeats, and report a range, not what one answer said.
  • Google's Generative AI performance report covers impressions for AI Overviews and AI Mode, while clicks and position are counted in the Performance report under set rules.
  • Bing Webmaster Tools reports citations for Microsoft's own AI surfaces, and Microsoft calls its Citation Share metric observational.
  • OpenAI says ChatGPT adds utm_source=chatgpt.com to referral links, which gives you a traffic signal to track.
  • The 30 to 50 prompt, monthly method in this guide is our own recommendation, and the uncertainty ranges are our own arithmetic.

You measure AI search visibility by running a fixed set of prompts many times and recording how often your brand is mentioned or cited, then adding referral traffic and the first-party reports from Google and Bing. No single number captures it, and a single screenshot proves nothing.

That’s the unglamorous truth. AI answers are generated fresh each time, most engines publish no reporting, and the tools that fill the gap vary in quality.

Here’s what can be measured, why it’s hard, and how to build a routine you can repeat. The prompt-set method is our own recommendation, built on published research and labelled as such, and it sits alongside our wider coverage of generative engine optimisation.

What Can You Measure?

Five signals: citations, brand mentions, your share against competitors, referral traffic and impressions. Each comes from a different place, and none is complete on its own.

Signal What it tells you Where it comes from Main weakness
Citations Your URL was linked as a source Prompt set; Bing AI Performance Varies run to run
Mentions Your brand was named in the answer text Prompt set Needs reading, not just counting
Share against competitors How often you appear compared with rivals Prompt set; Bing Citation Share Only as good as your prompt list
Referral traffic People clicked through to your site Analytics Misses answers nobody clicks
Impressions Links to your site were shown Google’s Generative AI performance report Google only; no clicks column described

What’s the difference between a citation and a mention?

A citation is a link to your page. A mention is your brand named in the text, with or without a link.

Log them separately. An engine can describe you accurately without linking, or link to you without naming you.

What does share against competitors add?

It tells you where you stand rather than just whether you appear. Count how often you show up against the same named rivals across the whole prompt set, which is the nearest thing to a share-of-voice figure.

It’s only as good as your prompt list, so build that carefully. Bing’s Citation Share is a different thing: the percentage of citations for a grounding query that go to your site.

Why track referral traffic as well?

Because it’s the only signal that proves someone acted. It undercounts visibility, since a brand can be named in an answer with no click, but it tells you which pages earn visits.

Why Is AI Visibility So Hard to Measure?

Because answers vary from run to run, people phrase prompts in very different ways, and most engines don’t report to site owners. A result you see once is a sample of one.

How much do answers vary?

SparkToro’s Rand Fishkin and Patrick O’Donnell of Gumshoe.ai published a study on 27 January 2026. Their 600 volunteers ran 12 prompts through ChatGPT, Claude and Google’s AI features, producing 2,961 runs in November and December 2025.

Fishkin reported a less than 1 in 100 chance that ChatGPT or Google’s AI would give the same list of brands in any two responses. For the same order it was less than 1 in 1,000, and the number of brands in each list varied too.

How differently do people ask?

The study also collected 142 human-written prompts from volunteers and reported an overall semantic similarity of 0.081, which Fishkin likened to the gap between Kung Pao chicken and peanut butter. A handful of tidy test prompts won’t mirror real customers.

Yet when 142 human headphone prompts were run for 994 responses, the same top brands kept returning. Fishkin concluded that visibility percentage across dozens to hundreds of prompts, run multiple times, is a reasonable metric.

What can’t the study tell you?

Plenty. It isn’t peer reviewed, O’Donnell works at Gumshoe, an AI tracking vendor (Fishkin discloses this), and the tools were the ones popular in the US.

Volunteers also used their default settings with no controls on memory, device or country. SparkToro itself says that how many times you need to run a prompt for statistically sound answers is still an open question.

What do the engines report to you?

Some do, and they differ. OpenAI documents a UTM parameter but, in the pages we read, describes no reporting dashboard, while Google and Microsoft offer first-party reports that we cover below.

For other engines you’re relying on your own testing or on third-party tools, whose methods you should check.

How Do You Build a Repeatable Prompt Set?

List the questions your customers actually ask, fix the wording, and run each prompt several times under the same conditions.

The structure below is our recommended approach. It leans on SparkToro’s finding that visibility percentage across dozens to hundreds of prompts, run multiple times, is a reasonable metric, and the specific numbers are our judgement, not a published standard.

1. Choose prompts by intent

Start with 30 to 50 prompts across your main products or services, in three groups. Use unbranded category questions, comparison questions that name competitors, and branded questions that name you.

Write some of them the way people really talk, long and specific. Customer emails, sales calls and your Search Console queries are good raw material.

2. Control the conditions

Use the same account type and location every time. OpenAI’s help centre says ChatGPT may use an approximate location from your IP address and, if memory is enabled, relevant saved memories when rewriting a search query, and that a VPN can change the inferred location.

So test from a clean account with memory off. Remember that SparkToro’s volunteers used default settings, so a clean baseline is tidier than what any one customer sees, and a few logged-in spot checks are a sensible complement.

3. Repeat each prompt

Run every prompt several times and record the proportion of runs in which you appear. We suggest five runs per prompt per month as a minimum, which gives 150 to 250 runs per engine, but the more runs, the narrower your uncertainty.

Here’s why. The table uses a standard binomial (Wilson) calculation, which treats runs as independent, so it’s our arithmetic and a best case.

Appearances Runs Observed rate 95% range
3 10 30% 11 to 60%
5 10 50% 24 to 76%
15 30 50% 33 to 67%
30 100 30% 22 to 40%
62 200 31% 25 to 38%

Ten runs can’t reliably tell 30% from 60%. Two hundred can separate 31% from 50%.

4. Record more than “yes or no”

Log the date, engine, prompt, run number, whether you were mentioned, whether you were cited, which URL, which competitors appeared and how you were described. Treat any ranking position as noise.

Column Why you need it
Date and engine Answers drift, and engines differ
Prompt and run number Lets you rebuild the denominator
Mentioned / cited (URL) The two core signals, kept separate
Competitors named Needed for share against rivals
How you were described, and whether it’s accurate Catches errors a count would hide

5. Review on a schedule

Re-run monthly, and again after any change you want to test. There’s no published guidance on cadence, so choose one you can sustain.

Want a professional baseline before you build your own routine? Book an AI visibility audit, from £500.

How Do You Track Referral Traffic From AI Engines?

Filter your analytics for the referral source each engine uses, starting with the one OpenAI documents: utm_source=chatgpt.com. OpenAI’s publisher FAQ says publishers who allow OAI-SearchBot can track ChatGPT referrals in tools such as Google Analytics.

How do you set it up in Google Analytics?

In GA4, open the Traffic acquisition report, filter by session source or medium for the values you find, and view the same data by landing page. Menu names change, so look for the report that breaks sessions down by source.

Save the filter or group the AI sources into a custom channel so you can compare month on month. The landing-page view shows which pages earn visits from AI answers.

Which sources should you look for?

OpenAI documents one parameter. We haven’t found official documentation for other engines, so list the referrers that actually appear in your own data and group the AI ones.

Some visits may arrive with no referrer and show up as direct. We found no published figure for how many.

Related: How Do You Get Your Brand Cited by ChatGPT?

What can your server logs add?

Crawler hits from AI bots show that your pages are being fetched, which is useful for checking access. They don’t show that you were cited.

Related: Should You Block AI Crawlers Like GPTBot?

What pattern should you watch in Search Console?

Look for pages with high impressions and falling click-through. Google counts an AI Overview as one position shared by every link in it, so a page cited there can gain impressions without gaining clicks.

That’s our reading of Google’s counting rules rather than a Google statement, so confirm it against the Generative AI performance report for the same pages.

What Do Google and Bing Report?

Google reports impressions for AI Overviews and AI Mode, and Bing reports citations for Microsoft’s AI surfaces. Neither is documented as covering ChatGPT, Perplexity or Claude.

Report Surfaces covered What it shows Caveats
Google Generative AI performance report AI Overviews, AI Mode Impressions by page, country, date, device Impressions only as described; excludes Search Labs
Google Performance report Web search, including AI features Clicks, impressions, position AI Overview links share one position
Bing AI Performance Copilot, AI summaries in Bing, select partners Total citations, average cited pages, grounding queries, page-level citations Grounding queries are a sample
Bing Citation Share (preview) The same Bing AI surfaces Your share of citations per grounding query Observational; competitors not shown

What does Google’s Generative AI performance report show?

Google’s help page says it covers impressions for AI Overviews and AI Mode, grouped by page, country, date and device, with a filter for text-based and multimodal searches. A note on it says Google rolled the report out to all websites worldwide as of 31 August 2026.

The usual Search limits apply, including the 1,000-row cap, and Search Labs experiments aren’t included. The page describes impressions only, so don’t expect clicks there.

Where do AI clicks show up?

In the standard Performance report. Google’s help page on impressions, position and clicks says clicking an external link in an AI Overview or in AI Mode counts as a click.

An AI Overview takes a single position shared by all its links, AI Mode follows the usual Search position methodology, and a follow-up question in AI Mode counts as a new query. The page doesn’t describe a way to isolate AI clicks in the Performance report, so don’t assume you can.

Related: How Do Google AI Overviews Choose Their Sources?

What does Bing’s AI Performance report show?

Microsoft’s AI Performance report launched in public preview on 10 February 2026. It covers Microsoft Copilot, AI-generated summaries in Bing and select partner integrations.

It shows total citations, average cited pages per day, page-level citation counts, a timeline and grounding queries. Microsoft says the grounding queries are a sample of overall citation activity, and that citation counts don’t indicate placement, ranking or importance.

What did Bing add in June 2026?

A 16 June 2026 update added four preview features: Intents, Topics, Citation Share and Compare. Citation Share is the percentage of citations for a grounding query that go to your site.

Microsoft calls it an observational metric, not a ranking or competitive scoreboard, and says it doesn’t expose competitor domains or represent traffic share. Compare overlays a previous period, such as the prior 30 days, on the current view.

What Should Go in a Monthly AI Visibility Report?

Report mention rate, citation rate, share against named competitors and accuracy, each with a range, and keep traffic and impressions alongside them. Keep the definitions the same every month.

Metric How to calculate it Caveat
Mention rate Runs naming your brand ÷ total runs Depends on prompt set
Citation rate Runs linking to your site ÷ total runs Needs the answer’s sources logged
Share against rivals Your appearances ÷ (yours + each rival’s) on the same prompts Relative, not absolute
Accuracy Answers that describe you correctly ÷ answers that name you Manual review
AI referrals Sessions from AI sources in analytics Excludes no-click visibility
Google and Bing data Generative AI impressions; Bing citations and Citation Share Different engines, not additive

What does a worked example look like?

These numbers are illustrative, not client data. Say 40 prompts run five times each on one engine gives 200 runs, and your brand is named in 62 of them: 31%, with a 95% range of roughly 25 to 38%.

A month later it’s named in 74 runs: 37%, range 31 to 44%. The ranges overlap, so you can’t say it moved.

Three months in, 100 of 200 is 50%, range 43 to 57%. That no longer overlaps with 25 to 38%, so the change is likely real.

Which metrics should you distrust?

Distrust any single “AI rank” or position. Fishkin called any tool that offers a ranking position in AI “full of baloney”, because order changed far more than whether a brand appeared.

Also distrust a number reported without its run count, and a screenshot offered as evidence.

Why track accuracy as well?

Because being named wrongly can be worse than not being named. OpenAI says search results and citations can be incomplete, outdated or incorrect, so log errors and fix the pages the engine is likely to read.

What Are the Limits of AI Visibility Tools?

Third-party tools run prompts for you, but they can’t see inside the engines. Google’s guide says no third-party tool has access to its internal ranking or AI systems, so be wary of any that claim otherwise.

What do tools do well?

They run prompts on a schedule, keep history and compare you with rivals at a scale a spreadsheet can’t match. That saves real time once the method is sound.

What should you ask a vendor?

  • How many times is each prompt run per period, and do you show ranges?
  • Are the prompts synthetic or written by people, and can I add my own?
  • Which account settings and locations do you use?
  • Do you collect answers through the app or an API, and is there evidence they match?
  • Which engines and surfaces are covered, and can I export the raw answers?

SparkToro asks tool sellers to publish transparent, reviewable research, and those questions are a fair test of whether they have.

When is a tool worth paying for?

When you have a working method, enough prompts that manual runs are a chore, or several markets to cover. Until then a spreadsheet does the job, and if you want an outside baseline first, an AI visibility audit is the quicker route.

How Do You Know Whether a Change Worked?

Change one thing at a time, keep a comparison set of unchanged pages and prompts, and judge the result against the ranges rather than single numbers. That’s how you separate your effect from the engines’ own drift.

How do you set a baseline?

Run the full prompt set before you touch anything, and run a second batch a week later. The gap between the two shows you the natural wobble for your market.

Why keep a control group?

Pick prompts that map to pages you won’t change. If those move as much as the prompts you did change, the engine moved and you didn’t.

How do you read the result?

If the new range overlaps the old one, treat the change as unproven and keep collecting runs. At Utterly Digital we test ranking factors in single-variable experiments before applying them to client sites, and our AI search work uses the same approach.

Ready to turn this into a baseline for your brand? Start with an AI visibility audit.

1Build the prompt setFix 30 to 50 prompts across your topics.
2Run it repeatedlyRepeat each prompt several times, same settings.
3Record the resultsLog mentions, citations and competitors.
4Compare and actReview monthly, change one thing, re-run.
A repeatable AI visibility routine

FAQs

Can I test signed out?

For ChatGPT, yes. OpenAI's help centre says people who aren't signed in can use web search, so a signed-out session is another way to get a clean baseline without personal memory. We haven't verified the position for other engines.

How do I measure visibility for a local business?

Put your area in the prompts and run them from the area you serve. OpenAI says ChatGPT may use an approximate location from the IP address for local results, and suggests adding a city, neighbourhood or postcode when results look wrong.

What is a good visibility percentage?

There's no benchmark, and the right number depends on how crowded the market is. In SparkToro's study, top brands in a narrow headphones market appeared in 55 to 77% of responses, while a broad brand-design-agency prompt gave top figures in the 30s and 40s, so compare yourself with named rivals on the same prompts.

Does running prompts through an API give the same answers as the app?

Nobody has settled it. SparkToro lists it as an open question, saying early data suggests API caching may not be the insurmountable problem it presumed, but it didn't test it exhaustively. Ask any tool vendor which method they use.

What should I do if an AI engine describes us wrongly?

Log it, then check what the engine could have read. OpenAI says search results and citations can be incomplete, outdated or incorrect, so fix or date the pages it retrieves and keep your business details consistent wherever they appear.

Can I do all this for free?

Mostly. Analytics, Search Console, Bing Webmaster Tools and a spreadsheet of prompts cost nothing, though the manual runs take time. Paid tools mainly save effort and add history, which is only worth buying once you have a method.

Sources

  1. NEW Research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility, SparkToro
  2. Generative AI performance report (Search), Google Search Console Help
  3. What are impressions, position, and clicks?, Google Search Console Help
  4. Optimizing your website for generative AI features on Google Search, Google Search Central
  5. Publishers and Developers - FAQ, OpenAI Help Center
  6. Searching the web with ChatGPT, OpenAI Help Center
  7. Introducing AI Performance in Bing Webmaster Tools Public Preview, Microsoft Bing Webmaster Blog
  8. New AI Visibility Insights in Bing Webmaster Tools: Intents, Topics, Citation Share, Compare, Microsoft Bing Search Blog

Related services

Want this done for you? Let’s talk.

Tell us about your site and your goals. You get a clear scope and price before any work starts.

Talk to a specialistGet an AI visibility audit