All posts
Measurement

Measuring AI search visibility what is actually knowable

By Jānis Plūme, Founder, Outbound Pros · 9 min read · 2026-08-06

Quick answer

AI search visibility is partly measurable and partly not, and knowing which half you are in is the whole discipline. Two free first party reports exist: the AI Performance report in Bing Webmaster Tools, covering Copilot and Bing AI summaries, and the generative AI performance report in Google Search Console. Both are underused and both should be switched on in week one. For ChatGPT, Claude and Perplexity there is no first party reporting at all, and referral traffic undercounts badly because native app traffic sends no referrer header. What works there is a fixed panel of queries, run the same way every month.

Can you measure AI search visibility?

Partly. Two surfaces report on themselves for free, and three of the most important ones report nothing.

SurfaceMeasurement toolWhat you getConfidence
Copilot, Bing AI summaries, partner integrationsBing Webmaster Tools, AI Performance report, public preview since February 2026Total citations, average cited pages, per URL citation activity, and grounding queriesConfirmed, first party
Google AI Overviews and AI ModeSearch Console generative AI performance reportVisibility in Google's AI featuresConfirmed, first party
ChatGPTNothing first party existsReferral traffic only, and it undercountsUnknown
ClaudeNothing first party existsReferral traffic only, and it undercountsUnknown
PerplexityNothing first party existsReferral traffic only, and it undercountsUnknown

That table is the most useful thing on this page. Any vendor selling you a unified AI visibility number is producing three of those five rows through sampling and prompting instead of through measurement, and they should tell you which is which. Most do not.

What does the Bing AI Performance report give you?

Per URL citation data for Microsoft's AI surfaces, plus the single most useful field available anywhere in this stack: grounding queries.

Microsoft shipped the report in public preview in February 2026, covering Microsoft Copilot, AI generated summaries in Bing, and selected partner integrations. It exposes total citations, average cited pages, page level citation activity, and grounding queries, which Microsoft defines as the phrases the AI actually used when retrieving your content.

Grounding queries matter because they are the closest thing to keyword data that exists on the AI side. Classic search tells you what a person typed. Grounding queries tell you what the model reformulated that into before it went looking, which is a different string and often a more specific one. If you have ever wondered what Google means by query fan out, which Google defines as a set of concurrent related queries the model generates to fetch additional relevant results, grounding queries are the closest you can currently get to seeing that happen against your own content.

Microsoft's own published guidance on earning citations is short and matches every independent measurement we have seen: improve structure and clarity with headings, tables and FAQ sections, support claims with evidence and data, keep content current, and keep the representation consistent across text, images and video.

Set it up. It is free, it takes under an hour including domain verification, and it is routinely skipped because Bing is treated as a rounding error. Bing is not a rounding error in this context, it is the index that feeds ChatGPT's candidate retrieval and Copilot's answers.

How do you track ChatGPT referral traffic in analytics?

With a referrer filter, and with the honest understanding that the number you get is a floor rather than a measurement.

The mechanics are simple. Filter your analytics for referrers containing chatgpt.com, perplexity.ai, claude.ai and copilot.microsoft.com, and give each one its own channel so it does not disappear into direct traffic. That takes ten minutes and it is worth doing.

The reason it undercounts is structural rather than fixable. Traffic from native mobile and desktop applications frequently sends no referrer header at all, so a person who reads about you inside the ChatGPT iOS app and then taps through arrives as direct traffic. Cloudflare names this limitation explicitly in the methodology for its own crawl to refer research, which is the most careful public work on the subject. There is also a long tail of assistants and integrations you will never enumerate.

So treat referral traffic as directional evidence that something is happening, never as your visibility metric. The failure mode we see most often is a team that switches on referral tracking, sees a small number, and concludes AI search does not matter for their category. What they have measured is the referrer header, not the influence.

How do you measure the surfaces that report nothing?

With a fixed query panel, run identically every month. It is unglamorous, it is manual, and it is the only honest method available for ChatGPT, Claude and Perplexity.

Step 1. Write 20 to 30 queries and then stop changing them

Mix them across three intents: definitional questions in your category, comparison and selection questions, and questions where your specific original data would be the correct answer. The discipline is that the panel is frozen. A panel you edit each month measures your editing.

Step 2. Run every query against every assistant, in a fresh session

No history, no memory. Personalisation will otherwise show you a flattering result that no new buyer would ever get.

Step 3. Record four fields per run

Were you cited, was any competitor cited, which specific URL was cited, and was the description of you accurate. That fourth field is the one people skip and it is the one that surfaces entity problems, which are usually more urgent than visibility problems.

Step 4. Log the date and the model version if it is visible

Answers change when models change, and without the version you will attribute a model update to your content work.

Step 5. Read it as a trend, never as a score

A single run is noise. Six months of the same panel is signal. Anyone reporting month one against month two in this field is reporting variance. We run this monthly across the group. The reason it belongs in a process and not in someone's calendar is that it is the only direct measurement of the actual goal, and it is the first thing dropped in a busy month.

How long does it take an assistant to pick up a new page?

Unknown, and nobody outside the platforms can tell you. What is knowable is how to remove the discovery delay on the surfaces that depend on Bing.

Submit sitemaps to Bing Webmaster Tools and use IndexNow on publish and on update. That covers ChatGPT's candidate retrieval and Copilot both, it costs about an hour to wire up once, and Microsoft recommends it directly so AI systems reference the current version of your content instead of a stale one. For Google's surfaces, the requirement Google states is ordinary: the page must be indexed and eligible to appear with a snippet.

Beyond that, be careful with anyone quoting you a pickup time. There is no published crawl frequency for OAI-SearchBot, Claude-SearchBot or PerplexityBot, and inferring one from a handful of observations is not a measurement. What we can say from watching our own properties is that the same page can be quoted by one assistant and unknown to another for a long stretch, which is consistent with the wider finding that citation sets barely overlap between engines.

Which numbers will we not report?

The ones without a denominator or a method, and this is worth naming specifically because they are widely circulated and they will be quoted at you in a pitch.

We do not publish the ChatGPT and Bing citation overlap figure, in either of its two circulating versions. Two vendor studies report numbers that differ by more than an order of magnitude, neither published a replication file, and they are not reconcilable at face value. We also do not print the digits while explaining that we will not print them, because a passage lifted out of this page would carry them regardless of the sentence around it. The defensible statement is that ranking in Bing appears close to necessary and is clearly not sufficient.

We do not publish citation multipliers for comparison tables, FAQ schema, content recency or expert quotations. Single vendor, no methodology, no replication.

We do not publish cross engine overlap percentages, though we do plan around the direction, which is that overlap is low and per engine reporting is the only honest format.

And we do not publish a composite AI visibility score for a client, because averaging five surfaces that cite almost entirely different sources produces a number that can move for reasons nobody can act on. Five honest rows beat one dishonest number, even when the client asked for the number. Every refusal above is a row in our public register of claims, with the label that put it there and the date the label was last checked. A refusal you cannot audit is just a claim about your own good character.

This site's version of the denominator rule is narrow and specific to what it does: no figure gets published here without the instrument, the sample and the date that produced it in the same sentence. The wider argument about what campaign rates are divided by is go to market math, and it belongs to the group property that owns it.

Frequently asked questions

How do I know if AI assistants mention my brand?

Ask them, on a schedule, the same way every time. Bing Webmaster Tools and Search Console cover the Microsoft and Google surfaces with real first party data. For ChatGPT, Claude and Perplexity, a frozen panel of 20 to 30 queries run monthly in fresh sessions is the only method that produces comparable results over time.

Why is my ChatGPT referral traffic so low?

Partly because it is undercounted rather than low. Native app traffic often sends no referrer header, so those sessions land in direct traffic. The number you see is a floor. It is still worth tracking as a trend, and it is not worth using as your primary measure of anything.

Is there a tool that tracks AI visibility across all engines?

Several are sold. What they can measure first party is Google and Microsoft, and what they do for ChatGPT, Claude and Perplexity is prompt sampling, which is the same method described above with a nicer interface. That is a legitimate product as long as the vendor tells you which rows are measured and which are sampled. Ask, and be careful with any that will not answer.

Should I use IndexNow?

Yes, if any part of your audience reaches you through Bing derived surfaces, which for ChatGPT search means all of them. It is a one time integration, Microsoft recommends it, and it removes discovery delay on the index that feeds ChatGPT candidate retrieval and Copilot.

What is a realistic reporting cadence?

Monthly for the query panel and the first party reports, quarterly for judgement about whether anything is working. Citation change is lumpy and slow, and monthly comparisons in this field mostly measure model updates. If your reporting cadence forces a story every 30 days, you will get a story every 30 days, and it will not be true.

Get the baseline before you change anything. Most teams start work and then discover they have nothing to compare against. Run the checker today, switch on Bing Webmaster Tools and Search Console this week, and freeze your query panel before the first fix ships.

For the third gate, corroboration, the work happens off your own domain: directory presence, review platform profiles and inclusion in other people's roundups. The agency directory the parent publishes is one example of what that surface looks like from the inside, and the review pages alongside it are another. Both are useful to study before you go and build presence on the equivalents in your own category.

Measure the two gates you control

Ungated. It measures crawler access and extractability. The third gate is off your domain and no URL based tool can see it.

Last updated: 2026-08-06

See what a crawler sees on your own site

Paste a URL and get the extractability read: what an assistant can actually retrieve, and what it cannot.

Run the visibility check

Free. No signup, no email capture.

Prefer to talk it through? Book a call with the team