How to track ChatGPT visibility

Tarun Kapoor, founder of Context Hints, seated at a wooden desk with a soft city light behind him.Tarun Kapoor Updated July 25, 2026 12 min read

You cannot track ChatGPT visibility in analytics, because most AI mentions produce no click and therefore no referrer — the brand is named, the user absorbs it, and nothing reaches your server. Measuring it means sampling the conversation from the outside: presence rate across a fixed prompt panel, citation share, description accuracy, and the referral traffic that does arrive. This guide covers the free methods, the first-party report almost nobody uses, the tool landscape, and the cadence that keeps the numbers meaningful.

The 2-sentence answer

You cannot track ChatGPT visibility in Google Analytics, because most AI mentions produce no click and therefore no referrer — the brand is named, the user absorbs it, and nothing reaches your server. Tracking it instead means measuring four things directly: how often you appear across a fixed set of buying prompts, what share of citations you hold, whether you are described accurately, and what referral traffic does arrive — using a mix of free first-party reports and dedicated monitoring tools.

The short version

  • Most AI visibility is invisible to analytics. A mention without a click leaves no trace on your site.
  • Measure four things: presence rate, citation share, description accuracy, and referral traffic.
  • The prompt set is the instrument. Get it wrong and every number downstream is meaningless.
  • Bing Webmaster Tools publishes free AI citation data — page-level citations and the grounding queries behind them.
  • Tools start at $29/month and run to enterprise; the differentiator is engine coverage and prompt volume, not dashboards.
  • Segment by query stage. Aggregate visibility hides the commercial queries that actually matter.

Why analytics can't see AI visibility

The instinct is to look in Google Analytics for a ChatGPT referral source. That will show you something, but it will systematically understate your actual visibility — often by an order of magnitude — for a structural reason.

When a model names your brand in an answer, three outcomes are possible. The user clicks a citation link, in which case you get a referral. The user reads the recommendation and searches your brand separately, in which case the visit is attributed to organic or direct. Or the user simply absorbs the information and acts later, or not at all — in which case nothing whatsoever reaches your servers.

The second and third outcomes dominate. Zero-click behaviour was already the norm in classical search — clickstream analysis put roughly 68% of Google searches ending without a click in early 2026 — and a conversational interface that answers the question directly increases that share rather than reducing it.

So the measurement problem is not that tracking is badly configured. It is that the event you care about happens somewhere you have no instrumentation: inside a conversation you cannot see. Every method below is a way of sampling that conversation from the outside.

The reframe

Stop asking "how much traffic does ChatGPT send us?" and start asking "when our buyers ask the questions that matter, how often are we in the answer, and what share of the answer is ours?" The first question measures a leaky side-effect. The second measures the thing itself.

The four things worth measuring

MetricQuestion it answersHow to get it
Presence rateIn what % of our target prompts are we mentioned at all?Prompt panel, run repeatedly
Citation shareOf all sources cited on those prompts, what share is ours?Monitoring tool or Bing AI reports
Description accuracyWhen named, are we described correctly and favourably?Manual review of captured answers
Referral trafficWhat arrives when someone does click?Analytics, segmented by AI referrer

These are deliberately ordered. Presence rate is the primary metric because it maps directly to consideration-set membership — models name only a handful of brands per answer, so being present at all is the binary that matters. Citation share refines it: appearing once among eight sources is weaker than appearing twice among four.

Description accuracy is the one teams skip and later regret. Being named with an outdated price, a discontinued feature, or a competitor's positioning is worse than not being named — and it is invisible unless someone reads the answers rather than the dashboard.

Referral traffic sits last deliberately. It is real, it is the easiest to measure, and it is the least representative. Treating it as your AI visibility metric is the single most common analytical error in this area. For the full metric definitions, see ChatGPT visibility metrics.

A translucent blue glass measuring instrument on a white field with a single fine needle resting against a delicate arc of hairline gradations.
Presence rate and citation share are the two readings that matter. Referral traffic measures only the clicking minority.

Building the prompt set (the hard part)

Every method here depends on a fixed set of prompts you test repeatedly. The prompt set is the instrument, and a badly constructed one produces confident numbers about nothing.

Four rules make it trustworthy:

  1. Write prompts a buyer would actually type, not queries a marketer would. Real prompts are long, conversational, and full of constraints: "we're a 12-person dental practice and our front desk is drowning in appointment calls — what should we look at?" Keyword-shaped prompts produce keyword-shaped answers that no real user sees.
  2. Cover the whole funnel, then segment. Include early exploratory prompts, mid-funnel comparison prompts, and late-stage specific prompts. Report them separately — aggregate visibility hides the commercial queries where third-party weighting punishes vendor content hardest.
  3. Include prompts you expect to lose. A panel of prompts you already win measures nothing but your own comfort. Deliberately include competitor-favouring and category-generic prompts; those are where the growth is.
  4. Freeze it, then change it deliberately. The panel must stay constant to be comparable over time. When you add prompts, version the set and report the old and new panels separately for a period rather than silently redefining your baseline.

A workable starting panel is 30–50 prompts: roughly 40% evaluation-stage ("best X for Y", "X vs Z"), 30% problem-stage ("how do I solve Y"), 20% category-generic, 10% brand-specific ("is X any good"). Scale up once the workflow is running; precision improves with panel size but so does cost.

Expect variance, and design for it The same prompt can return different brands on different runs. This is inherent to generative systems, not a bug in your process. Two implications: never draw a conclusion from a single run, and always report presence as a rate across repeated runs rather than a binary. Three to five runs per prompt per measurement cycle is a reasonable minimum.

Free methods: what you can do today

Before buying anything, a manual baseline is worth building — partly for cost, mostly because reading actual answers teaches you things a dashboard never will.

The manual prompt audit. Run your panel by hand in a logged-out or temporary chat, capture each answer, and record: were you mentioned, which competitors were mentioned, which sources were cited, and was your description accurate. A spreadsheet with one row per prompt-run is entirely sufficient.

Two methodological cautions. Use a fresh or temporary session so personalisation and memory do not contaminate results — a logged-in account that has discussed your brand before will over-report your presence. And keep location consistent, because answers vary by market.

This is slow: 40 prompts × 3 runs is a couple of hours. But it produces a genuine baseline, it surfaces the qualitative problems tools miss, and it tells you whether paying for automation is worth it before you pay.

The free report almost nobody uses

Bing Webmaster Tools publishes AI citation data for verified sites, and it is the most underused free instrument in this field.

Two reports matter. The AI Page Stats report lists your pages by citation count — how often each was used to ground an AI answer. The AI Search Queries report lists the grounding queries that produced those citations, along with intent classification, topic, your citation count, and — most valuable of all — your citation share on each query.

That last column is the metric everything else in this guide approximates, delivered as first-party data at no cost. It tells you not just that you were cited, but what proportion of all citations for that query were yours — which is the difference between "we appear sometimes" and "we own this question."

Three ways to use it that most teams miss:

The obvious caveat: this is Bing and Copilot grounding data, not ChatGPT's own. Because ChatGPT's search has drawn on Bing's index, it is a reasonable directional proxy — but it is a proxy, and it should be triangulated against direct prompt testing rather than treated as ground truth.

Tracking the traffic that does arrive

Referral traffic understates visibility, but it is worth capturing properly because it is the only layer that connects to revenue.

The monitoring tool landscape

Dedicated tools automate the prompt panel across multiple engines and track results over time. Representative options as of 2026:

ToolEntry pricePositioning
Otterly.ai~$29/monthEntry level; ChatGPT, AI Overviews, Perplexity, Copilot in core plans
Peec AI~€75/month (25 prompts)Mid-market analytics; multi-country tracking, competitor share of voice
ProfoundEnterpriseBroadest engine coverage — around ten platforms including Gemini, Claude, Grok

What actually differentiates them, in order of practical importance:

  1. Prompt volume included. Pricing is usually banded by tracked prompts, and a 25-prompt tier will not cover a serious panel. This is the constraint that most often forces an upgrade.
  2. Engine coverage. If your buyers use Perplexity or Gemini, ChatGPT-only tracking is a partial view.
  3. Geographic coverage. Answers differ by market; single-market tracking misleads international brands.
  4. Competitor benchmarking. Share of voice against a named competitor set is the reporting most teams end up needing.
  5. Raw answer export. Underrated. If you cannot read the actual answers, you cannot audit description accuracy.

A reasonable adoption path: run the manual audit for one cycle to establish a baseline and learn the qualitative picture, cross-check against free Bing AI reports, then buy the cheapest tool that covers your prompt volume and engines. Buying before you have a considered prompt panel usually means paying to automate the wrong questions.

Measuring citation share correctly

Citation share is the most decision-useful metric in this discipline and the easiest to compute incorrectly.

The definition that matters: of all sources the engine cited when answering a given query, what proportion were yours? Not what proportion of answers mentioned you — that is presence rate. Share measures depth of grounding, and it moves independently of presence.

Why the distinction pays: presence rate tells you whether you are in the consideration set; citation share tells you how load-bearing you are within it. A brand mentioned in 80% of answers but supplying 10% of citations is being name-checked while competitors supply the substance — a fragile position that collapses the moment a competitor publishes something better.

Two practical rules. Segment share by query stage, because informational and commercial queries behave differently and a healthy aggregate routinely conceals a weak commercial position. And track the addressable gap, not just your share: citations ÷ share gives the total available, and the difference is the prize. A query where you hold 33% of 470 available citations has more unclaimed ground than one where you hold 60% of 40.

Cadence: how often to measure

ActivityFrequencyWhy
Automated prompt panelWeeklyEnough to see trend without drowning in variance
Manual answer reviewMonthlyCatches description accuracy and tone drift
Bing AI citation reportsMonthlyFirst-party share data; slow-moving
Competitor share benchmarkQuarterlyEntity-level signals move over months, not weeks
Prompt panel reviewQuarterlyCategories evolve; the panel must too — versioned, not silently edited

Resist daily measurement. Run-to-run variance is large enough that daily numbers are mostly noise, and the underlying signals — third-party mentions, corroboration, domain credibility — change on a timescale of months.

Six measurement mistakes

  1. Using referral traffic as the visibility metric. It measures the small clicking minority and misses the majority who never leave the conversation.
  2. Testing in a logged-in account. Personalisation and chat memory inflate your own presence. Use temporary or logged-out sessions.
  3. Drawing conclusions from one run. Generative variance is real. Report rates across repeated runs.
  4. Measuring only prompts you win. Comfortable panels produce comfortable numbers and no growth.
  5. Reporting aggregate visibility. It hides the commercial-query weakness that matters most, because informational wins mask evaluation-stage absence.
  6. Editing the prompt panel silently. Every historical comparison becomes invalid. Version it.

Turning measurement into action

Measurement only earns its cost if it changes what you publish. The routing is fairly mechanical once you have presence rate and citation share segmented by query.

What the data showsDiagnosisAction
Absent entirely, no relevant pageCoverage gapBuild a page answering that sub-query directly
Page exists, indexed, never citedSelection failureRestructure for extraction; add original data
Present informationally, absent on "best X"Entity credibility gapThird-party mentions, reviews, comparisons
High presence, low citation shareName-checked, not load-bearingPublish the specific, citable substance others lack
Named but described wronglyStale corroborationCorrect third-party sources; publish canonical facts
Share fallingDisplacementIdentify who took the slot and what they published

The third and fourth rows are where most B2B brands actually sit, and both are fixed off your own domain rather than on it. Why that is, and what to do about it, is covered in why ChatGPT recommends your competitor. For the mechanism underneath all of this, see how brands appear in ChatGPT responses.

Frequently asked questions

How do I track ChatGPT visibility?

Measure four things directly: presence rate across a fixed panel of buying prompts, citation share on those prompts, whether you are described accurately, and referral traffic. Analytics alone will not work, because most AI mentions never produce a click and therefore leave no trace on your site.

Can I see ChatGPT traffic in Google Analytics?

Partially, and it understates reality badly. You will capture the minority who click a citation link, but not the users who read your brand name and search for you separately, and not those who simply absorb the information. Build a dedicated AI channel group, then treat the number as a floor rather than a measure.

Is there free data on AI citations?

Yes. Bing Webmaster Tools publishes AI citation reports for verified sites — an AI Page Stats report showing citations per page, and an AI Search Queries report showing the grounding queries, their intent, your citation count, and your citation share on each. It is the most underused free instrument in this field.

What is citation share?

Of all the sources an engine cited when answering a given query, the proportion that were yours. It is different from presence rate, which measures whether you were mentioned at all. Share tells you how load-bearing you are within the answer — a brand mentioned often but supplying few citations is being name-checked while competitors supply the substance.

How many prompts should I track?

A workable starting panel is 30–50, weighted roughly 40% evaluation-stage, 30% problem-stage, 20% category-generic and 10% brand-specific. Include prompts you expect to lose — a panel of questions you already win measures nothing but your own comfort.

Why do I get different answers to the same prompt?

Run-to-run variance is inherent to generative systems, not a fault in your process. Never draw a conclusion from a single run: report presence as a rate across repeated runs, with three to five runs per prompt per cycle as a reasonable minimum.

What are the best AI visibility tracking tools?

Otterly.ai starts around $29/month for entry-level monitoring; Peec AI sits in the mid-market from about €75/month with multi-country tracking and competitor share of voice; Profound targets enterprise with the broadest engine coverage. The differentiators that matter are included prompt volume, engine coverage, geographic coverage, and whether you can export raw answers.

How often should I measure ChatGPT visibility?

Automated prompt panels weekly, manual answer review and Bing AI reports monthly, competitor benchmarking and panel revision quarterly. Avoid daily measurement — run-to-run variance makes daily numbers mostly noise, and the underlying signals move over months.

Should I test prompts in my own logged-in ChatGPT account?

No. Personalisation and chat memory will inflate your own brand's apparent presence, because the system has seen you discuss it before. Use temporary or logged-out sessions, and keep location consistent, since answers vary by market.

Does Bing AI citation data reflect ChatGPT?

It is a reasonable directional proxy rather than ground truth. The reports cover Bing and Copilot grounding, and ChatGPT's search has drawn on Bing's index — but they are different systems. Triangulate the free Bing data against direct prompt testing rather than relying on either alone.

Sources and further reading

Want a visibility baseline built for your category?

30 minutes with Tarun. We will design your prompt panel, establish a presence-rate and citation-share baseline, and show you where the unclaimed volume sits.

Book a discovery call
Tarun Kapoor, founder of Context Hints, seated at a wooden desk with a soft city light behind him.
Tarun Kapoor
Founder & CEO, Context Hints

Twelve years of media buying across GroupM, WPP, Ogilvy & Mather, and Neil Patel Digital. Has personally owned media for Nestlé, Sage, Qualcomm, Aetna, Weight Watchers, Chubb and Novotel.