All articles

AI search visibility metrics: what should you track?

Devesh KhanalDevesh KhanalAugust 24, 202611 minutes read
AI search visibility metrics: what should you track?
Share

AI search has created a new measurement problem. Traditional SEO is organized around ranking. A page holds a position for a keyword, and you track that position and the traffic it sends, on the assumption that the slot is stable enough to occupy and defend. AI search doesn’t work that way. An AI answer usually resolves the query in the response itself, often with no click, and there's no fixed position to hold, because the same question can return a different set of brands each time it's asked. 

So the useful unit of measurement shifts from ranking position to visibility, meaning how often a brand appears when people ask about its category. A set of metrics has grown up around that idea, and most guides present them as a menu: e.g., brand mention rate, citation share, share of voice, a headline visibility score, sentiment, and so on.

We built Traqer because we run GEO/AEO for dozens of clients and need to measure results. What we found is that the most vital question is which of these metrics are most useful to measure over time. Several commonly-tracked metrics move for reasons that are not within your influence and have nothing to do with your actual visibility, while a headline percentage can look healthy while telling you very little. 

This guide covers the AI search visibility metrics in common use, what each one measures, and how to read them without fooling yourself.

What AI visibility metrics are measuring

Almost every metric in this category comes down to two underlying insights: whether an AI answer names your brand (a brand mention), and whether it links to your site as a source (a citation). These are separate things, and the difference is important for your measurement and reporting. 

A brand mention means the model names you in its answer, for example listing your product when someone asks for recommendations in your category. This is the outcome that drives leads, because the user sees your name in the response itself. 

A citation means your URL appears as a source link the model drew on. Citations are useful as a foundation signal and can be used as an outreach map, since they show which pages the model is pulling from for a topic. But a citation on its own is not a recommendation.

In their default view, most AI visibility tools lead with a combination of mentions and citations in a single visibility figure. That inflates the overall number and hides which of the two is actually happening. 

A brand that is cited often but rarely named has a different problem from one that isn't appearing at all, and a blended score can't tell them apart. In Traqer, we default to the brand mention view because this is the most valuable metric you can track, and there’s a toggle at the top of every page for brand mentions only, citations only, or both, so you can read each independently (or together, if you prefer).

We go into the distinction in more detail in our guides to tracking brand mentions in AI search and LLM citation tracking.

Why one blended score across models misleads

A single visibility percentage averaged across every model is one of the most common headline metrics people use to report on content performance in LLMs, yet it is one of the least useful. The various engines behave differently enough that averaging them into one number hides the thing you need to act on.

Perplexity and Google's AI Overviews work largely as search summarizers, so a brand with strong traditional SEO tends to have decent visibility there. ChatGPT leans more heavily on its training data and its own browsing behavior, so strong SEO doesn't seem to influence its answers as much. Read as one blended figure, a brand that looks strong on Perplexity and weak on ChatGPT shows up as a middling average that describes neither. The independent research points the same way. An analysis by Ahrefs found that Google's AI Mode and AI Overviews cite different sources about 87% of the time for the same query, so even two features from a single company diverge widely.

For this reason, Traqer reports visibility per LLM and doesn’t blend the models by default. We cover the model-by-model view in our AI visibility tool overview, and some of the individual LLMs in our Perplexity rank tracker, AI Overviews tracker, and AI Mode tracker articles. 

The metric that moves when you change what you track

The single visibility percentage has a further problem beyond blending. The figure is the share of your tracked prompts where your brand appears, which means it moves whenever you change the set of prompts, not only when your visibility changes.

Add ten ambitious prompts you don't yet appear for, and the percentage drops. Remove the prompts where you're weakest, and it climbs. Neither action reflects anything that has actually happened in the models. This creates an incentive to avoid tracking the hard topics, or to quietly prune the weak ones before a client or leadership report, and the headline number ends up rewarding the behavior you least want.

This is where a percentage on its own stops being a dependable measure of progress. It works as one view among several, but if it's the only number you watch, you can show improvement without anything having improved.

Metrics that only move when something real happens

To get around the gameability problem, Traqer reports two brand-level metrics alongside the percentage:

  • LLM visibility count. The raw number of prompts where your brand appears, per model. Because it counts appearances rather than a proportion, adding new prompts can't drag it down. It rises only when you genuinely appear somewhere that you didn't before.

  • Topic visibility. The number of topics where you have high (over 50%) or some (over 0%) visibility across the models you've selected. An ambitious new topic doesn't reduce your existing scores, so the metric improves only when your coverage genuinely widens.

Both are straightforward to put in front of a client or an executive, because they move in one direction and only when something has actually changed. That's a large part of why we recommend tracking visibility over time rather than as a single snapshot, which we cover in how to track ChatGPT rankings over time.

Why we measure topics rather than single prompts

Several tools still report a “ranking position” for an individual prompt, as though an AI answer were a stable search result you could hold a spot in. It isn't, and the evidence on this is now fairly clear.

Research published by SparkToro ran 2,961 prompts across ChatGPT, Claude, and Google's AI Overviews and AI Mode. It found less than a 1-in-100 chance of getting the same list of brands back from two runs of the same prompt. The same study noted that how often a brand appears across many related prompts is far more stable than its position in any single answer.

There are two reasons one prompt is unreliable. On conversational platforms like ChatGPT, answers are shaped by chat history, account context, and the model's own randomness, so no two users see quite the same thing and no single phrasing can be reproduced. On search-based services like Google’s AI Overviews, the model fans a query out into several underlying searches, so even a fixed wording pulls from a shifting set of sources. Either way, a single prompt is a noisy sample.

Traqer organizes tracking around topics, each containing several prompts that approach the same buying-intent question from different angles, and reports the percentage of those prompts where you appear, per model. That gives a signal you can read week to week instead of the noise of a single response.

How to understand share of voice without overprioritizing it

Share of voice, which is sometimes called “answer share of voice” is your brand’s portion of total category visibility measured against a set of competitors across the same tracked prompts. It’s a useful metric for reporting to senior management because it describes broader market position rather than only your own progress, and it can explain a situation where your own visibility is rising but the business still feels stuck, because competitors may be gaining faster.

The caveat is that share of voice inherits every limitation mentioned above. It’s only as sound as the prompt set that it's built on, it should be read per model rather than blended, and different tools weight it in different ways. All this means that scores from separate products aren’t really comparable. 

Treat it as a directional read on where you stand in the category rather than a precise ranking.

Brand mention probability: which prompts are worth tracking?

Not every prompt query generates a response where a brand is likely to appear. When someone asks an informational question like “what is project management software?” the model will usually answer directly and probably won’t recommend any brands in particular, so tracking that prompt doesn’t tell you much about your brand visibility. The prompts that matter are the buying-intent ones, where a user is asking for a recommendation and the model responds with a list of companies, products, or services.

Traqer's Analyze and Improve view gives each prompt a brand mention probability of high, medium, or low, based on how many brands the models list for it. Here’s what it looks like:

A low-probability prompt with few brands is usually informational noise you can drop and replace with something more product-centric. This keeps your tracked set focused on the prompts where visibility is achievable, which in turn keeps every metric above more meaningful.

The AI search metrics we don’t emphasize (and why)

A few metrics you'll see elsewhere demand less emphasis. The first is the “AI readiness” or “site health” score that some tools report. These are generally built from on-site technical signals such as an llms.txt file, question-style headings, FAQ blocks, or AI-specific schema. 

We've tested those signals across our client base and measured no meaningful effect on visibility. Put simply: focusing too much on this distracts from the question of whether your brand is getting mentioned in answers. Rather, these metrics describe how cleanly a model might parse a page once it arrives. It’s not that this data is completely useless; but it’s not what should be tracked as a priority. 

The second is “sentiment” and “accuracy” scores. Granted, it does matter how an LLM describes you, and a confidently wrong description is a reputational issue to be aware of. The difficulty here is that these insights are drawn from answers that change from one run to the next. One reading isn't a trend. So, Traqer surfaces sentiment as context alongside the mention count rather than as a headline number, and we treat an inaccurate description as an event to correct rather than a metric to optimize continuously.

One caveat applies to all of these numbers, our own included. Traqer measures a neutral, logged-out user, captured from the actual web interface rather than an API response. That gives a baseline that stays consistent from one check to the next, which is what lets a trend mean something. But it isn't a fully personalized user with their own chat history and account context. No tool can access that.

How to approach AI visibility metrics, and how they fit together

In summary, this is how we recommend approaching your AI visibility metrics: 

  • Prioritize brand mentions as the most important metric; treat them as the outcome you’re optimizing for because they are most likely to generate customers for your business.

  • Separate brand mentions from citations, and review them independently.

  • Read all your metrics per model, because each of the LLMs behaves differently and relying on a single average visibility score doesn’t tell you much about your strategy. 

  • Focus on topic-level visibility rather than tracking single-prompt performance, so you’re measuring a pattern.

  • Keep your visibility percentage as one view, but rely on LLM visibility count and topic visibility when reporting progress, because they move only when something has genuinely changed.

The numbers, viewed in this context, will tell you where to act: which topics to produce content for, which models you're weak on, and which pages the AI already trusts. For the deeper strategy behind improving those numbers rather than only tracking them, our Toro TMS and Constitution Lending case studies show the approach was applied by Grow and Convert to boost client LLM visibility.

Traqer is built by Grow and Convert. For more on the strategy behind improving AI visibility, Topic-Based GEO and Prioritized GEO explain the thinking behind how Traqer was built. You can try Traqer for free here.