Most guides on tracking AI brand mentions point you toward manual prompting or one of the growing number of AI visibility tools. We tried both before we built Traqer and found neither to be a complete solution: manual tracking isn't realistic at scale, and most tools share a measurement problem worth understanding before you commit to one.
The default metric across most tools out there right now is a single blended visibility percentage: one number that shows how often your brand appears across AI-generated answers. This is often displayed front and center in the form of a graph, showing changes in this metric over time.
This might sound helpful, but in practice, it a) gives you a skewed view of your actual visibility in different platforms due to the blending and b) hides the details you need to make decisions. To understand where your brand is actually showing up across ChatGPT, Perplexity, Google AI Overviews, Gemini, and Claude, you need a different approach to measurement.
At our agency (Grow & Convert), we track AI visibility for our client, and we built Traqer because the tools available weren't measuring it the way we think it should be measured. This article explains what that actually means in practice:
Where common tracking approaches fall short
What topic-based tracking looks like
How to use tracking data to understand and improve your AI search visibility
The typical approaches to tracking brand visibility across AI (and why they don't work)
Why manual tracking isn't worth doing
In theory, you could track your AI brand mentions manually. In practice, it's not realistic. To do it properly, you'd need to build a large bank of prompts, run them across multiple AI platforms, record every brand and competitor mention, track rankings, sentiment, citations, and positioning, then repeat the whole process regularly to identify trends over time.
Even if you were willing to invest that effort, the results would still be unreliable. AI answers are inconsistent from one query to the next, and they’re personalized to each user in ways a prompt bank you run yourself can’t capture.
For most marketing teams, manual tracking is both too time-consuming and too inconsistent to produce meaningful data. Using a dedicated AI visibility platform is the only realistic way to monitor AI brand mentions at scale.
Why most AI visibility tools fall short
An AI visibility tool will solve the tracking problem, but, as mentioned above, most have two measurement issues worth knowing before you commit to one.
The first is the single visibility score. Most tools report one percentage that shows how often your brand appears across AI-generated answers. But AI visibility isn't a single number. It exists at the topic level. Your brand might be recommended frequently for one topic and rarely for another, even when both matter to your business. One percentage can't tell you that.
A rising overall score doesn't necessarily mean visibility is improving where it matters most. Strong visibility in one area can mask weaker visibility in another. And the metric itself is easy to manipulate: just delete the prompts where you're not appearing and your percentage improves, even though your actual brand visibility in the real world hasn’t changed.
The second problem is prompt-level rank tracking. Some tools report your brand’s “position” for a given prompt, as if AI search returned a fixed results page you could rank on. It doesn’t, and there are two reasons why.
The first is that you can’t predict the prompts your customers actually type. People don’t query AI tools with short, repeatable keywords the way they search Google. They have long, contextual conversations, and each one is pretty much one of a kind. The prompt a tracking tool monitors is rarely the prompt a real customer would use, so a ranking for it tells you little about what people see. (Some tools claim to have data on the real prompts people type, drawn from millions of chats. That data comes from people who’ve opted into having their browsing tracked, and are unlikely to resemble your actual buyer. Also it can tell you what some people typed, but as we’ll explain next, not the heavily personalized answer they actually received.)
The deeper reason is personalization. Even if you could guess the exact prompt a real customer used, you wouldn’t reproduce the answer they got. The model factors in everything it knows about that person: their industry, their company size, what they’ve tried before, and the context built up across months of conversations. That hidden context is often what decides whether you get recommended. The literal prompt someone types is short; the effective prompt the model actually answers is far longer and specific to them.
We call these invisible prompts, and it’s the main reason a “position 2” ranking for a single prompt is a false sense of precision. We tested it directly by taking the exact prompt a real lead had used, the one where ChatGPT had recommended us, and ran it again in a fresh, logged-out session. The answer came back completely different, and didn’t mention us at all. Even setting personalization aside, the same prompt produces different answers from one run to the next. Rand Fishkin’s research at SparkToro found that Claude would need to be asked the same question 1,429 times before producing two answers with the same brands in the same order.
Tracking any single prompt, then, gives you a single data point rather than a picture of your visibility. These limitations encourage us to think differently about AI visibility measurement.
(For more detail on where standard tools fall short for agencies specifically, see our article on AI visibility tracking tools for agencies.)
Solution: Track brand mentions by topic for a more accurate picture of your AI visibility
If overall visibility scores obscure the detail you need, and prompt-level rankings are unreliable, the answer is to track visibility by topic.
By grouping your related prompts together and measuring how often your brand appears across all of them, you’ll get a much clearer view of where AI platforms see your brand as relevant. It also makes the data actionable: instead of watching a single percentage move up and down, you can see exactly which topics you're strong in, and which ones need some work.
(Also read: Topic-Based GEO: A Content Strategy that Gets LLMs to Recommend Your Brand)
This is the approach we built Traqer around. It's one of the only AI visibility tracking tools built specifically to measure brand visibility by topic, per LLM, with brand mentions and citations reported separately. The metrics it tracks include:
Visibility by topic: The percentage of prompts within each topic where your brand appears (50%+ indicates strong visibility)
Topic coverage: The number of topics where your brand is visible
Prompt opportunities: Suggested prompts to monitor and improve visibility
Citation frequency: Which pieces of content are cited and how often
Prompt-level visibility: The specific prompts where your brand is mentioned or recommended
LLM coverage: Which AI models include your brand in their responses
Competitive visibility: How your visibility compares to competitors across the same topics
Visibility trends: How your visibility within each topic changes over time
The sections below explain how to use each of these in practice.
How to more accurately track your brand mentions in AI search with Traqer
One thing to flag before going further: most AI tracking tools pull data from LLM APIs rather than the real web interfaces. This matters, because API responses differ materially from what actual users see. LLM web products layer system prompts, interface tuning, and model configurations on top of the base model, which changes which brands appear and how responses are structured.
Traqer runs queries through real browser sessions across ChatGPT, Claude, Perplexity, Google AI Overviews, and Gemini, capturing what a user would actually see.

A logged-out, context-free session can never perfectly replicate a real user's personalized experience, but it's a meaningful step closer to reality than an API call.
Understand how visible you are across your chosen topics
Traqer measures not just whether your brand appears for a specific prompt, but how much coverage you have across a whole topic. That's a better indicator of whether you'll show up reliably when users ask related questions.
To get started, you'll need 5 to 10 topics that matter to your business. We recommend starting with bottom-of-funnel topics; buying-intent areas where the user is actively looking for a product or service, rather than just researching a concept. The more specific the topic is to what you actually sell, the better your chances of being mentioned.
For our content marketing agency, Grow & Convert, we track topics like SaaS content marketing agencies and B2B content marketing agency. We wouldn't track content marketing on its own. It's too broad, and the people searching that term are mostly looking for information rather than a specific service provider.
Once you define your topics, Traqer generates five prompts per topic that approach it from different angles. You can also add your own. Traqer then gives you a visibility percentage for each topic. 50% and above indicates strong visibility, meaning your brand is likely to appear no matter how a user phrases their question within that area.
Understand which prompts are most likely to mention your brand
Not all prompts are worth tracking. Some questions, particularly informational ones, produce answers that don't mention brands at all. Tracking those is noise.
Traqer's Brand Mention Probability feature shows the likelihood of each prompt generating a response with brand recommendations, rated high, medium, or low. Start with high-probability prompts only. Once you have solid visibility across those, you can expand into medium-probability ones.
This is a more efficient use of your tracking budget and produces cleaner data for reporting. For a broader discussion of why bottom-of-funnel prompts are the only ones worth prioritising, see our article on AI search content strategy.
The multi-prompt approach also matters for reliability. When we tracked Level AI's visibility for the topic "AI call monitoring solutions," Traqer ran the same topic through six differently-worded prompts. Brand mention counts varied across them. This is exactly the point: AI search isn’t reproducible the way traditional search is, and a single prompt gives you a single data point, not a picture of your real world visibility. Averaging across multiple phrasings of the same topic produces a number that more accurately represents the likelihood of you being mentioned when a real user asks an LLM about that topic.
Distinguish brand mentions from citations
Most tools combine brand mentions and citations into a single visibility number.
But they're different things, and should be reported separately.
A brand mention is when an LLM names your brand as a recommendation: "for that use case, you might want to look at [your brand]." That's what drives leads. A citation is when an LLM uses content from your site to inform its answer, but doesn't necessarily name you as a recommendation. Your URL appears as a source, but your brand may not appear in the actual response at all.
To use a concrete example: we work with a client called Toro, a trucking software company. When an LLM cites their article on "The 13 best trucking companies for small business" as a source but doesn't list Toro in the response itself, that's a citation without a brand mention. Being cited as a reference is useful as an indicator that your content is influencing LLMs, but is far less likely to drive actual leads compared to a brand mention.
Traqer tracks both separately. You can toggle between brand mentions, citations, or both at the top of every view. The distinction matters for understanding what your visibility data actually means.
We saw this play out clearly with our client, Level AI. For the topic "call center real time reporting," Traqer showed strong brand mention visibility across ChatGPT, Perplexity, and Google AIO. It also showed that the article G&C produced on that topic was cited across those 6 prompts a combined 15 times; nearly twice as often as any competing article.

Two different signals, both useful, but for different reasons: the citation data told us the content was working as an LLM source; the brand mention data told us Level AI was actually being recommended to users.
Analyse your overall brand presence across LLMs
Beyond topic-level data, Traqer consolidates your performance across all topics into a company-level view. This is useful for leadership reporting: it lets you quickly answer questions like which articles are getting picked up across LLMs, which topic areas have the most coverage, which LLMs are most likely to mention you, and how you compare to competitors.
Agencies find this view particularly useful for toggling between clients without losing context. It gives a quick summary without flattening the topic-level data that makes the analysis useful in the first place.
Understand how you stack up against your competitors
Traqer shows which brands are mentioned most often across your chosen topics and prompts, with mention counts and citation counts broken out by platform. You can click into a specific competitor to see which prompts they appear in and which they don't, which tells you where there are gaps in your own coverage. This isn't a vanity exercise; knowing that a competitor has strong visibility on Perplexity but weaker visibility on ChatGPT for the same topic points directly to where content investment can pay off.
See how your visibility changes over time by platform
Visibility data at a single point in time tells you very little. The value comes from tracking changes week over week. The clearest way to do this in Traqer is to view trends by topic rather than by overall brand-level visibility. The overall LLM visibility chart tends to look sporadic, with fluctuations that don't map to anything meaningful. By topic, patterns become readable: you can see which topics are improving, which have stalled, and click into a specific week to understand what changed.
To start tracking your own brand mentions, start for free here.
How to improve your AI brand visibility: topic-based GEO
Tracking is one side of the problem. The other is knowing what to do with the data.
SEO improvement is relatively mechanical: find a keyword with volume and intent, create content that meets the search intent, build links. At some point, the content has a good chance of ranking in the SERPs. You know where you stand, and usually there are clues about what to do next.
GEO is generally harder to act on, and the reason comes down to how AI search actually works. Traditional SEO optimizes for a fixed results page; users enter a variation of similar short queries and see broadly similar results, depending on what Google defaults to. But AI search has no fixed page. Every user is having their own personalized conversation, shaped by prior interactions, stated preferences, and context the LLM has built up over time. The same prompt asked by two different people will produce different responses. While Google customizes somewhat, LLMs are at a different level.
This means you can't create one piece of content targeting one prompt and expect that to drive AI visibility the way a single article might rank for a keyword. LLMs need a broader body of evidence to consistently recommend your brand: content that covers who you are, who you serve, what problems you solve, and how you're different from the alternatives. We call this your GEO Topic Map.

Here's how Grow and Convert approaches GEO for clients:
Map out every product angle that exists. This includes categories, use cases, competitors, value propositions, target customers, and the specific pain points you solve. This can't come from existing marketing collateral alone. You need to interview the people who know the product and customers best: the founders, the sales team, and/or customer service. The goal is to surface the specific, differentiated details that make your product the right answer for specific situations.
Produce highly specific content around each angle. The content that works in AI search is not beginner-level educational material. LLMs need the kind of detail you'd get in a product demo: specific features, use cases, outcomes, and comparisons. Generic "what is X" content won't move the needle. If those product interviews surface keywords with search volume, you can create content that targets both traditional SEO and AI search simultaneously. For more on why that connection still holds, see our article on Prioritized GEO.
Produce context-rich case studies. First-hand customer stories with specific details, the pain points they faced, the alternatives they considered, and the outcome after using your product, give LLMs the contextual evidence to connect your brand to users in similar situations. Thin case studies don't do this. The more specific the narrative, the more useful it is.
Use Traqer data to inform what to produce next. The Analyse & Improve view within each topic shows which content types LLMs are most often citing for that topic. If the cited articles are predominantly comparison posts or list articles, that tells you something specific about what to create. Try it free here.
Get started with Traqer for free
Traqer tracks your AI brand mentions across ChatGPT, Perplexity, Google AI Overviews, Gemini, and Claude, by topic, per LLM, with brand mentions and citations reported separately. Start your free 7-day trial or schedule a demo to see how it works.
