A citation in AI search is your page showing up as a linked source beneath an answer in an LLM. It's a sign that these tools are drawing on your content when they answer buyers, and because it rarely surfaces in your normal analytics (most LLM sources don’t get clicked), tracking it directly is the only way to know it's happening. Doing that reliably is harder than it looks. Citations shift from one run to the next, they're easily confused with brand mentions, and most tools report them in a way that isn't helpful.
The obvious first move, checking by hand, doesn't get you far: the sources behind an answer change every time you re-run the prompt, so a manual spot-check is out of date almost immediately, and doing it across every engine and query you care about isn't realistic.
We ran into this ourselves at our agency, Grow and Convert, while tracking AI visibility for our clients, which is why we ended up building Traqer, an AI visibility tracker of our own. In this post we'll cover how citations differ from brand mentions, why they're harder to track than search rankings, which metric to track instead, how to set up citation tracking in practice, and what to do with the data once you have it.
Citations and brand mentions are different signals
Before you track anything, separate two signals that most AI visibility tools collapse into a single number: citations and brand mentions.
A citation is your URL used as a linked source behind an answer.
A brand mention is the model naming you in the recommendation itself, as in “QuickBooks is a common choice for small businesses.” They often don't coincide. Your page can be cited as the source while a competitor is the brand being recommended, which tells you something precise: your content was good enough to inform the answer, but it isn't (yet) the thing being recommended.
A brand that's cited often but rarely recommended has a different problem to solve than one that doesn't appear at all, which is why the two are worth tracking separately. This article is about the citation side. For the recommendation side, see how to track brand mentions in AI search.
Why tracking citations is harder than tracking Google rankings
SEO rank tracking works because Google is broadly stable: one keyword, roughly the same ten results (even accounting for some of Google's personalization), and a position you can check over time. AI search offers none of those guarantees, and citations inherit every reason why.
There's no fixed ranking to track. LLMs generate answers by sampling from a probability distribution, so they produce variation by design. Research has shown a 1-in-100 chance of getting the same list of recommended brands twice from the same prompt. Citations move for the same reason: run a prompt, note the sources, run it again, and the cited set will often be different.
Chat responses are shaped by context you can't see. When someone asks an LLM about products in your category, the model factors in their chat history, stated preferences, and account context, then searches and cites accordingly. Your tool, running the same prompt from a clean session, isn't seeing exactly what that specific user sees. We call this the problem of invisible prompts.
Google AI Overviews behaves differently, but is no more stable. AI Overviews responds to fairly standard keyword searches rather than long personalized conversations, so the invisible-prompt effect is weaker here. The instability comes from query fan-out instead. Rather than answering your search directly, Google expands it into a set of related sub-queries, runs each one in the background, and assembles the overview from the sources that come back across all of them.
For instance, a search for "best project management software for small teams" might fan out into separate queries for affordable options, tools aimed at small business, and recent reviews, each surfacing its own pages. The overview you see is stitched together from that wider pool, not from the single query you typed. Because the fan of sub-queries isn't fixed, the same search can pull a different set of cited sources from one run to the next, even when nothing about your content or your original query has changed.
Citations vary across platforms. A page ChatGPT cites constantly may barely register in Perplexity, which leans on live web results more than ChatGPT does. Blending citation data across engines into one average hides these differences, which are often the most useful thing in the data. This is why we track and report per LLM rather than combining everything, and treat each engine as its own surface. (We go deeper on individual engines in our guides to the Perplexity rank tracker and rank tracking in ChatGPT.)
API data isn't what users see. Many tracking tools query LLMs through their APIs because it's cheaper and easier. The catch is that API responses run differently than web interfaces, without the system prompts and product-level tuning that shape a real answer. A source cited through the API may not get cited in the interface a customer uses. This was the single biggest reason we rebuilt our first version of Traqer, switching to responses captured from real logged-out browser sessions instead.
Citation rate at the topic level
If a single prompt result is unreliable, the answer isn't to track single prompts more carefully. Rather, the best approach is to stop reading them individually and start measuring patterns across many related prompts.
While any single result is close to random, how often a brand appears across dozens of runs of similar prompts is meaningfully stable. A page cited in 70% of the prompts within a topic area is in a stronger position than one cited in 15%, and that difference remains reliable once you measure across enough prompts.
In Traqer, the topic is the top level of that structure. A topic (for example “best project management software for small teams”) sits above a set of prompts that approach the same buying question from different angles, so instead of tracking one phrasing you track the group.

Traqer generates those prompts, runs each across the LLMs you care about, and reports citations and brand mentions separately. Your citation rate for a topic is the share of its prompts where your domain appears as a source, reported per LLM rather than blended into one number. Measured that way, it's far more reliable than any single check, because it samples across the variability rather than pretending it isn't there.
Reporting this way also protects you from a metric that quietly games itself. A raw visibility percentage improves the moment you stop tracking the prompts where you don't appear, without anything real changing. A citation count at the topic level doesn't move that way. It only goes up when your domain starts getting cited on more prompts, which is the thing you care about.
How to set up citation tracking
Here's the process Traqer is designed to support:
Start with bottom-of-funnel topics. LLMs rarely cite or recommend specific brands for broad informational questions; they answer those directly. Sources and brands show up on buying-intent queries: e.g., "best X for Y" or "alternatives to Z". If you're already running a Pain Point SEO strategy, this list of queries will look pretty familiar.
Use multiple prompts per topic, not one. One phrasing only tells you what came back for that exact wording on that run. Several realistic prompts per topic give you a citation rate you can trust. Traqer generates these automatically during setup, including a plain keyword version for AI Overviews users who still search in short phrases, plus natural-language questions for the chat engines.
Track the engines your buyers use. ChatGPT, Perplexity, Google AI Overviews, Gemini, and Claude cite differently, so run your prompts across the ones that matter for your audience and keep the results separate per engine.
Capture the real interface, not the API. Whatever you use, confirm it reads what users see rather than raw API output. Traqer scrapes the live web interfaces and attaches a screenshot of the actual response to every prompt, so you can see the citation in context.
Separate citations from mentions from the start. Decide whether you're looking at who gets cited as a source or who gets recommended by name, and read each on its own. Blending them will flatter or mislead you depending on which way the gap runs.
Measure change over time. Citation data is only useful as a trend. Traqer updates weekly and lets you compare any two dates, so you can see which topics gained cited pages and which lost them between reporting periods.
One caveat is that because a fully personalized session is invisible to any external tool, Traqer measures a neutral, logged-out user rather than one specific person's contextualized experience. That's an accurate baseline for tracking change over time, but it's not a claim to reproduce exactly what every individual sees.
Using citation data for your content strategy
Tracking only helps if you can act on what it shows you. While your citation rate is worth knowing, the data that shapes what you produce is which sources the models cite for a topic, and where your site is missing from that set. Traqer's Analyze and Improve view shows, for each topic, which domains the LLMs cite most and which specific pages get cited more than once.

That gives you two clear approaches, which line up with the hierarchy we lay out in Prioritized GEO:
Produce owned content that ranks. Content on your own site that ranks in traditional search for a buying-intent keyword tends to get pulled in and cited when a model answers a related product query. This is the foundation, and it's the same bottom-of-funnel content that works for AI search generally. (Our Constitution Lending case study walks through how specific, product-level content earned citations against much larger competitors.)
Get mentioned on the pages LLMs already cite. The "cited more than once" list is effectively an outreach target list. If a roundup or comparison article keeps getting cited for your topic and you're not on it, that's a gap you can pursue through a guest contribution, an expert quote, or a listing. Our Toro TMS case study shows this playing out in a B2B software category.
There are limits worth noting, however. Neither lever guarantees a result. Ranking in Google doesn't necessarily mean a model will cite you, and getting onto a frequently cited page doesn't mean it will then recommend you by name. These are well-supported hypotheses we've seen work repeatedly across clients, not levers you pull for a fixed output. You're influencing the inputs, not controlling the answer.
One tactic we can be clearer about: the on-site hacks that get hyped every week, such as adding an llms.txt file, rewriting headings as questions, or bolting on FAQ schema, have shown no measurable effect in our testing. They address whether a model understands your page once it arrives, not whether it ever encounters and cites it. The topic-based GEO approach above is where the results come from.
A final note on AI visibility tools
Many AI visibility tools make citation data difficult to trust: they blend citations with brand mentions, they track individual prompts as if the results were stable, and they pull data through APIs rather than the real interfaces. Enterprise platforms are capable but priced for large in-house teams, with paid tiers that start around a few hundred dollars a month and rise into custom enterprise contracts.
We built Traqer to track citations the way we think they should be measured: at the topic level, separated from brand mentions, broken out per LLM, and captured from the real web interfaces rather than APIs.
Pricing is per prompt with unlimited brands and users on every plan, which is why it works for agencies tracking many clients at once. If that framework fits how you want to measure AI search, you can compare it against others in our overviews of the best AI visibility tool and AI visibility tracking tools for agencies.
Traqer is built by Grow and Convert. For the strategy behind how we measure AI visibility, Topic-Based GEO and Invisible Prompts explain the thinking.
