All articles

How to track ChatGPT rankings over time

Devesh KhanalDevesh KhanalJuly 9, 202612 minutes read
How to track ChatGPT rankings over time
Share

As marketing teams focus more on their visibility in AI tools, the natural response is to track ranking positions in the way they once tracked Google; set up the queries, watch the line on the chart, and report whether you went up or down week over week.

The problem is that there is no fixed position in ChatGPT to plot. Ask ChatGPT the same product question ten times and you will get ten different answers, with different brands named, in a different order, sometimes with no recommendations at all. There is no equivalent of a stable rank that holds still long enough to track. A chart that shows your ChatGPT “ranking” moving from third to second to fifth is not recording a real change in your position, it is recording the model's normal variation between runs.

But that doesn’t mean tracking is hopeless. It just means you have to be more strategic in what you track and how you track it. Over the past year we have been measuring AI visibility for dozens clients at Grow and Convert, and we built our own tool, Traqer, to track brand visibility in AI tools like ChatGPT properly, so we measure real changes in visibility not random AI noise. This article shares our experience from the data we’ve collected for dozens of brands on what actually moves over time in ChatGPT, what you can measure reliably, and how to set up tracking so the trend you are looking at means something.

Why there’s no "ranking" in ChatGPT to track

Traditional rank tracking works because Google is relatively stable. The same query produces roughly the same set of results for roughly the same user. There is some personalization, but a keyword has a position you can record today and check again next week, and the comparison is meaningful.

ChatGPT does not behave that way, and the reason matters for how you measure it.

Real prompts are effectively unique

When a real person asks ChatGPT for a recommendation, the words they type are only a fraction of what the model is working with. ChatGPT factors in everything it already knows about that user from saved memory, past conversations, and the context of the current chat. A six-word question on screen can carry pages of background the model is quietly drawing on.

We have started thinking about this as the difference between the literal prompt and the effective prompt. The literal prompt is what the user types. The effective prompt is that text plus all the personal context ChatGPT adds before it answers. For example, one of our clients found us through ChatGPT after asking for the best SEO companies in her space. Her literal question was short and casual, but ChatGPT already knew her industry, her company size, and her goals from months of earlier conversations, and it recommended us based on all that hidden context. When we asked the same question from our own account, the answer was completely different.

We call these invisible prompts. The practical consequence for tracking is that even if you track the exact wording a real user typed, you will not reproduce their result, because you cannot reproduce the context underneath it. The prompts that actually generate your leads are ones you will never see, and their effective search volume is one. 

That is not a problem you can solve by adding more prompt variations to a tracking tool. It is a feature of how these tools work, and it is getting more pronounced as people stay logged in and the models remember more about them.

The same prompt gives different answers each time

There is a second reason a single position is not trackable, separate from personalization. LLMs generate responses probabilistically, so even the identical prompt run twice will vary.

This has now been measured. A study of around 2,961 prompts across ChatGPT, Claude, and Google's AI, testing 12 brand-recommendation questions dozens of times each. It found less than a 1-in-100 chance of getting the same list of brands twice, and less than a 1-in-1,000 chance of getting that list in the same order. The number of brands named varied too, sometimes two or three, sometimes ten or more.

The study found one more thing that matters for measurement. While the exact list and order were close to random, the underlying set of brands was not. In a given category the same familiar names kept appearing, just in different combinations and positions. This is why appearance frequency is worth measuring even when rank is not. You cannot reliably track where you come in the list, but you can track how often you make it into the list at all.

What you can track over time is visibility at the topic level

Because any single prompt result is unreliable, what you measure instead is how often your brand appears across a range of prompts about the same buying-intent topic. We call this topic-based visibility, and it’s the unit that holds up over time.

A topic is essentially a buying-intent area your customers ask about, something like “wedding photographer in London,” “commercial roofing contractor near me,” or “best meal delivery service for families.” Within that topic you track multiple prompts that approach it from different angles. Your visibility is the percentage of those prompts where your brand is mentioned, measured separately for each LLM. “We appeared in 7 of 10 prompts on Perplexity and 3 of 10 on ChatGPT for this topic” is a number you can record this week, check again next week, and act on.

This matters because AI visibility isn’t a single brand-wide figure. It is specific to each topic. A brand can have strong visibility for one topic and almost none for a closely related one. 

For example, Grow and Convert’s visibility for “B2B content marketing services” looks different from visibility for “B2C content marketing services,” because more has been published about the B2B work and the LLMs know the brand better in that context. Averaging the two into one brand score would hide the only information that tells us what to do next.

Tracking at the topic level also handles both problems described above. Output variability evens out across a set of prompts rather than swinging wildly on a single one. And while you cannot see the invisible prompts your real buyers use, you can still see whether your presence is growing for the topics those conversations are about, even though the individual prompts remain out of view.

How to set up tracking you can trust over time

  1. The prompts you track should reflect how your customers actually ask, not be left entirely to the tool. Suggested prompts are just a starting point. The thing to avoid is treating that initial list as final. Review what gets suggested, cut anything no real person would type, and add the phrasings your customers actually use, so the tool supports your strategy rather than setting it. You decide the topics and prompts based on your bottom-of-funnel keyword strategy.

  2. Use a metric that doesn’t move when you add prompts. Many teams report a single visibility percentage across all their prompts, and it creates a predictable problem. The number drops the moment you start tracking new topics you are not yet visible for, even though nothing about your actual visibility has changed.

    We watched this happen in another tool. On one date a brand showed 26.7% visibility. The next day we added two new prompts that were relevant but had no visibility yet, and the score fell to 20%. In reality, the brand had lost nothing. The metric only looked worse because we were tracking more ground. A percentage that punishes you for being ambitious creates a bad incentive, because it pushes teams to avoid tracking the topics they most want to win.

    A raw count of prompts where you appear behaves better over time. It only goes up when you genuinely gain visibility, and adding aspirational topics does not drag it down. Track the count alongside the percentage so you can distinguish real movement. 

  3. Separate brand mentions from citations. Two different things get called “visibility,” and blending them hides which one is improving. A brand mention is ChatGPT naming your brand in its answer, the recommendation itself. A citation is your URL appearing as a source link the model drew on. Being recommended by name is what generates leads. Being cited is useful but is not the same thing and nowhere near as valuable. If your tracking combines them, a rising line could mean you’re getting recommended more, or it could mean you’re getting cited more while recommendations stay flat. Those call for different responses, so keep the two metrics separate and watch each trend on its own.

  4. Consider LLMs as separate. Visibility on ChatGPT and visibility on Perplexity are not interchangeable. Perplexity and Google On the other hand, ChatGPT leans more on training data and on how widely your brand is mentioned across the web, so the same content tends to help there too, but more slowly and less directly. The practical result is that a brand can be well established on the search-based tools while still building presence on ChatGPT, and a blended cross-LLM score hides exactly that gap. Track each platform on its own line.

Reading the trend and connecting it to what you did

With a stable topic-level trend for each LLM in place, you can start to read what the movement means and respond to it.

The first thing to know is that visibility responds slowly. When you publish content or earn a mention on a site ChatGPT draws from, the effect tends to show up weeks later rather than days. For that reason it is better to read the trend over a span of months than to treat each week as a verdict on whatever you did last.

When you want to improve a topic where visibility is low, two levers do most of the work in our experience, but neither is guaranteed.

The first is owned content that ranks in traditional search. Content on your own site that ranks for bottom-of-funnel keywords related to a topic tends to get pulled into AI answers, particularly on the search-based tools. This is the foundation of the GEO Priorities Pyramid, where owned content comes first. 

If your domain is not among the sources LLMs cite for a topic, producing content that ranks for those queries is the place to start.

The second is getting mentioned on the sites LLMs already cite (Tier 2 above). Look at which specific pages and publications keep appearing as sources for a topic, then work to get your brand mentioned in that content through outreach or by being added to existing listicles and comparisons. In traditional SEO a link signals authority. In AI search, a third-party mention signals that sources other than you recognize your brand in the category, and the models appear to weight that external validation heavily.

The caveat is that neither lever guarantees a result. In our analysis, the majority (60%) of sources ChatGPT cites seem to be ranking for topically relevant SEO keywords. But that leaves 40% of sources it cites as having little to no SEO presence. So ranking in Google does not guarantee that  ChatGPT will mention you, and getting onto a frequently cited page does not guarantee the model will then recommend you. 

But being visible when ChatGPT searches the web to answer a user’s question is a very solid, evidence-backed way of increasing your brand exposure to ChatGPT, which then improves your brand’s chances of being recommended to the right user in the right context.  The reason to track at the topic level is precisely so you can see, over time, whether the actions you are taking are actually moving your visibility for the topics that matter.

How Traqer tracks visibility over time

We built Traqer to measure visibility this way, because the tools we tried were either priced way too high to track multiple brands with dozens of prompts each, or reported it in ways that simply didn’t make sense. 

In Traqer, tracking is organized by topic, with multiple prompts per topic, and visibility is shown per LLM across ChatGPT, Perplexity, Google AI Overviews, Gemini, and Claude rather than blended into one figure. 

You set the topics and prompts based on your own bottom-of-funnel strategy. A toggle at the top of every view lets you see brand mentions only, citations only, or both, so you are never guessing which one a number represents. Alongside the per-LLM visibility percentage, Traqer reports a raw count of prompts where you appear and a topic-level breakdown of where visibility is high, some, or none, so adding ambitious new topics never makes your progress look worse than it is.

For tracking over time specifically, Traqer refreshes on a weekly cycle and keeps the history. You can pick any two dates and compare them side by side, at the brand level to see how each LLM moved and at the topic level to see which topics gained or lost ground. Every tracked prompt links to a screenshot of the real response as it appeared that week, so the trend is backed by evidence you can show a client or an executive.

If you want to go deeper on the measurement approach, our guide to AI visibility tools covers what to look for in any tool, and our piece on AI visibility tracking tools for agencies covers the multi-brand economics that matter if you are tracking visibility for clients.

The right way to think about tracking ChatGPT over time

You cannot track a ranking in ChatGPT, because there is no stable ranking to track. What you can track is how often your brand appears across the prompts that matter for a topic, measured the same way each week, per LLM, with mentions and citations kept apart. Read over months and set against the content and outreach you actually did, that trend will show whether your AI visibility is improving, which is more than a single “ranking” figure that moves up and down with the model's normal variation can tell you.

See how often your brand appears across ChatGPT, Perplexity, Gemini, Google AI Overviews, and Claude. Start tracking your AI visibility with Traqer.