Prompt Tracking: measuring AI visibility without counting noise
The same question returns different answers at the same moment. What a measurement setup that survives this looks like, and why numbers from two tools cannot be compared.
Prompt tracking is the repeated asking of a fixed set of questions to AI systems in order to measure whether and how a brand appears in the answers. The systems observed include ChatGPT, Perplexity, Google AI Overviews, Google's AI Mode and Gemini. Each run records whether the brand is named, in what context, and which sources the system draws on.
The difference from classic rank tracking lies in the stability of what is being measured. A position in a list of results stays the same for hours. An AI answer does not, not even for an identical question asked at the same moment.
What prompt tracking measures
The measurement has three parts. A fixed list of questions, a fixed procedure for asking them, and a fixed record of what comes back.
Each run records:
whether your brand is named
in what context it is named
which competitors appear alongside it
which sources the system links or names
whether the statement about the brand is accurate
The metrics are built from this raw data. Getting from a single answer to a reliable metric is what this article is about.
The second meaning of the term
Prompt tracking also refers to something else entirely. In software development the term means logging your own prompts inside your own application, covering versioning, cost control through token usage, A/B tests and debugging.
The two meanings share a name and have nothing else in common. Anyone searching for tools runs into both kinds. This article covers the measurement of brand visibility.
Why the same question returns different answers
This is where most measurement setups fall apart. Asking a question once and noting the result draws a sample of one.
How large the spread is
A study from the University of St. Gallen quantified this in spring 2026. Four systems, 32 questions across four industries, up to ten simultaneous runs per pairing.
The result is stark. Across simultaneous repetitions of the same question, the cited sources overlap by only 32 to 43 percent on average. This is not about different days. This is the same question, at the same moment, to the same system.
From one day to the next, around 65 percent of cited sources change.
The study is a preprint with no visible peer review, it covers 32 questions, and the lead author is affiliated with a vendor alongside the university. The order of magnitude does match what other measurements show.
What happens on top of that over weeks
SISTRIX measured the movement over time separately, across 82,619 questions and roughly 1.5 million snapshots between December 2025 and April 2026.
System
Share of cited domains that change weekly
Google AI Overviews
5 percent
Google AI Mode
56 percent
ChatGPT with web search
74 percent
The same study shows two findings that belong together. 86.5 percent of questions have a stable core of one to five domains. Of the remaining domains, 89 percent rotate weekly.
So there is a fixed base with constant churn around it. Looking only at the total number of sources shows movement where a stable core sits.
One more figure stands out. For an identical query, Google AI Overviews and AI Mode name different domains in 83 percent of cases. Two surfaces from the same provider, two answers.
Where the spread comes from technically
SEO guides usually say the spread comes from the model's temperature setting. That falls short.
Thinking Machines Lab measured in September 2025 what happens with random selection switched off. At temperature 0, meaning theoretically deterministic selection, 1,000 runs of the same prompt produced 80 different answers. The most frequent one appeared 78 times.
The cause lies in fluctuating server load. How many requests a compute node handles at once changes the result of the calculation in the final decimal places. Across hundreds of tokens that adds up to different wording.
OpenAI says as much for its own interface. On the parameter for a fixed random seed, the documentation states that determinism is not guaranteed.
Systems with web search add a second cause. The system searches at runtime and gets different results depending on the moment.
What the user context shifts on top
Beyond the model's own spread, the context of the measurement has an effect.
Google describes personalisation in AI Mode in its own help pages. Earlier searches and saved activity feed into the answer if history is switched on. Personalisation can be turned off.
ChatGPT can store memories from earlier conversations and draw on them in new answers.
One hard requirement follows for measurement. Measuring while signed in measures visibility for that one account with that one history. Comparable numbers require measuring signed out, without history, with a fixed region and language.
What this means for a single measurement
The authors of the St. Gallen study derive a concrete number from their data. They recommend at least seven runs per question per day for tracking brand mentions, and at least eight when the sources matter. From seven runs onward, the standard error of the detection rate drops below 0.10.
That is the practical consequence of the spread. Asking a question once and writing the result into a reporting sheet produces a number that carries no meaning.
Building the prompt list
The list is the measuring instrument. If it changes, the comparison with last month is worthless.
How many questions make sense
The number follows from arithmetic rather than instinct. Every question is asked several times, across several systems, at a fixed interval.
At seven runs, four systems and daily measurement, a single question produces 196 requests per week. Tracking 50 questions means almost 10,000 requests per week.
Less is usually more here. A list of 20 to 40 questions covering the key buying decisions is more reliable than 200 questions asked once each.
Branded and unbranded questions
Both kinds belong on the list, and they answer different things.
Unbranded questions do not contain your name. They measure whether the brand is considered at all. Example: „Which providers offer guest posts in German-speaking markets?"
Branded questions contain the name. They measure what a system says about the brand when asked directly. Example: „What is GetMentioned?"
Unbranded questions are the harder test. A brand that only appears when asked by name is invisible during the selection phase.
Where the questions come from
the questions sales and support actually receive
the follow-up questions from first calls
queries from Search Console, filtered for question words
comparison questions involving competitors
questions about pricing, selection and fit
Which questions are not worth tracking
questions with no bearing on a buying decision, because a mention there leads nowhere
questions aimed at a factual topic where no provider would be named
questions whose answer changes daily, such as those about current events
questions phrased so narrowly that they force a particular answer
Setting up the measurement
How often to measure
Weekly is enough for most purposes. The churn figures above explain why. With AI Overviews only 5 percent of sources change per week, so daily measurement yields little. With ChatGPT and web search it is 74 percent, where daily measurement would be mostly noise.
The measurement setup in five steps.
More runs per measurement point beat more measurement points.
What has to stay the same between measurements
the wording of the question, character for character
the system and, where visible, the model version
language and region
the sign-in state
the number of runs
If one of these changes, the series is broken. Record it, or a measurement artefact will later be read as a trend.
The most common break comes from outside. When a provider swaps the model behind its service, the answers change without anything in your setup being different. Such changes are rarely announced. A note in the log about when a sharp jump occurred helps more later than any explanation constructed after the fact.
Which systems are worth tracking
The available numbers diverge widely here, and the gap is itself informative.
Providers report large reach. ChatGPT states 900 million weekly active users, the Gemini app 950 million monthly active users, and Google's AI features in search over one billion monthly.
Panel measurements paint a different picture. According to analyses from early 2026, AI tools account for roughly 3 percent of search activity and send less than 1 percent of referral traffic.
The contradiction dissolves in the counting method. Google's billion counts anyone who saw an AI answer, not anyone who deliberately chose a tool. Weekly and monthly active users are not comparable to begin with.
For Germany, a Bitkom survey from November 2025 offers the most usable reference point. Half of respondents at least occasionally use AI chats instead of a search engine. Five percent rely on AI exclusively, another seven percent predominantly.
Two things follow for the selection. Google's AI features belong in every measurement because of their reach. ChatGPT belongs there because that is where deliberate use happens. Everything beyond depends on the market.
What each run records
For every combination of question, system and run, these fields belong in the log:
Field
Content
Timestamp
date and time
System
ChatGPT, Perplexity, AI Mode and so on
Brand named
yes or no
Position
where in the answer
Context
recommendation, example, list entry, warning
Competitors
which appear alongside
Sources
linked or named domains
Accuracy
whether the statement about the brand holds
The accuracy field is often forgotten. A false statement about your brand in an answer matters more than a missing mention.
The metrics and what they mean
Mention rate
The share of runs in which the brand is named. This is the base metric, and it always needs a reference. 40 percent across 7 runs over 20 questions says something different from 40 percent across one run.
Share of the mention
How often your brand is named relative to all named providers. Depending on the vendor this is called share of voice or share of model.
It says more than the plain mention rate because it accounts for competition. It is also more sensitive. If a system suddenly names ten providers instead of three, the share drops without your visibility having changed.
Citation rate
The share of answers in which your own domain is linked or named as a source. Mention and citation come apart. A brand can be recommended without its website appearing as a source.
Context and tone
Whether the mention comes as a recommendation, as one example among many, or with a caveat. This assessment is the hardest to automate and often the most useful.
Why two tools produce different numbers
There is no shared definition. Every vendor calculates differently, and the differences are larger than the effects being measured.
One example of counting rules. One vendor counts a mention once per answer, no matter how often the brand appears in the text. The same vendor counts citations individually. Two tools with different counting rules report different values for the same answer.
The measurement channel adds to this. One vendor documents explicitly that it measures through the official interfaces of the AI providers. That is clean and repeatable, but it does not show what a signed-in user sees in the interface. Other vendors say nothing on the matter.
The practical consequence. Numbers from different tools do not belong in one table. Within a single tool, the trend over time is meaningful. The comparison between tools is not.
Measuring yourself or buying a tool
Measuring yourself means keeping the questions in a spreadsheet, retrieving the answers manually and entering the results. That works up to around ten questions and two systems. At seven runs per question, that is already 140 retrievals per round.
Tools take that work off your hands. Four questions are worth asking when choosing one:
Is measurement done through the interface or through the user-facing product?
How many runs per question, and is that number stated?
How is the metric defined, with a formula?
How many requests does the plan actually include?
The last question deserves arithmetic. One vendor states consumption per answer and includes 2,500 units per year in a mid-tier plan. On an expensive system that comes to roughly 1,250 answers per year, which at daily measurement is about three questions.
The open disagreement between vendors
Two vendors hold publicly opposing positions.
SISTRIX deliberately skips repeat runs, arguing that repeated asking changes the wording but not the facts, because the underlying model stays the same.
seoClarity makes repetition a selling point and speaks of a statistically significant presence rate across thousands of variations. The vendor names no number of repetitions and no significance level.
The St. Gallen measurement argues against the first position, at least at source level. If simultaneous runs overlap by only a third to two fifths of their sources, more than the wording is changing.
One first screening question follows for your own selection. How many runs per prompt does the tool perform, and does it state the number?
What prompt tracking does not tell you
It names no causes
The measurement shows whether a brand is named. It does not show why. A rising mention rate can come from your own work, from a model change at the provider, or from a competitor dropping out.
It proves no business effect
No study with a documented method has measured a causal link between being named in AI answers and revenue. Correlations exist, and they come mostly from vendors selling measurement tools.
That is not an argument against measuring. It is an argument against the promises derived from it.
It is not complete
What gets measured is a selection of questions in a selection of systems at one point in time. Nobody knows the full picture of what users actually ask, because the providers do not release that data.
Measuring improves nothing
Prompt tracking is observation. Whether a brand appears in answers is decided elsewhere, in the clarity of your own information and in what independent sources report about you. The article on AI visibility covers that work.
Conclusion: understand the spread before you measure
Prompt tracking measures a moving target. The same question, asked at the same moment to the same system, produces answers whose sources overlap by only a third to two fifths. Measuring once measures chance.
A workable setup follows almost by itself. A small, fixed list of questions. Several runs per question. Constant conditions. Metrics with a stated reference. And the willingness to read small movements as what they usually are, which is noise.
Frequently asked questions
What is prompt tracking?
Repeatedly asking a fixed set of questions to AI systems to measure whether and how a brand appears in the answers. Each run records the mention, its context, the competitors named and the sources used.
How often should you measure?
Weekly is usually enough. More runs per measurement point matter more than more frequent measurement, because the same question returns different answers.
How many runs does one question need?
A University of St. Gallen study from April 2026 recommends at least seven runs per question per day, and at least eight when the cited sources are being analysed.
Why do I get different answers to the same question?
The output of a language model is not repeatable even with random selection switched off. The main reason is fluctuating server load, which changes the calculation in the final decimal places. Systems with web search add changing search results on top.
How many prompts should I track?
Usually 20 to 40. Each question is asked several times across several systems, so the effort grows quickly. A small, well-chosen list is more reliable than a long one asked once each.
What is the difference between branded and unbranded prompts?
Branded questions contain the brand name and measure what a system says about the brand. Unbranded questions leave it out and measure whether the brand is considered at all.
Does being signed in affect the result?
Yes. Google draws on earlier searches in AI Mode when history is enabled. ChatGPT can use memories from earlier conversations. Comparable measurements should be taken signed out and without history.
Are numbers from two tracking tools comparable?
No. There is no shared definition of the metrics, and vendors measure partly through interfaces and partly through the user-facing product. What is meaningful is the trend within one tool.
What does prompt tracking cost?
Prices range from around 119 euros per month at a vendor with a public price list to plans quoted only on request. The number of requests included matters more than the headline price.
Which systems should I track?
Google's AI features for reach and ChatGPT for deliberate use. In Germany, a Bitkom survey from November 2025 found around 5 percent using AI chats exclusively instead of a search engine, with another 7 percent using them predominantly.
Does prompt tracking show whether my work is paying off?
Only to a limited extent. It shows changes but names no causes. A rise can come from your own work, from a model change, or from a competitor dropping out.
Does prompt tracking improve my visibility?
No, it observes it. Visibility comes from clear information of your own and from mentions on sites the systems follow.
David has been building link acquisition and digital PR processes since 2016, first as an agency under SEO Galaxy, today as a platform with GetMentioned. He has scaled his own projects from zero to seven-figure monthly traffic and delivered thousands of campaigns for clients. Here he writes about what works in practice, and about what only costs budget.
Your next link does not have to be a blind buy
Compare publishers, SEO data and prices in one place and book the placements that fit.
AI answers recommend providers before anyone clicks a ranking. How to measure AI visibility, which signals carry it, and where the external evidence models lean on comes from.
GEO targets being named in an AI answer rather than a position in the list of results. What is evidenced, which levers work and what the circulating figures are worth.
AEO works on the individual section, not the whole page. What Google says about the term, which levers hold up and how to build a section an answer engine can use.