SEO Galaxy is now GetMentioned. The platform stays the same. Book your slot here →

GetMentioned
Strategy

Prompt Tracking: measuring AI visibility without counting noise

The same question returns different answers at the same moment. What a measurement setup that survives this looks like, and why numbers from two tools cannot be compared.

David Hahn

David Hahn · August 15, 2026 · 15 min read

One question fans out into seven runs, four of which name the brand

The short version

  • Prompt tracking measures a moving target, not a fixed ranking.
  • A single run per question mostly shows chance.
  • One study recommends at least seven runs per question per day.
  • Vendors openly disagree on whether repeat runs are needed at all.
  • Numbers from two different tools cannot be compared.

Prompt tracking is the repeated asking of a fixed set of questions to AI systems in order to measure whether and how a brand appears in the answers. The systems observed include ChatGPT, Perplexity, Google AI Overviews, Google's AI Mode and Gemini. Each run records whether the brand is named, in what context, and which sources the system draws on.

The difference from classic rank tracking lies in the stability of what is being measured. A position in a list of results stays the same for hours. An AI answer does not, not even for an identical question asked at the same moment.

What prompt tracking measures

The measurement has three parts. A fixed list of questions, a fixed procedure for asking them, and a fixed record of what comes back.

Each run records:

  • whether your brand is named
  • in what context it is named
  • which competitors appear alongside it
  • which sources the system links or names
  • whether the statement about the brand is accurate

The metrics are built from this raw data. Getting from a single answer to a reliable metric is what this article is about.

The second meaning of the term

Prompt tracking also refers to something else entirely. In software development the term means logging your own prompts inside your own application, covering versioning, cost control through token usage, A/B tests and debugging.

The two meanings share a name and have nothing else in common. Anyone searching for tools runs into both kinds. This article covers the measurement of brand visibility.

Why the same question returns different answers

This is where most measurement setups fall apart. Asking a question once and noting the result draws a sample of one.

How large the spread is

A study from the University of St. Gallen quantified this in spring 2026. Four systems, 32 questions across four industries, up to ten simultaneous runs per pairing.

The result is stark. Across simultaneous repetitions of the same question, the cited sources overlap by only 32 to 43 percent on average. This is not about different days. This is the same question, at the same moment, to the same system.

From one day to the next, around 65 percent of cited sources change.

The study is a preprint with no visible peer review, it covers 32 questions, and the lead author is affiliated with a vendor alongside the university. The order of magnitude does match what other measurements show.

What happens on top of that over weeks

SISTRIX measured the movement over time separately, across 82,619 questions and roughly 1.5 million snapshots between December 2025 and April 2026.

SystemShare of cited domains that change weekly
Google AI Overviews5 percent
Google AI Mode56 percent
ChatGPT with web search74 percent

The same study shows two findings that belong together. 86.5 percent of questions have a stable core of one to five domains. Of the remaining domains, 89 percent rotate weekly.

So there is a fixed base with constant churn around it. Looking only at the total number of sources shows movement where a stable core sits.

One more figure stands out. For an identical query, Google AI Overviews and AI Mode name different domains in 83 percent of cases. Two surfaces from the same provider, two answers.

Where the spread comes from technically

SEO guides usually say the spread comes from the model's temperature setting. That falls short.

Thinking Machines Lab measured in September 2025 what happens with random selection switched off. At temperature 0, meaning theoretically deterministic selection, 1,000 runs of the same prompt produced 80 different answers. The most frequent one appeared 78 times.

The cause lies in fluctuating server load. How many requests a compute node handles at once changes the result of the calculation in the final decimal places. Across hundreds of tokens that adds up to different wording.

OpenAI says as much for its own interface. On the parameter for a fixed random seed, the documentation states that determinism is not guaranteed.

Systems with web search add a second cause. The system searches at runtime and gets different results depending on the moment.

What the user context shifts on top

Beyond the model's own spread, the context of the measurement has an effect.

Google describes personalisation in AI Mode in its own help pages. Earlier searches and saved activity feed into the answer if history is switched on. Personalisation can be turned off.

ChatGPT can store memories from earlier conversations and draw on them in new answers.

One hard requirement follows for measurement. Measuring while signed in measures visibility for that one account with that one history. Comparable numbers require measuring signed out, without history, with a fixed region and language.

What this means for a single measurement

The authors of the St. Gallen study derive a concrete number from their data. They recommend at least seven runs per question per day for tracking brand mentions, and at least eight when the sources matter. From seven runs onward, the standard error of the detection rate drops below 0.10.

That is the practical consequence of the spread. Asking a question once and writing the result into a reporting sheet produces a number that carries no meaning.

Building the prompt list

The list is the measuring instrument. If it changes, the comparison with last month is worthless.

How many questions make sense

The number follows from arithmetic rather than instinct. Every question is asked several times, across several systems, at a fixed interval.

At seven runs, four systems and daily measurement, a single question produces 196 requests per week. Tracking 50 questions means almost 10,000 requests per week.

Less is usually more here. A list of 20 to 40 questions covering the key buying decisions is more reliable than 200 questions asked once each.

Branded and unbranded questions

Both kinds belong on the list, and they answer different things.

  • Unbranded questions do not contain your name. They measure whether the brand is considered at all. Example: „Which providers offer guest posts in German-speaking markets?"
  • Branded questions contain the name. They measure what a system says about the brand when asked directly. Example: „What is GetMentioned?"

Unbranded questions are the harder test. A brand that only appears when asked by name is invisible during the selection phase.

Where the questions come from

  • the questions sales and support actually receive
  • the follow-up questions from first calls
  • queries from Search Console, filtered for question words
  • comparison questions involving competitors
  • questions about pricing, selection and fit

Which questions are not worth tracking

  • questions with no bearing on a buying decision, because a mention there leads nowhere
  • questions aimed at a factual topic where no provider would be named
  • questions whose answer changes daily, such as those about current events
  • questions phrased so narrowly that they force a particular answer

Setting up the measurement

How often to measure

Weekly is enough for most purposes. The churn figures above explain why. With AI Overviews only 5 percent of sources change per week, so daily measurement yields little. With ChatGPT and web search it is 74 percent, where daily measurement would be mostly noise.

The measurement setup in five steps.

More runs per measurement point beat more measurement points.

What has to stay the same between measurements

  • the wording of the question, character for character
  • the system and, where visible, the model version
  • language and region
  • the sign-in state
  • the number of runs

If one of these changes, the series is broken. Record it, or a measurement artefact will later be read as a trend.

The most common break comes from outside. When a provider swaps the model behind its service, the answers change without anything in your setup being different. Such changes are rarely announced. A note in the log about when a sharp jump occurred helps more later than any explanation constructed after the fact.

Which systems are worth tracking

The available numbers diverge widely here, and the gap is itself informative.

Providers report large reach. ChatGPT states 900 million weekly active users, the Gemini app 950 million monthly active users, and Google's AI features in search over one billion monthly.

Panel measurements paint a different picture. According to analyses from early 2026, AI tools account for roughly 3 percent of search activity and send less than 1 percent of referral traffic.

The contradiction dissolves in the counting method. Google's billion counts anyone who saw an AI answer, not anyone who deliberately chose a tool. Weekly and monthly active users are not comparable to begin with.

For Germany, a Bitkom survey from November 2025 offers the most usable reference point. Half of respondents at least occasionally use AI chats instead of a search engine. Five percent rely on AI exclusively, another seven percent predominantly.

Two things follow for the selection. Google's AI features belong in every measurement because of their reach. ChatGPT belongs there because that is where deliberate use happens. Everything beyond depends on the market.

What each run records

For every combination of question, system and run, these fields belong in the log:

FieldContent
Timestampdate and time
SystemChatGPT, Perplexity, AI Mode and so on
Brand namedyes or no
Positionwhere in the answer
Contextrecommendation, example, list entry, warning
Competitorswhich appear alongside
Sourceslinked or named domains
Accuracywhether the statement about the brand holds

The accuracy field is often forgotten. A false statement about your brand in an answer matters more than a missing mention.

The metrics and what they mean

Mention rate

The share of runs in which the brand is named. This is the base metric, and it always needs a reference. 40 percent across 7 runs over 20 questions says something different from 40 percent across one run.

Share of the mention

How often your brand is named relative to all named providers. Depending on the vendor this is called share of voice or share of model.

It says more than the plain mention rate because it accounts for competition. It is also more sensitive. If a system suddenly names ten providers instead of three, the share drops without your visibility having changed.

Citation rate

The share of answers in which your own domain is linked or named as a source. Mention and citation come apart. A brand can be recommended without its website appearing as a source.

Context and tone

Whether the mention comes as a recommendation, as one example among many, or with a caveat. This assessment is the hardest to automate and often the most useful.

Why two tools produce different numbers

There is no shared definition. Every vendor calculates differently, and the differences are larger than the effects being measured.

One example of counting rules. One vendor counts a mention once per answer, no matter how often the brand appears in the text. The same vendor counts citations individually. Two tools with different counting rules report different values for the same answer.

The measurement channel adds to this. One vendor documents explicitly that it measures through the official interfaces of the AI providers. That is clean and repeatable, but it does not show what a signed-in user sees in the interface. Other vendors say nothing on the matter.

The practical consequence. Numbers from different tools do not belong in one table. Within a single tool, the trend over time is meaningful. The comparison between tools is not.

Measuring yourself or buying a tool

Measuring yourself means keeping the questions in a spreadsheet, retrieving the answers manually and entering the results. That works up to around ten questions and two systems. At seven runs per question, that is already 140 retrievals per round.

Tools take that work off your hands. Four questions are worth asking when choosing one:

  • Is measurement done through the interface or through the user-facing product?
  • How many runs per question, and is that number stated?
  • How is the metric defined, with a formula?
  • How many requests does the plan actually include?

The last question deserves arithmetic. One vendor states consumption per answer and includes 2,500 units per year in a mid-tier plan. On an expensive system that comes to roughly 1,250 answers per year, which at daily measurement is about three questions.

The open disagreement between vendors

Two vendors hold publicly opposing positions.

SISTRIX deliberately skips repeat runs, arguing that repeated asking changes the wording but not the facts, because the underlying model stays the same.

seoClarity makes repetition a selling point and speaks of a statistically significant presence rate across thousands of variations. The vendor names no number of repetitions and no significance level.

The St. Gallen measurement argues against the first position, at least at source level. If simultaneous runs overlap by only a third to two fifths of their sources, more than the wording is changing.

One first screening question follows for your own selection. How many runs per prompt does the tool perform, and does it state the number?

What prompt tracking does not tell you

It names no causes

The measurement shows whether a brand is named. It does not show why. A rising mention rate can come from your own work, from a model change at the provider, or from a competitor dropping out.

It proves no business effect

No study with a documented method has measured a causal link between being named in AI answers and revenue. Correlations exist, and they come mostly from vendors selling measurement tools.

That is not an argument against measuring. It is an argument against the promises derived from it.

It is not complete

What gets measured is a selection of questions in a selection of systems at one point in time. Nobody knows the full picture of what users actually ask, because the providers do not release that data.

Measuring improves nothing

Prompt tracking is observation. Whether a brand appears in answers is decided elsewhere, in the clarity of your own information and in what independent sources report about you. The article on AI visibility covers that work.

Conclusion: understand the spread before you measure

Prompt tracking measures a moving target. The same question, asked at the same moment to the same system, produces answers whose sources overlap by only a third to two fifths. Measuring once measures chance.

A workable setup follows almost by itself. A small, fixed list of questions. Several runs per question. Constant conditions. Metrics with a stated reference. And the willingness to read small movements as what they usually are, which is noise.

Frequently asked questions

What is prompt tracking?

Repeatedly asking a fixed set of questions to AI systems to measure whether and how a brand appears in the answers. Each run records the mention, its context, the competitors named and the sources used.

How often should you measure?

Weekly is usually enough. More runs per measurement point matter more than more frequent measurement, because the same question returns different answers.

How many runs does one question need?

A University of St. Gallen study from April 2026 recommends at least seven runs per question per day, and at least eight when the cited sources are being analysed.

Why do I get different answers to the same question?

The output of a language model is not repeatable even with random selection switched off. The main reason is fluctuating server load, which changes the calculation in the final decimal places. Systems with web search add changing search results on top.

How many prompts should I track?

Usually 20 to 40. Each question is asked several times across several systems, so the effort grows quickly. A small, well-chosen list is more reliable than a long one asked once each.

What is the difference between branded and unbranded prompts?

Branded questions contain the brand name and measure what a system says about the brand. Unbranded questions leave it out and measure whether the brand is considered at all.

Does being signed in affect the result?

Yes. Google draws on earlier searches in AI Mode when history is enabled. ChatGPT can use memories from earlier conversations. Comparable measurements should be taken signed out and without history.

Are numbers from two tracking tools comparable?

No. There is no shared definition of the metrics, and vendors measure partly through interfaces and partly through the user-facing product. What is meaningful is the trend within one tool.

What does prompt tracking cost?

Prices range from around 119 euros per month at a vendor with a public price list to plans quoted only on request. The number of requests included matters more than the headline price.

Which systems should I track?

Google's AI features for reach and ChatGPT for deliberate use. In Germany, a Bitkom survey from November 2025 found around 5 percent using AI chats exclusively instead of a search engine, with another 7 percent using them predominantly.

Does prompt tracking show whether my work is paying off?

Only to a limited extent. It shows changes but names no causes. A rise can come from your own work, from a model change, or from a competitor dropping out.

Does prompt tracking improve my visibility?

No, it observes it. Visibility comes from clear information of your own and from mentions on sites the systems follow.

David Hahn

About the author

David Hahn

Managing Director, GetMentioned

David has been building link acquisition and digital PR processes since 2016, first as an agency under SEO Galaxy, today as a platform with GetMentioned. He has scaled his own projects from zero to seven-figure monthly traffic and delivered thousands of campaigns for clients. Here he writes about what works in practice, and about what only costs budget.

Your next link does not have to be a blind buy

Compare publishers, SEO data and prices in one place and book the placements that fit.

You might also like these

All articles →
An AI answer names three brands and leans on external sources to do it
Strategy

AI visibility: building presence in AI search systems

AI answers recommend providers before anyone clicks a ranking. How to measure AI visibility, which signals carry it, and where the external evidence models lean on comes from.

August 10, 2026 · 12 min read