The measurement problem of AI search

Classic rank trackers don't answer the question about AI visibility: generated answers have no positions, outputs are non-deterministic, partly personalized and different per system. Even so, AI visibility can be measured reliably — not as an exact number but as a trend under controlled conditions.

I developed and refined the following method in GEO projects for a regulated broker and several online stores. It needs no special tool — a spreadsheet is enough to start with.

The method at a glance

  • Step 1: Define the prompt set (20–50 real customer questions)
  • Step 2: Define systems and measurement conditions
  • Step 3: Query regularly and log
  • Step 4: Evaluate — share of voice and trends
  • Step 5: Derive actions

Step 1: Define the prompt set

Collect 20–50 questions your customers really ask — not the ones you would like them to ask. The best sources: sales and support conversations, Search Console queries and the FAQ sections of competitors. Phrase each question the way a person would type it, and cover four categories:

  • Brand prompts: “Is [brand] trustworthy?”, “[brand] reviews”
  • Comparison prompts: “[brand] or [competitor] — which is better?”
  • Purchase-intent prompts: “Which provider for [service] in Germany?”
  • Knowledge prompts: “How does [topic] work?” — here content wins, not brands
Example · Regulated broker

“Which forex broker is BaFin-regulated in Germany?” · “What does trading cost at [brand]?” · “forex.com or [competitor] for beginners?” — the more concretely the prompts are tied to real decision situations, the more valuable the result.

Step 2: Define systems and conditions

As a rule, four systems are relevant: ChatGPT (with web search enabled), Gemini, Perplexity and Google's AI Overviews / AI Mode. What matters is keeping the conditions constant: a new chat window without history, German language, German location, documented model version. Only then are measurements comparable over months.

Step 3: Query and log

A monthly cycle is enough for most brands. For each prompt and system, exactly four fields are recorded:

  • Mention: Does the brand appear in the answer? (yes/no)
  • Role: recommendation, neutral mention or side note?
  • Statement: Are the facts stated correct?
  • Source: Which URLs does the system cite?

No more fields are needed — every additional column lowers the probability that the measurement will be sustained. The most valuable field is the source: it shows where you need to be present in order to be cited.

Step 4: Evaluate — share of voice and trends

The central metric is share of voice: the share of prompts in which the brand appears — per system and over time, ideally with the two most important competitors as a benchmark. In addition: the error rate of statements (how often does the system misrepresent the brand?) and referral traffic from AI systems, easy to segment in GA4 (chatgpt.com, perplexity.ai, gemini.google.com).

Because answers are non-deterministic: the trend counts, not the single measurement. For critical prompts, query 2–3 times per measurement and average the result.

Step 5: Derive actions

Only the derivation turns monitoring into a management tool. The three most common finding-action pairs:

  • No mention for purchase-intent prompts: content gap — build answer-first blocks and comparison content for exactly these questions.
  • Wrong statements about the brand: establish fact consistency — bring website, profiles and directories to identical figures and wording.
  • Systems cite third-party sources in which you are missing: source work — build presence in exactly these industry publications and directories.

Limits of the method

The method delivers no absolute truth: answers vary by wording, location and model version, and a prompt set is a sample, not a census. But it delivers what counts for decisions — directionally reliable trends under the same conditions. And it enforces discipline: process beats software, especially as long as the tool landscape changes monthly.

Conclusion

AI visibility is measurable — today, without a budget request. A clean prompt set with a clear protocol answers the questions that matter after two or three measurements: do we appear? Is what it says correct? And where do we need to be present? If you want to read up on the basics: What is GEO? explains the terms behind it.