Skip to content
GEOMeasurementMethod

How to run an AI citation baseline yourself

The measurement half of GEO, written so you can run it without hiring anyone: how to build the prompt set, what to record, how often to sample, and the mistakes that make the numbers meaningless.

SearchSynth Engineering9 min read
Analytics dashboard on a laptop showing bar charts and bounce-rate metrics
Photo: Luke Chesser on Unsplash

Most AI visibility reporting is not wrong so much as unfalsifiable. A composite "visibility score" moves and nobody can say why, or whether it would move again if you ran it twice. This is the method we use instead. It is not proprietary and you can run it without us.

Step 1: build the prompt set

The prompt set is the whole experiment. Get it wrong and nothing downstream means anything.

Use the phrasing buyers actually use with a chatbot. People type differently into an assistant than into a search box — longer, more conversational, and loaded with constraints. "crm" is a search query. "what CRM should a 12-person agency use if we already run Xero" is a prompt. Your rank tracker contains the former and almost none of the latter.

Good sources for real phrasing:

  • Sales call recordings — the questions asked before a demo is booked
  • Support tickets, especially pre-sales ones
  • The "what did you consider" answers in win/loss notes

Cover the families that precede a purchase. In practice these are: comparison (X vs Y), alternatives (alternatives to X), constrained recommendation (best X for Y under Z), and validation (is X any good, is X worth it). Informational top-of-funnel prompts are cheap to win and rarely worth anything.

Size it for repeatability, not coverage. Sixty to two hundred prompts is enough. The number matters less than never changing it once sampling starts — adding prompts mid-programme makes the trend line meaningless, so new prompts go into a clearly separated second cohort.

Step 2: decide what you record

Per prompt, per engine, per run. This is the part people cut, and it is the part that makes the result diagnostic rather than decorative.

FieldWhy it earns its place
MentionedNamed at all, with or without a link
CitedNamed with an attributed source link
Source URLWhich page earned it — tells you what to make more of
Competitors namedShare of voice needs a denominator
Verbatim answerLets you diff week to week, and catch misstatements
SentimentBeing named unfavourably is a distinct problem

Recording the full answer text is the highest-value field and the one most often skipped. Without it you can see that a number moved but never why.

Step 3: sample properly

Every engine, not one. Results diverge substantially between them, and a baseline drawn from a single assistant tells you about that assistant.

On a fixed cadence. Weekly is enough for most programmes. The cadence matters more than the frequency, because you are looking for a delta.

With sessions isolated. Personalisation and conversation history will contaminate results. Fresh session per prompt, no logged-in account carrying memory, and be aware that geography changes answers — pin a location if you serve one.

Repeat before you trust a change. These systems are non-deterministic. A prompt that named you last week and does not this week may be noise. Sample each prompt more than once per run, or treat single-run changes as unactionable.

Step 4: reduce it to two numbers

Everything above collapses to:

  • Citation rate — of the prompts sampled, the share that named you. Per engine and overall.
  • Share of voice — your mentions against those of a fixed, named competitor set. Fixed matters; a competitor set that changes when convenient is a reporting artefact.

Trend both monthly. Report the query-level detail beneath them so a movement can always be traced to specific prompts.

The mistakes that invalidate the result

Changing the prompt set. The most common one, usually well-intentioned.

Sampling once and calling it a baseline. One run is a snapshot of a stochastic system.

Counting brand-name prompts. Asking an engine about your company by name will usually get you named. It measures nothing about discovery. These belong in a separate accuracy check, not in the citation rate.

Reporting a composite score. If a number cannot be decomposed into which prompts moved on which engine, it cannot be acted on.

Attributing through last-touch analytics. Referral traffic from AI surfaces exists and is worth tracking, but it undercounts badly. Most of the effect arrives as a buyer who is already convinced. Add a "how did you hear about us" field — for most companies that is where it first becomes visible.

What to do with the result

The baseline tells you which of three situations you are in.

If you are absent across the board, check crawler access before anything else. Robots rules excluding the retrieval agents are the most common single cause and cost nothing to fix.

If you are mentioned but rarely cited with a link, the gap is usually extractability — the model knows of you but your pages give it nothing specific to attribute.

If competitors are consistently named and you are not, look at what is being cited. It is frequently a third-party page neither of you owns, and the work is to be present on those sources rather than to publish more of your own.

SearchSynth Engineering

The team that runs the engagements

Written by the engineers who do the work rather than a content team briefed on it. Methods described here are the ones we run on client engagements.

Want this run against your own queries?

We'll build the prompt set with your team and hand back the baseline matrix. Fixed fee, two weeks, no retainer attached.

30-minute strategy call

With an engineer, not a closer

Book a strategy call

Typical reply time: under 4 business hours.

hello@searchsynth.ai