Methodology

How We Measure AI Visibility

How often an AI assistant names you when buyers ask. This is the entire method, including the parts that stop it being a full answer.

What we measure

AI citation rate is one number: out of a fixed set of questions you should own, how many answers name your business. It is a share over a denominator you control, run the same way every time, so the trend carries meaning.

It is not a ranking. Nobody publishes AI answer rankings, and any vendor implying they have access to them is guessing. What this is, precisely, is a repeated sample of a system that varies. That's why the method has to stay frozen while the inputs change.

The fixed question set

It starts with a question list built from your real buying questions, not from a keyword tool. We draft it during the audit, you approve it, and then it barely changes.

  • Category questions: who does this work, best providers of the service, top firms for the job.
  • Geography questions, one set per market you actually sell into.
  • Comparison questions naming you against the competitors you lose to.
  • Problem-first questions, phrased the way a buyer talks before they know the category name.
  • Brand questions: what an engine says about you when asked about you directly.

Why the list stays locked

Change the questions and the number stops being comparable. New questions go into a separate tracking group and stay out of the baseline series until we deliberately re-baseline.

Twenty to forty questions covers most businesses. Multi-location operators run a set per market, which is how a 40-location restaurant group ends up with a number per market instead of one blended average that hides the weak cities.

One assistant, and what that costs you

We put the question set to one AI assistant, not to every engine on the market. That is a real limitation and it is worth being plain about, because plenty of vendors will quote you a per-engine scorecard.

  • Running one assistant well beats running four badly. A number you can reproduce is worth more than four you cannot.
  • You can check it. Take a question off your own list, ask it yourself, and see whether the answer matches what we reported.
  • What it costs: an assistant we do not measure could be answering differently, and we will not pretend to know.
  • If that changes and we add one, we re-baseline and say so in writing rather than quietly moving the number.

How each run is captured

Every run uses the same questions in a clean session, with no prior conversation to bias the answer, and we keep the raw response text next to the number. The text is the useful half. A number tells you that you slipped. The response tells you the assistant now describes you as a regional player when you sell nationally.

These systems change, and the method will change with them. When something gets added or dropped we re-baseline and say so in writing, instead of quietly moving the number.

How often we run it

On a set cadence we agree at the start, same method every time. Spacing the runs out is deliberate: these systems drift for reasons that have nothing to do with your marketing, and measuring too often manufactures noise that people then act on.

  • Month zero: a baseline before any work ships, so every later number has something to sit against.
  • Every run after: the full question set, raw responses kept.
  • Quarterly: a written read on what moved, what caused it, and what we do next.
  • On demand: a re-run after a large content push or a migration, tagged separately from the baseline series.

How the number works

It is a count over a fixed denominator, not a weighted index. Out of the questions in your set, how many answers name you. There are no weights to publish because there is no formula, and that is the point: a metric nobody can audit is a metric nobody should trust.

  • Presence: does the answer name you at all. Binary per question, and it is the headline, because absence is the one failure that makes everything else irrelevant.
  • Description accuracy: when you are named, does the assistant get your services, markets, and positioning right. A confidently wrong description can cost more than an omission.
  • Competitive set: who else appears, and in what order. This is the share-of-answer view, and it is the part clients argue with most productively.

Three records, one rate

Presence is what makes the number. Accuracy and the competitive set are recorded alongside it rather than blended into it, because rolling three different things into a single score is how a metric stops being checkable.

Multi-market clients get a rate per market plus a rolled-up figure. The rolled-up figure is for the board. The per-market numbers are what the work actually runs on.

We report the parts, never just the total. A total that moved five points is trivia. Presence up in nine markets with accuracy down in two is a work plan.

What this cannot tell you

This is the part most vendors skip, so here it is in full.

  • It is one assistant, not the whole field. Another could be answering differently, and this measurement will not see it.
  • It is not traffic and it is not revenue. AI answers rarely pass a referrer you can trust, so treat it as a leading indicator, not an attribution model.
  • It is a sample, not a census. Real buyers ask in their own words, with their own chat history behind them, and personalization means their answer is not exactly your answer. It also cannot isolate cause on its own: these systems update, competitors publish, and your content ships, often in the same month.
  • It does not prove a purchase. Nobody can currently trace an AI recommendation to a closed deal end to end. Anyone claiming they can is selling a model, not a measurement.
  • Run-to-run wobble is normal. Judge it on a quarter. A three-point dip between two runs is usually the system, not you.

Why we publish the limits

The number is only useful if you trust it, and trust needs the boundaries stated out loud. Used properly it does one job well: it tells you whether an AI assistant is getting more likely to recommend you, and what it says when it does.

The national restaurant group we run this for started from a baseline where most markets returned no mention at all. After 90 days of entity, citation, and content work the share of answers naming it was up 11 points. The pattern underneath mattered more than the number: the markets that moved first were the markets where the listings got fixed first.

Every number and every raw answer sits in your portal, on the same day we see it. There is no monthly PDF, and there is no version of this you have to take on faith.

Case metrics are illustrative placeholders pending client approval to publish named results.

Keep readingThe AI search practice this measurement runs insideThe case: per-market measurement across 40+ locations
Next step

Find out how you actually show up in AI answers.

You just read the method. A systems audit runs it on your business: a baseline score on every engine your buyers use, plus the entity and citation gaps behind it. You keep the findings whether you hire us or not.

Book a systems audit