Skip to main content

AI Visibility | 11 min read

AI Search Analytics: Measuring Something That Leaves No Trail

By ยท Updated ยท 11 min read

What AI search analytics actually means

AI search analytics is the practice of measuring whether assistants recommend your business, when nothing about that recommendation reaches your existing reports. It goes under several names. AI search monitoring, AI citation tracking, AI visibility measurement and share of voice in AI search all describe roughly the same activity, and the differences between the labels are smaller than the volume of writing about them suggests.

The reason it needs a name at all is that the usual measurement layer is missing. Google gives you Search Console, where an impression is logged whether or not anyone clicks. When an assistant names your business inside an answer, nothing is logged anywhere. There is no impression count, no position, and no report you can open tomorrow to see what happened yesterday.

So the work is not analysis in the ordinary sense. It is closer to building the instrument first and then reading it. You decide what questions to ask, ask them on a schedule, and record the answers yourself, because no one is recording them for you.

The short version

You are not pulling a report. You are running an experiment on a schedule and writing down what happened, because the report does not exist.

Why none of it appears in your analytics

Three separate things break the trail, and they break it at different points.

The answer itself is never logged

When ChatGPT tells someone that your store is a good place to buy running shoes, that event exists only in that one conversation. It is not aggregated anywhere you can reach. A business can be recommended hundreds of times a month and have no way of knowing.

Most recommendations produce no click

This is the point people find hardest to accept. The answer often is the destination. Someone asks what to buy, gets three names, and goes to a shop or searches for one of those names directly later. Your analytics sees a direct visit, if it sees anything at all, and attributes it to nothing.

The clicks that do happen arrive badly labelled

Where a click does occur, the referrer that comes with it is frequently missing or mangled past recognition. Some of it can be recovered with care, and tracking referral traffic from ChatGPT in GA4 covers the practical side of that. The important thing to hold onto is that whatever you recover is a floor and not a total. It is the small visible portion of something larger.

What can genuinely be measured

Four things are measurable with reasonable honesty. Everything else on offer in this category is some presentation of these four.

Whether you are named

Put a question to an assistant and read whether your business appears in the answer. Binary, unambiguous, and the foundation of everything else. This is the same operation a visibility checker performs, and doing it once is a check rather than analytics. Doing it repeatedly on a fixed question set is where measurement starts.

Share of voice

Across your question set, how often are you named compared with the businesses you compete with. This is the most genuinely useful number in the category, because being absent matters far less than being absent while three rivals are present. It also survives the volatility better than any single result does.

Which sources the answer leaned on

Some surfaces name their sources. Perplexity numbers them on every answer, which makes it the one place you can reliably check your work, and how Perplexity decides what to cite goes into what it weighs. Where sources are visible, they tell you something no ranking report would: which pages, yours or anybody else's, the answer was actually built from.

Referral traffic, as a floor

Worth recording, worth never confusing with the total.

Of those four, share of voice is the one to build the reporting around, and it is worth being precise about why. An absolute count of how often you were named moves for reasons that have nothing to do with you. A model updates. A surface changes how many sources it lists. Your count halves and you have learned nothing about your own business. A comparative figure absorbs all of that, because whatever shifted shifted for your competitors too. If you were named in six answers out of twenty and the two rivals you care about were named in fifteen, that gap is a real fact about your position, and it stays a real fact next month when every raw number has moved.

Where a recommendation stops being measurable A shopper asks, an assistant answers, and three of the five steps leave no record behind. Shopper asks A question in their own words Assistant answers Names three or four businesses You are named Or you are not. No middle ground Maybe a click Often none at all from the answer A sale Attributed to something else No impression is logged Search Console records nothing. The answer happened invisibly Usually no click The answer was the destination. Nothing to count Referrer is unusable Arrives as direct, or mangled past recognition
Three of the five steps leave no record. Analytics here means rebuilding the missing ones deliberately.

What no tool can tell you, whatever it claims

Being straight about the ceiling is more useful than a longer feature list.

How many people saw it. There is no impression count. A tool reporting that you appeared in some number of AI answers is reporting how many of its own test queries named you, not how many real people were told about you. Those are different quantities and the second one is not available to anybody.

What real shoppers asked. Your question set is a model of how you think buyers speak. It is a good model or a poor one, and it is never the actual distribution of what people typed.

Whether a recommendation caused a sale. With no click and no referrer, the causal chain is broken at the first link. Anyone presenting revenue attribution from AI answers is presenting a model, and the model deserves the same scepticism you would apply to any other.

What the assistant will say tomorrow. The same question asked twice can produce different names. That is not a fault in the measurement, it is the thing being measured.

Why one run is not a result

This is the single most common way people mislead themselves here, and it is worth spelling out because the mistake feels like diligence.

You ask an assistant a question. You are not named. You conclude you have a visibility problem, and you spend a quarter on it. But assistants are not deterministic. The same question, asked again ten minutes later, can produce a different set of names, because the retrieval step ran again and returned slightly different pages, or because the model composed the answer differently from the same material.

A single absence is not evidence of anything. Neither is a single appearance, which is the more dangerous version, because it is the one people screenshot and send to their boss.

What produces a usable signal is repetition. The same question set, asked on a schedule, over enough runs that a pattern separates from the noise. Being named in one run out of one tells you nothing. Being named in two runs out of twenty tells you something real, and so does eighteen out of twenty.

The rule

A rate needs a denominator. If you cannot say how many times you asked, you do not have a measurement, you have an anecdote.

Building a question set that is worth asking

The question set decides the quality of everything downstream. A poor one produces confident numbers about nothing.

Use the words buyers use, not the words you use. Your category name and your product names are the language of your business. Shoppers describe problems. Somebody does not ask for a mid-weight merino base layer, they ask what to wear running when it is close to freezing.

Do not include your own brand name. An assistant asked about your brand will find your brand. That measures nothing except that your website exists. The useful questions are the ones where you would have to earn the mention.

Cover the question types that actually trigger answers. Not every search produces an AI answer, and the query patterns that trigger them are reasonably well understood. Comparisons, recommendations and how-to questions dominate.

Keep it fixed. A question set you edit every month cannot show you a trend, because every change resets the baseline. Add to it deliberately and rarely, and keep the original questions running unchanged alongside anything new.

Twenty to fifty questions is a sensible size for most businesses. Small enough to run often, wide enough that one odd result does not dominate.

Why two tools report different numbers for the same business

Run two products against the same business in the same week and you will often get different answers. That is usually not a defect in either one. It is a consequence of four choices each of them made.

They asked different questions. The largest single source of divergence. Each tool has its own set, and the sets are rarely published.

They asked different assistants. The surfaces do not agree with each other. In one run across thirty operator questions, Claude cited 213 distinct domains where ChatGPT cited 105. A tool weighted toward one surface reports a different world from one weighted toward another.

They asked at different times. Perplexity retrieves live, so its answer reflects the web today. ChatGPT shifts over longer periods. AI Overviews move roughly with Google rankings. A Tuesday reading and a Friday reading are not the same reading.

They count differently. A passing mention, a recommendation, and a linked citation are three different events. Whether all three count as a hit, and whether a mention in second place counts as much as one in first, is a choice each tool makes quietly.

None of this makes the numbers useless. It makes them incomparable across tools. Pick one, understand what it counts, and read the direction of travel rather than the absolute figure. What the tools do and what to ask them covers the buying side of this.

What to write down, and how often

A workable setup is smaller than most people expect. Four columns and a schedule.

The question, the assistant, the date, and whether you were named. That is the whole core of it. Add the competitors named in the same answer and you can compute share of voice, which is the number worth watching. Add the sources cited, where the surface shows them, and you learn which pages the answers are built from.

Weekly is usually the right cadence. Daily produces noise you will over-read. Monthly is too slow to connect a change on your site to a change in the answers.

Record what you changed, in the same place. This is the part almost everybody skips, and it is the part that makes the record worth keeping. Six weeks from now, the only question that matters is whether the thing you did moved anything, and you cannot answer it unless the two histories sit side by side.

Keep the runs, not just the summary. The temptation is to store a weekly percentage and throw the underlying answers away. Do not. When a figure moves, the only way to understand why is to read the answers from either side of the move, and see whether a competitor appeared, whether the assistant changed which sources it leaned on, or whether the question started being interpreted differently. A percentage on its own tells you that something happened and gives you no way to find out what.

Expect the slow half to be slow. Fixing whether crawlers can reach your pages can show up within days. Building the authority that gets you retrieved in the first place takes quarters. If your measurement runs for three weeks and shows nothing, the honest reading is usually that not enough time has passed, not that the work failed. An audit with pass and fail criteria is the faster diagnostic while you wait.

Whether you need to buy anything for this

Honestly, not at first.

A spreadsheet and an hour a week reproduces most of the value for a business with one site and one market. You ask your questions, you write down who was named, and after six weeks you have a real baseline that belongs to you and that you understand completely, because you built it.

Paying for something starts to make sense at the point where the manual version stops being feasible. That is usually one of three situations. You are running enough questions across enough assistants that the hour becomes a day. You need the history to be consistent across people, so it cannot live in one person's notebook. Or you want the surfaces watched more often than a person reasonably can.

What you should not buy is a number you cannot interrogate. If a product reports a score and cannot tell you which questions produced it, which assistants were asked, and how many runs it represents, then it is reporting a feeling. Ask those three things before anything else. What separates a platform from a dashboard goes further into that distinction.

The first real decision is not which product to buy. It is whether you have a baseline at all, because without one, nothing you do afterwards can be shown to have worked.

Frequently asked questions

What is AI search analytics?

It is the practice of measuring whether AI assistants recommend your business, when none of that activity appears in your normal reports. Because there is no impression log for an AI answer, it means running a fixed set of questions on a schedule and recording who gets named, rather than opening a dashboard someone else populated.

Can you track AI citations properly?

Partly. You can reliably track whether you are named across a question set you control, how often you appear compared with competitors, and which sources an answer used on surfaces that show them. You cannot track how many real people saw an answer, because no impression count exists anywhere.

Why do two AI visibility tools give different numbers?

Because they asked different questions, asked different assistants, asked at different times, and count a mention differently from a citation. All four are quiet choices made inside each product. The numbers are not comparable across tools, so pick one and read the direction of travel rather than the absolute figure.

How often should I measure AI visibility?

Weekly suits most businesses. Daily produces variation you will over-interpret, because the same question can return different names ten minutes apart. Monthly is too slow to connect a change on your site to a change in the answers.

Does AI search traffic show up in Google Analytics?

Some of it, and always less than the real figure. Most recommendations produce no click at all, and many of the clicks that do happen arrive with a missing or mangled referrer. Treat whatever you can recover as a floor rather than a total.

MG
Written by

Matt is the founder of RunOctopus. He built All Angles Creatures from zero to page-1 rankings in reptile feeder insects using exactly this method. Turning a hard, entrenched niche into RunOctopus's proof store for programmatic SEO and AI search citation.

Connect on LinkedIn →

Ollie builds this for your store automatically

A complete launch build . 8 expert guides, 6 collection pages, and an interactive tool. Structured for both Google and AI search. Live on your store in 48 hours.

See What Ollie Builds →

See what Ollie builds before you pay. Cancel anytime.

Trusted by store owners in 20+ niches