What this category actually is
AI visibility tools ask AI assistants questions on a schedule and tell you whether your brand was named. That is the whole product in most cases, and it is a genuinely new thing to be able to measure, because nothing in your analytics reveals it.
The category appeared quickly and is unusually young. Most of the products in it are under two years old, several are under a year, and the vocabulary has not settled. The same product might describe itself as AI visibility, AI search monitoring, brand mention tracking or answer engine optimization depending on who wrote the homepage.
That youth matters when you evaluate them. There are no long track records to check and no established benchmarks to compare against. What you can do is interrogate the method, which is what this page is for.
Nearly every tool in this category measures the same thing in slightly different ways. The differences that matter are the question set, the engine coverage, and how honestly each handles variance.
The two kinds of tool, and why the difference matters
Everything in this space falls on a spectrum between two poles, and knowing which end you are buying prevents most of the disappointment.
Monitoring tools
These tell you what is happening. They run a question set, record who gets named, chart it over time, and often break it down by engine and by competitor. That is genuinely useful and it is where the large majority of the category sits. What they are counting is AI citations, and the counting is the easy half.
Their limit is that measurement does not change anything. A dashboard telling you every week that you are absent from your category is accurate and, on its own, inert. If you buy a monitoring tool, budget for the work it will point at as well as the tool itself.
Execution tools
These help produce the thing the measurement says is missing, usually content. They are rarer, harder to build, and easier to get wrong, because generating pages against a gap is only useful if the gap is real and the pages are good.
The combination of both is unusual and is the thing genuinely worth paying for: something that finds where you are absent and then helps close it, with the measurement feeding back so you can tell whether closing it worked.
There is a third group worth naming because it is easy to mistake for the first two. Some products in this space are really search tools that have added an AI panel, and some are agencies selling a service with a dashboard attached. Neither is a problem in itself, but they behave differently. A search tool with a bolted-on panel usually checks one engine shallowly and treats AI as a feature rather than the subject. An agency dashboard is often thin software wrapped around people doing the work manually, which can be exactly what a small team wants, as long as everyone is clear that is what is being bought.
The way to tell quickly is to ask how many questions the product runs, against how many engines, how often. Anything vague there is usually vague for a reason, and a vendor who will not name a number is telling you something about the number.
Why two tools give you different numbers
Run two tools on the same brand in the same week and they will disagree, sometimes wildly. Four reasons, all of them legitimate.
Different questions. Each tool picks or generates its own set. A tool asking twenty questions about your product category and one asking twenty about your brand name are measuring different things entirely.
Different engines. Coverage varies, and so does how each tool talks to each engine. Asking ChatGPT through its interface, through its API, and with search explicitly enabled can all produce different answers.
Different weighting. Some tools weight questions, some treat them all equally, and the vendors rarely say which. It is worth asking, because it changes what the headline number is actually summarising.
Different sampling. Once a week and once a day produce different pictures, especially in a category where the answers move.
None of this means the tools are broken. It means a score is only comparable to itself, and switching vendors means starting your history again.
The practical consequence is worth stating plainly, because people run into it after they have already committed. If you are ever asked to justify a number to somebody outside your team, you cannot point at an industry benchmark, because there is not one. What you can do is show the direction of your own line over enough months to be meaningful, and show the competitor lines from the same tool alongside it. Comparison within one tool is legitimate and useful. Comparison across tools is not, and presenting it as though it is will eventually be challenged by somebody who checks.
What to ask before you buy anything
Six questions. A tool that answers all six clearly is worth considering, whatever it costs.
Can I write my own questions, and can I see the whole set? If not, the most consequential decision has been made for you and you cannot audit it.
Which engines, and can I see them separately? A blended number hides the case where you are strong in one and absent from another, which is the most useful thing to know.
Do you keep the full answer text? Presence is a weak signal on its own. Being named dismissively and being recommended are very different outcomes that both count as present.
How do you handle run-to-run variance? Ask what the noise floor is. A vendor who has not measured their own variance is reporting precision they do not have.
What does a zero mean in your product? Genuinely absent and every call failed should not look identical.
What happens when I leave? Can you export the raw question-and-answer history, or only the charts. The raw history is the asset.
Whether you need a tool at all
A tool is worth it at a certain scale and genuinely is not below it, which vendors rarely tell you.
You probably do not need one yet if you are checking for the first time. Twenty questions, two engines, a spreadsheet, one afternoon. That gives you a real baseline and teaches you what the tools are compressing. Doing this once makes you much harder to sell to.
You probably do want one if you are tracking monthly across several engines and competitors. The manual version stops being reasonable somewhere around a few hundred question-and-engine combinations a month, and consistency matters more than heroics.
You definitely want one if somebody other than you needs to see the trend. Reporting to a client or a board is a real use case and a chart nobody had to assemble by hand is worth paying for.
You do not need one to know what to do first. The first actions are almost always the same regardless of what any tool says: confirm the AI crawlers can read your site by checking your robots.txt, and write down the twenty questions your buyers actually ask. The store-side guide to AI search covers what to do with the answers.
Setting one up so it earns its keep
A tool bought and configured badly produces a chart nobody trusts and everybody ignores within two months. Four decisions at setup determine whether that happens.
Write the question set before you connect anything. Twenty questions in a buyer's words, covering shortlists, comparisons and constraints rather than twenty variations of the same one. If you let the tool generate them, you will be tracking a set shaped by your own marketing language.
Add your real competitors by name, not the ones you wish you had. The useful comparison is against whoever is actually being named in your category's answers today. Run a manual check first to find out who that is, because it is often not the brands you benchmark against commercially.
Pick a cadence you will not react to. Weekly tracking on a monthly-moving signal produces meetings about noise. Monthly is usually right, with a longer view for anything you report externally.
Decide up front what would make you change course. Write it down before you have data. Without a threshold agreed in advance, every reading gets interpreted to fit whatever anyone already believed, and the tool becomes decoration.
The setup that goes wrong most often is the one where somebody connects the tool in an afternoon, accepts every default, and discovers three months later that the questions never matched how buyers speak. Redoing the question set restarts the history, so the afternoon you spend on it at the start is the cheapest hour in the whole exercise.
What this actually costs, and what drives the price
Pricing in this category is unsettled and the headline number rarely tells you what you will pay, because the thing that scales is not the thing on the pricing page.
What actually drives cost is questions times engines times frequency. Twenty questions across four engines once a week is roughly three hundred and twenty answers a month. Two hundred questions across four engines daily is around twenty-four thousand. Those are different products at the same nominal subscription, and most plans are shaped by that volume even when they are presented as tiers of features.
Competitor tracking usually multiplies it. Watching five competitors across the same set is often five times the underlying work, and some vendors price it that way while others fold it in. Ask directly.
The cost that catches people out is the work the tool creates. A monitoring product that tells you every week which questions you are absent from is generating a content backlog. Budget for the writing, or the subscription buys you a recurring reminder of a problem you are not addressing.
Worth checking before signing anything: whether historical data survives a downgrade. Losing a year of trend because you moved to a cheaper plan is a genuinely painful way to learn that question.
There is also a build-versus-buy calculation that is less obvious than it looks. Every engine in this category has an interface you can call directly, so running your own checks is technically straightforward, and for a modest question set the direct cost is small. What you are actually buying from a vendor is not access. It is the scheduling, the storage, the handling of engines that fail or rate-limit mid-run, the parsing of an answer into whether a brand was named, and the charting. Those are all unglamorous and all genuinely fiddly, which is why the category exists at a price that looks high against the raw cost of asking the questions.
The honest guide is scale. Below a few hundred answers a month, running it yourself is usually cheaper, and it teaches you far more about what the numbers mean. Above a few thousand, the operational side is real work and paying somebody to do it properly is reasonable. In between it comes down to whether the person who would maintain it has anything better to do, which for most small teams is a short conversation.
What none of them do yet
Three genuine gaps in the category as it stands, worth knowing so you do not go looking for a product that does not exist.
Attribution. No tool can reliably connect an AI mention to a sale, because the path from recommendation to purchase usually goes through a separate branded search with nothing carrying the attribution. Anyone claiming otherwise is modelling rather than measuring.
Cross-market data. Almost every tool measures one brand for one customer. Very few publish anything about how citation behaves across a whole market, which means the benchmarks everyone would like do not really exist yet.
Telling you why. A tool can tell you that you are absent. Working out whether that is a crawler problem, a retrieval problem born of thin topical authority, or a quotability problem is still manual, and the optimization guide covers how to separate them. It is also the part where the answer actually changes what you do next.
The category measures well and acts badly. Buy for measurement, budget separately for the work, and write your own questions whatever you buy.