Skip to content
LumiRank
Choosing a vendor
Software · Choosing a tool

How to evaluate an AI visibility tracking tool

No ranked product table, and a reason for that. What the measurement has to get right, the four traps, and when building it yourself is the better answer.

MonitoringToolsMeasurementBuyer's guide
Craft tools laid out on a teal desk — rolls of masking tape, scissors, a rotary cutter, a scalpel, a cutting mat and sticky notes

Tools for tracking whether AI engines mention your brand went from roughly nothing to a crowded category in about two years. Several are good. The problem for a buyer is that they are hard to tell apart from a demo, because every one of them produces a chart that goes up.

This page does not rank named products. Pricing and feature sets in this category have moved materially within the last two quarters, a comparison table would be wrong within one more, and most tool round-ups on the open web are funded by referral fees from the products being ranked.

What does not go stale is the measurement itself. Below is what a tool has to get right to be worth paying for, the four traps that make a weak one look rigorous, and the honest version of the build-versus-buy calculation.

At a glance

Frozen panels or nothing

If the prompt list can be edited silently, the trend line means nothing. Adding prompts you have started winning is the easiest way to manufacture progress.

Capture the description, not just the mention

Whether you were named matters less than how you were characterised and who was named beside you. Yes-or-no tracking throws away most of the signal.

Per engine, never blended

A single composite score across engines that behave differently hides the only actionable detail: which engine you are losing and why.

Export or it is not yours

Panel definitions and historical runs should leave in a usable format. A measurement you cannot take with you is a subscription with lock-in attached.

Why this page ranks nothing

Three reasons, all practical. The category is moving fast enough that feature and price comparisons published a quarter ago are now partly wrong, and a buyer acting on a stale table makes a worse decision than one acting on criteria. We sell a managed service that overlaps with what these tools do, so our ordering of them would not be disinterested. And the dominant business model for tool round-ups is affiliate referral, which reliably produces rankings shaped like commission rates.

The criteria below do not decay the same way. What a measurement of AI visibility has to do to be sound is a property of the problem, not of this quarter's products, and it will still be true when half the current names have merged or closed.

If you want a shortlist of products, the fastest honest route is to search the category, take the first five names that appear repeatedly across sources with no referral disclosure, and run every one of them through the table above. That takes an afternoon and produces a better answer than any ranking, including one we could write.

The four traps

The composite score with no published formula. Almost every product in this category offers a single number summarising your AI visibility. Some are reasonable weighted aggregates. Others are unfalsifiable. The test is simple: ask how it is calculated. A vendor who will not say has given you a number you cannot audit, cannot reproduce and cannot act on, and its main function is to move reassuringly.

The single run presented as a position. Ask an engine the same question three times and you can get three different answers. A tool that samples once per period and draws a line between the points is charting noise with confidence. What is meaningful is the trend across many runs of a stable panel, with the variance visible rather than smoothed away.

The quietly mutable panel. If prompts can be added or removed without a version record, the baseline is not a baseline. This rarely happens through bad faith — someone notices the panel missing an obvious question and adds it — but the effect on the trend line is the same, and only a visible edit log prevents it.

Share of voice with no denominator. A share figure requires a defined universe: which prompts, which engines, which competitor set. Presented without those, it is a percentage of something unspecified, and it can be moved by changing the denominator rather than the performance.

Three categories of product, and what each is for

Dedicated AI-visibility platforms. Built for this problem specifically, generally the strongest on panel management, per-engine breakdown and competitor capture. They are also the youngest companies in your stack, which is a real procurement consideration if you need the historical data to still exist in three years. Export capability matters most here.

Established SEO suites with AI modules added. The advantage is consolidation: one contract, one login, and AI metrics sitting beside the Search Console data they have to be read against. The trade-off is depth, since these modules are typically newer and shallower than the dedicated products, and panel control is often the first thing simplified away.

Do it yourself. A scheduled job that runs a fixed prompt list against each engine's API, stores the raw responses, and extracts mentions and competitor names. Entirely feasible for a team with an engineer, and it gives you complete control of the panel and permanent ownership of the history.

Most businesses over a certain size end up with two of the three, usually a suite for the search data and something dedicated or homemade for the panel, because the suites are rarely deep enough on their own and the dedicated products do not replace Search Console.

Build versus buy, honestly

Building is cheaper than it looks and more work than it sounds. The engineering is genuinely modest: a scheduler, API calls to each engine, a store for raw responses, and extraction logic for mentions and competitor names. A capable developer has a first version running in days.

The costs people forget are the ongoing ones. API calls for a real panel across several engines, run often enough to see through the variance, are an actual monthly line item. Extraction logic breaks when response formats change, which they do without notice. And somebody has to read the output every week and care about it, which is the part that quietly fails first — a homemade tracker nobody opens is worse than no tracker, because it produces the feeling of measurement without any.

Buy if you want the reading habit enforced by an interface someone else maintains, or if nobody on the team will own it. Build if the panel needs to be unusual, if the data must be yours permanently, or if you are already running scheduled jobs and this is one more.

Either way the panel is the asset and the tool is the container. Write the prompts from recordings of your own sales calls, freeze them, and keep them somewhere that survives changing vendors. That list is worth more than any dashboard rendering it.

What to demand before you sign

Six capabilities that separate measurement from decoration, and how to verify each one inside a demo rather than after a year of invoices.

Capability, why it matters, and how to check it in a demo
CapabilityWhy it mattersHow to verify it in the demo
Frozen, versioned prompt panelAn editable panel can be tuned toward flattering results, deliberately or notAsk to see the panel's edit history. If there is no version log, there is no baseline
Repeated runs with variance shownThe same prompt returns different answers between runs; a single run is not a positionAsk how many runs per prompt per period, and whether the spread is displayed or averaged away
Per-engine separationEngines differ in retrieval and citation behaviour; a blended score hides which one you are losingAsk to switch the view to one engine. If the product resists, the composite is the product
Competitor captureWho gets named instead of you is the most actionable output there isAsk to see the named-competitor list for a single prompt over time, not an aggregate share chart
Description captureHow a model characterises you is where entity and content problems show up firstAsk to read the raw answer text a mention was extracted from
Full exportPanels and history should outlive the subscriptionAsk for a sample export file before signing, not a promise that one exists

Deliberately no product names or prices: both change faster than this page could be maintained, and a stale comparison is worse than none.

Frequently asked

How to evaluate an AI visibility tracking tool: common questions

Why does this page not name the best tools?

Because the category changes faster than a comparison table can be maintained, because we sell a managed service that competes with these products and our ranking would not be disinterested, and because most tool round-ups are funded by referral fees from the tools being ranked. The evaluation criteria do not go stale the way a product table does, so that is what this page offers.

Can I just track this manually?

For a handful of prompts, yes, and it is a reasonable way to start. Write down ten questions your buyers actually ask, run them across the major engines on the same day each week, and record who gets named. It is tedious and it is real measurement. Tools earn their cost when the panel grows past what a person will genuinely keep doing, which is usually sooner than people expect.

What is the single most important feature?

A frozen, versioned prompt panel. Everything else in a tracking product is presentation layered on top of that list, and if the list can change without a record, none of the presentation means anything. Competitor capture is a close second, because who is named instead of you is the most actionable output the measurement produces.

Do these tools tell me why I am not being mentioned?

Generally not. They measure the outcome rather than the cause. A tool can show that an engine names three competitors and never you; working out whether that is a crawler access problem, an entity resolution problem or an absence of anything worth citing is diagnostic work the dashboard does not do. That gap is the reason tracking alone rarely changes anything.

How often should the panel run?

Often enough that between-run variance averages out, which in practice means at least weekly and ideally several runs per prompt per week. Monthly sampling on a system this noisy produces a line you cannot draw conclusions from. If a tool's pricing tier limits you to monthly runs, that tier is not measuring anything.