Skip to content
LumiRank
SEO and GEO services
Service · 05

LLM Visibility Monitoring

Track how you appear inside ChatGPT, Claude, Gemini & Perplexity, then move the needle weekly.

MonitoringShare of voiceReporting
Overhead desk flat-lay: an open notebook, laptop, magnifying glass, cup of black coffee, pens and a film camera on dark wood

You can't improve what you can't see. LLM visibility monitoring measures how each major engine answers the questions that matter to you: whether you're named, how you're described, and who's named instead.

We turn that into a weekly operating rhythm: a clear share-of-voice baseline, movement tracking across engines, and the next set of actions to grow your presence.

At a glance

Prompt panels

A maintained set of the prompts your buyers use, run across every major engine on a schedule.

Share of voice

How often you're named versus competitors, measured, trended, and benchmarked.

Sentiment & accuracy

Not just whether you appear, but how correctly and favourably you're described.

Weekly reporting

Plain-language reports tied to the specific actions that move your numbers.

Want results like this for your brand?

Talk to a Growth Strategist

Setting a baseline before anyone touches anything

The single most valuable run is the first one, and it has to happen before any optimisation work starts. Once pages have been rewritten and crawler access has been fixed, there is no way to reconstruct where you actually stood.

So the opening fortnight of an engagement is measurement only. The panel gets built, run at least twice to establish how much the answers vary on their own, and recorded. That variance figure matters as much as the results: without it there is no way to tell a real improvement from an engine having a different afternoon.

Clients occasionally push to skip this and start fixing things immediately, which is understandable and a mistake. Two weeks of baseline buys the ability to prove everything that follows. Skipping it means every later report is an assertion.

Nobody can tell you where you stand without measuring it

Ask an agency whether ChatGPT recommends you and most will give you an impression. Impressions are worthless here, because the answers vary between runs, between accounts and between days, and a single lucky query proves nothing in either direction.

Monitoring exists to replace anecdote with a series. A fixed set of prompts, run on a schedule across every major engine, with each result recorded. One run is noise. Twelve weeks of runs against a stable panel is a trend, and a trend is something you can actually make decisions against.

This is also the only honest basis for reporting on GEO work. Without a baseline recorded before the work started, every later claim of improvement is unfalsifiable, which is convenient for the agency and useless to you.

What we actually measure

Four things, tracked separately, because they move independently and mixing them hides the story.

Presence: whether you are named at all in the answer to a given prompt. This is binary and it is the floor.

Share of voice: how often you are named against the competitors named instead. A prompt where three companies get listed and you are one of them is a different result from one where you are the only name, and both are different from absence.

Accuracy: whether what the engine says about you is correct. Being named with the wrong service list, a former location or a competitor's pricing is a problem that looks like success in a presence-only report.

Sentiment and framing: whether you are the recommendation or the also-ran in the same answer. Being mentioned as the expensive option is a mention, and it is not a win.

How the prompt panel gets built

The panel is the instrument, so its construction decides whether the whole exercise means anything. We build it from sources that reflect real demand rather than guesses: the questions your sales team fields on first calls, the language in support tickets, People Also Ask sets on your commercial queries, and the follow-up questions the engines themselves suggest.

Forty to sixty prompts is the usual size. Big enough to cover the decisions that matter, small enough to run often and read carefully.

The panel then stays fixed. This is the part that requires discipline, because a panel that gains prompts every month cannot show a trend, and quietly adding queries you have started winning is the oldest trick in visibility reporting. When the panel does need to change, the change is logged and the affected series is marked, rather than the numbers being silently rebased.

Engines behave differently and the report should say so

Treating AI visibility as one number averages away the information. The engines draw on different sources, update at different rates and cite differently, so a single score tells you nothing about what to do next.

Perplexity cites inline with visible links and still sends referral traffic, which makes it the surface where a win is most directly measurable. ChatGPT's search path and its training-derived answers behave differently from each other, and a brand can be strong in one and absent from the other. Gemini and AI Overviews lean on Google's index and on the Google-Extended permission, which is why a single robots.txt line can zero out that column. Claude has its own crawl and citation behaviour again.

Our reporting keeps them in separate columns for that reason. When a number moves, the first useful question is which engine moved and what changed in its inputs.

Turning measurement into a week of work

A dashboard nobody acts on is an expense. The point of the panel is that it produces a short list of specific things to do next, and that list is what the weekly report leads with.

The pattern is consistent. Prompts where you are absent but a competitor is named get read closely to find what that competitor has and you do not: usually a page answering the question directly, or a third-party source corroborating them. Prompts where you are named but described wrongly point at an entity problem rather than a content one. Prompts where you are named correctly and still not recommended usually point at missing proof.

Each of those has a different fix, which is why the four measures stay separate. A report that collapses them into one visibility score cannot tell you which of the three situations you are in.

What it costs and what it does not do

Monitoring runs inside the retainer for engagement clients and works as a standalone subscription for teams doing their own optimisation work. Retainers run from $299 to $1,999 CAD a month with custom scoping above that.

What it does not do is move anything by itself. Measurement is instrumentation. The gains come from the work the instrumentation points at, and a client who buys monitoring alone should expect a clear picture of the problem rather than a solution to it.

We would rather say that at the proposal stage than sell a dashboard as a growth programme.

Reading a bad result correctly

The most common mistake with this data is treating every absence as a content gap. Absences have at least four different causes and they need different responses.

Sometimes the engine cannot reach you, and the fix is a robots.txt line or a rendering change rather than anything editorial. Sometimes it reaches you and cannot tell which company you are, which is an entity problem. Sometimes it knows exactly who you are and has nothing from anyone but you, which is a corroboration problem and the slowest of the four to solve. And sometimes your page genuinely does not answer the question, which is the only one that a writing brief will fix.

Diagnosing which of the four you are looking at takes ten minutes and saves a quarter of misdirected work. It is the first thing we do with a new absence, and it is the reason the report names a cause rather than just a gap.

What a weekly report actually contains

One page, four blocks, and no chart that exists to fill space.

The panel result: presence, share of voice, accuracy and framing by engine, with the change from last week and a note on anything that moved more than normal variance. Then the classical picture from Search Console, kept in its own block. Then what shipped in the last seven days, named specifically enough that you could check it. Then the next actions, in priority order, with the reason each one is on the list.

What is deliberately absent is a composite visibility score. Those are popular because they always seem to be going up, and they are unfalsifiable by construction. If you want one number, the honest one is share of voice on the panel, and it does not move every week.

Clients get the underlying run data too, not just the summary. If we claim a prompt improved, you can read the answers from both runs and judge for yourself.

Frequently asked

LLM Visibility Monitoring: common questions

Why not just check ChatGPT myself?

You can, and it is worth doing for a feel. What it cannot give you is a trend. Answers vary between runs, personalisation and memory affect what you see, and a query you ran in a logged-in session is not the answer a stranger gets. Monitoring runs the same panel from a clean state on a schedule so the differences over time are real rather than artefacts of how you asked.

How often do you run the panel?

Weekly for the full panel on an active engagement, which matches how quickly the underlying sources actually change. Daily runs mostly produce noise. For monitoring-only subscriptions the cadence can be fortnightly or monthly, which is usually enough to catch a real shift without paying for movement that is within normal variance.

What is a good share of voice?

It depends entirely on the category, so any absolute number quoted at you is invented. What matters is the direction over a quarter and the gap to the specific competitors named instead of you. We benchmark against the names that actually appear in your panel's answers, because those are your real competitors in this channel, and they are frequently not the ones you would have listed.

Can you monitor competitors too?

Yes, and it is usually more useful than tracking yourself alone. The panel records every brand named in each answer, which shows you who the engines currently treat as the credible options in your category. That list is often surprising, and it is a better read on your real competitive set than anything a classical rank tracker will tell you.

Does this replace Search Console?

No, and we report the two side by side without merging them. Search Console measures classical search: impressions, positions, clicks. The panel measures whether AI engines name you. They move independently and a gain in one is not a gain in the other. Blending them into a single visibility figure is how agencies take credit for movement they did not cause.

Related

Case studies using this work

+1

Want to be the brand the models name?