Back to thinking
Technology 29 Jul 2026

How to Measure AI Visibility: Prompt Panels, Share of Voice, and What the Tools Do Not Tell You

There is no rank to track, because a generated answer has no positions and no two runs agree. The workable unit is share of voice across a fixed panel of prompts your buyers would actually type. How to build the panel, what to record, how often to run it, which four numbers to report, and why a screenshot of ChatGPT naming your brand is evidence of nothing.

How to Measure AI Visibility: Prompt Panels, Share of Voice, and What the Tools Do Not Tell You

You cannot rank-track a generative answer. There is no position one, the output changes between runs of the same prompt, and it varies by phrasing, account, and day. Every reporting habit built for the ten blue links breaks here.

What replaces it is share of voice across a fixed panel of prompts: across forty questions your buyers would plausibly type, how often are you named, how often are you cited with a link, and is what the model says about you actually true. That is measurable, it trends, and it survives the non-determinism. This is how we run it.

If the concepts are new, SEO vs GEO sets up the definitions and the 2026 playbook covers the work this measures.

Why rank tracking does not transfer

Three properties of generated answers break the old instrumentation.

Non-determinism. Ask the same question twice and you can get different sources. A single observation tells you almost nothing, which is why one flattering screenshot is not a result and one disappointing screenshot is not a crisis.

Fan-out. Google's own description of its generative features includes expanding one question into several internal searches. You are not competing on the query you typed. You are competing across a cloud of sub-queries you never see.

No position. An answer either names you or it does not. There is no third place. The gradient that rank tracking measured simply is not there, so the metric has to become frequency across many questions rather than position on one.

The consequence: sample size is the whole game. Forty prompts observed monthly is data. One prompt observed once is an anecdote.

Step 1: Build the prompt panel

Thirty to fifty prompts, written the way humans type, spread across four intent bands.

Unbranded research (roughly half the panel). "What does a brand identity cost in India?" "How do I choose a packaging designer?" These are where you are discovered by people who do not know you exist, and they are the hardest and most valuable to win.

Comparison (about a quarter). "Best branding agencies in Bangalore." "Agency vs freelancer for a website." Generative answers lean heavily on comparison and listicle content, so this band tells you whether you are present in the pages that get quoted.

Branded (about a sixth). "What is NOW Media?" "Is NOW Media any good?" This band measures accuracy rather than presence. A model that invents your founding year or your services is a live problem, because that invention is what a buyer sees.

Bottom-of-funnel (the remainder). "Who should I hire to redesign FMCG packaging in India?" Small volume, closest to money.

Source them from real inputs, not imagination: the questions in your sales calls, your support inbox, the queries in Search Console, and the phrasings competitors target. A panel you invented at a desk measures your imagination.

Freeze the panel once written. The point is comparability month over month, so prompts get added at review, never quietly edited.

Step 2: Decide what you record

Per prompt, per engine, four fields. Anything more and nobody maintains it.

  1. Named. Did the answer mention your brand at all? Yes or no.
  2. Cited. Did it link to your domain as a source? Yes or no. Named-without-cited is common and worth separating, because the fixes differ: being named is reputation, being cited is retrieval.
  3. Accurate. Was what it said about you true? Flag any error verbatim; these are your highest-priority work items.
  4. Competitors named. Who else appeared. This is how you get a denominator.

Run across the engines that matter for your buyers: ChatGPT, Google AI Mode and AI Overviews, Perplexity, and Claude. Log the date and the engine. A spreadsheet is genuinely adequate.

Step 3: Run it properly

Three runs per prompt, not one. Non-determinism means a single pass mixes signal with noise. Three passes and counting a mention if it appears in at least two gives you a stable reading at manageable cost.

Monthly. Weekly is noise for most businesses and burns the time you should be spending on the work. Anything slower than monthly and you cannot tell a fix from a fluctuation.

Logged out, clean session. Personalization contaminates the reading. Same conditions every time matters more than perfect conditions.

Same person or same script. Judgment calls about "was this a mention" need to be made consistently.

Step 4: The four numbers to report

Everything above compresses into four figures that a founder can act on.

Presence rate. Percentage of panel prompts where you were named at all. The headline number. Track the trend, not the absolute.

Citation rate. Percentage where you were named *and* linked. This is the one that produces traffic, and it is the one that responds to on-site work: crawler access, indexing, passage structure.

Accuracy rate. Percentage of branded prompts where everything said about you was true. This should be near 100%. Anything else is urgent, because the wrong answer is being delivered with total confidence.

Share of voice. Your mentions as a proportion of all brand mentions across the panel. This is the competitive number and the one to put in front of a board, because it moves independently of category growth.

Read them together. Presence rising while citation stays flat means you are being talked about but not read: an off-site reputation win that has not converted into retrieval. Citation rising while presence stays flat usually means you fixed a technical fault and the ceiling is now reputation.

Step 5: The instrumentation behind the panel

The panel measures outcomes. Two other sources tell you whether the machinery works.

Server logs are the leading indicator. Count hits from GPTBot, OAI-SearchBot, ClaudeBot, and PerplexityBot monthly. Crawler visits precede citations, which precede referrals. If crawler hits are zero, nothing downstream can improve, and you have found your problem before spending a rupee on content.

Referral traffic is the lagging indicator, and it undercounts. Segment AI referrers in analytics. Expect the number to look small and to understate reality: many AI answers end the journey, and some referrals arrive without a usable referrer and land in direct traffic. Judge these visitors on quality rather than volume, since AI-referred traffic converts at several times the rate of standard organic per Semrush's 2026 data.

What the tools do and do not do

Paid AI visibility platforms automate the panel: they run prompts on a schedule, parse mentions, and chart it. That is real work saved once a programme is running, and worth paying for at that point.

What they cannot do, and what nobody should let a dashboard imply:

  • They do not fix anything. No tool clears a WAF rule blocking GPTBot, restructures a pricing page, or earns a mention in trade press.
  • Their default panels measure their prompts, not your market. A generic prompt list produces a number that moves without meaning. Load your own.
  • They cannot see inside the model. Everything is inferred from sampled outputs, same as your spreadsheet, with better plumbing.
  • Their absolute scores are not comparable across vendors. Two tools will give you two different visibility scores for the same brand in the same week. Only the trend within one tool means anything.

Buy one after the work is underway. Buying one instead of the work is the most common way to spend money on this and get nothing, which we cover alongside the other failure modes in GEO mistakes.

The trap: optimizing the panel

The failure mode of every measurement system is that people optimize the measurement.

You will be tempted to add prompts you already win, drop the ones you lose, and watch the line go up. The panel is a sample of your market, and a sample you have curated for flattery is worth precisely nothing.

Two guards. Keep the unbranded research band at roughly half the panel, since it is the hardest and the one you would naturally prune. And review the panel once a quarter with someone who does not own the number.

The honest check is whether enquiries follow. If presence and citation both climb for two quarters and nothing reaches your inbox, either the panel is not your market or the pages the citations land on are not doing their job.

What good looks like

Month 1. Baseline recorded. Expect it to be worse than you hoped, and expect at least one factual error about your business in a branded answer.

Month 3. Technical faults fixed and structural rewrites landed. Citation rate should be moving first, because it is the part you control directly. Accuracy should be at or near 100% by now.

Month 6. Presence rate moving on unbranded research prompts, which is the slow one, because it depends on third-party sources accumulating. Share of voice starting to be a number worth showing someone.

Anyone promising a step change inside a month is either fixing a blocked crawler, which genuinely is that fast, or showing you a screenshot.

The one-sentence version

Freeze forty prompts your buyers would type, run them across four engines three times a month, and report presence, citation, accuracy, and share of voice: four numbers that survive the fact that no two answers are ever the same.

We set this up as part of website and AI visibility engagements, including the baseline and the monthly panel. See how we work, or start a scope.

FAQ

How do I track my brand in ChatGPT?

Build a fixed panel of 30 to 50 prompts your buyers would type, run them monthly in a logged-out session, three passes each, and record whether you were named, whether you were cited with a link, whether the claims were accurate, and which competitors appeared. Single spot checks are unreliable because the same prompt returns different sources between runs.

What is share of voice in AI search?

Your brand's mentions as a proportion of all brand mentions across a fixed prompt panel. It replaces keyword ranking as the headline metric, because generated answers have no positions to occupy. It is meaningful only against a frozen panel and a consistent method.

How often should I measure AI visibility?

Monthly for most businesses. Weekly is mostly noise given how much generated answers vary run to run, and it consumes time better spent on the underlying work. Slower than monthly and you cannot attribute a change to anything you did.

Why does ChatGPT give different answers to the same question?

Generated responses are non-deterministic and additionally vary with phrasing, session, personalization, and live retrieval results. This is why measurement has to be frequency across many prompts and repeated runs rather than a position on one query.

Are AI visibility tools worth paying for?

As instrumentation, once the actual work is underway, yes: they automate the panel and the mention tracking. They fix nothing, their default prompt lists measure their prompts rather than your market, and their absolute scores are not comparable between vendors. Only the trend inside one tool means anything.

Why is my AI referral traffic so low in analytics?

Two reasons, and both are expected. Many AI answers resolve the question so the user never clicks, and a share of the clicks that do happen arrive without a usable referrer and get bucketed as direct. Treat referral volume as a lagging, undercounting indicator and judge those sessions on conversion quality instead.

What do I do if an AI model says something false about my business?

Fix the sources rather than arguing with the output. Publish one canonical paragraph of facts on your own site, propagate it identically across LinkedIn, Crunchbase, Google Business Profile and your press coverage, and get corrections into the third-party pages being cited. Corrections typically surface over four to twelve weeks as those pages are recrawled.

NOW Media is a Bangalore creative studio founded in 2019, a brand of Bleep Design Private Limited. We run this panel on our own brand and products before selling the method to anyone.

More in this series

AI search and GEO

What generative engine optimization actually is, how to get retrieved and cited by ChatGPT, Gemini, Perplexity and Claude, and which tactics are a waste of money.

  1. SEO vs GEO: What Generative Engine Optimization Actually Is, and What Actually Changed Start here
  2. How to Rank in ChatGPT, Gemini, Perplexity and Claude: The 2026 GEO Playbook
  3. Do You Need an llms.txt File? The Honest Answer for 2026
  4. GEO Mistakes: 11 Things Not to Do, and How to Spot an Agency Selling You Nothing
  5. Building Products Designed to Be Cited by LLMs: A 2026 Playbook
  6. How Brandauditor.ai Hit Its First ChatGPT Citation in 2 Days (And the Exact Setup We Used)
  7. Brand Guidelines That AI Tools Can Actually Use (Without Breaking the System)