Back to blog
MethodologyJuly 5, 2026·9 min read

How to measure your pharma brand's visibility in AI answers — the five metrics that matter

AI visibility is not one number. Here are the five metrics that actually tell you how ChatGPT, Perplexity, Gemini, and Claude portray your brand — visibility, share of voice, position, sentiment, and the one most tools skip: accuracy — and what a good Share of Voice really looks like in pharma.

Every brand team we talk to arrives at the same question within the first few minutes: how visible are we in AI? It is the right instinct. Patients, prescribers, and payors are asking ChatGPT, Perplexity, Gemini, and Claude about conditions and treatments every day, and the answers they get shape decisions long before anyone lands on your website. But “how visible are we” is not one number, and treating it like one is how teams end up with a dashboard that looks reassuring and tells them nothing.

AI visibility is five measurements, not one. Each answers a different question, each can move independently, and in pharma one of them — the one most tools ignore — matters more than all the others. Here is the full set, what each one actually tells you, and how to read “good.”

1. Visibility: are you named at all?

The base layer. For a given question, does the engine mention your brand in its answer? Measured across a fixed set of prompts and a fixed cadence, visibility is the share of answers in which your brand appears at all — a simple percentage that tells you whether you are in the conversation or invisible to it.

The trap is measuring it once. AI answers are non-deterministic; ask the same question twice and you can get two different answers. A single screenshot is an anecdote. Visibility only means something as a rate, sampled repeatedly over a stable prompt set, so that a change in the number reflects a change in the engines rather than the roll of the dice. And in pharma it has to be split by audience: you can be highly visible in HCP-framed answers and absent in the patient-framed versions of the same question, which is a very different problem than a flat “we're at 60%.”

2. Share of voice: how much of the answer is yours?

Visibility tells you whether you show up. Share of voice tells you how much room you take up relative to everyone else who shows up. Of all the brands named across your prompt set, what fraction of those mentions are yours? This is the competitive metric — the one that turns “we're mentioned” into “we're winning, or losing, against these specific competitors.”

The question everyone asks next is “what is a good share of voice percentage,” and the honest answer is that there is no universal target. Share of voice is relative to how crowded the category is. In a class with two or three real players, a leader might hold 40–60%. In a crowded class where generics, guideline bodies, and patient forums all get named, being in 20–30% of answers can be category-leading. What matters is the trend against the competitors who actually appear in the same answers — not a round number borrowed from someone else's category.

And the pharma caveat is not optional: share of voice is worthless if the voice is wrong. A brand that is named in every answer with an off-label or outdated statement does not have enviable share of voice — it has a compliance exposure with excellent reach. Never read this number without the accuracy number beside it.

3. Position: how prominently are you named?

Being mentioned in the third sentence of a paragraph and being the headline recommendation are not the same outcome, and a visibility percentage flattens the two. Position captures where in the answer your brand lands — first choice, one of several options, or a caveat at the end. Tracked as a median across your prompt set, it tells you whether the engines treat you as the answer or as an also-ran, even when visibility looks healthy.

4. Sentiment: how are you framed?

The same brand can be named favorably (“a well-tolerated first-line option”) or unfavorably (“associated with significant side effects”) in answers that both count as a mention. Sentiment scores the framing — positive, neutral, negative — so you can see not just whether the engines talk about you but how they talk about you. In practice sentiment is where competitive positioning and safety perception show up first, often before anything changes in the market.

5. Accuracy: is what it says correct?

This is the metric that separates pharma AEO from every consumer version of it, and the one most tools do not measure at all. Accuracy asks whether the engine's statement about your brand is actually true — checked against your approved, referenced claims: the label, the pivotal-trial endpoints, the guideline positioning. A generic AI-visibility tool will happily report that you are named in 80% of answers and never notice that a third of those answers describe an indication you were never approved for.

In a regulated category, accuracy is not a nice-to-have layer on top of visibility — it is the whole reason to measure at all. An engine that names your brand in every answer while telling patients something the label does not support is not a marketing success with an asterisk; it is a pharmacovigilance and MLR problem that happens to have wide distribution. Accuracy is the metric that routes a finding to Medical Affairs instead of to the celebration channel.

Why one number is never enough

The reason to hold all five is that they interact, and the interactions are where the real signal lives. High visibility with low accuracy is the most dangerous quadrant in pharma — maximum reach for a wrong statement. High visibility with low position means the engines know you exist but do not lead with you. High share of voice with negative sentiment means you are winning the mention and losing the framing. No single metric surfaces any of these; only the combination does, read audience by audience.

That is also why “good” is a direction, not a threshold. The useful question is never “are we at the right number” but “are visibility and share of voice rising, is position improving, is sentiment holding, and is accuracy clean — for each audience that matters, over time.” A rising line on all five, against the specific competitors who share your answers, is what progress looks like. A high number on one, read alone, is how teams talk themselves into a false sense of coverage.

Common questions

How do you measure AI visibility?

Ask the AI engines the questions your audiences actually ask, on a fixed cadence, and score every answer on five things: whether your brand is named at all (visibility), how often you are named versus competitors (share of voice), how prominently you appear when named (position), how the answer frames you (sentiment), and whether what it says is correct (accuracy). One-off spot checks are not measurement; a stable, repeated prompt set is.

What is a good share of voice percentage in AI answers?

It depends entirely on how crowded the category is. In a two-or-three-brand class, a leader might hold 40–60%. In a crowded class with generics and guideline bodies in the mix, being named in 20–30% of answers can be category-leading. The honest benchmark is relative: are you gaining share against the specific competitors who show up in the same answers, over time. And a high number means nothing if the mentions are inaccurate — in pharma, 100% share of voice built on a wrong statement is a liability, not a win.

What does 50% share of voice mean?

It means that across the answers where any brand in your class is named, yours accounts for half of the mentions. It is a relative measure, not an absolute one — 50% in a two-brand race is parity, while 50% in a ten-brand class is dominance. Always read it next to visibility (how often the class is discussed at all) and against the named competitors, not a generic target.

What is the difference between visibility and share of voice?

Visibility asks: when someone asks this question, does the engine mention you at all? Share of voice asks: of all the brands it mentions, how much of that space is yours? You can have high visibility and low share of voice (you are always named, but so are five competitors) or low visibility and high share of voice (you are rarely discussed, but when you are, you own the answer). Pharma teams need both, split by audience.

How do you measure whether an AI is accurate about a drug?

You compare what the engine says against your approved, referenced claims — the label, the pivotal-trial data, the guideline positioning — and flag any statement that contradicts, overstates, or omits something material. Accuracy is the metric generic AI-visibility tools skip, and it is the one that matters most in a regulated category: an engine can name your brand in every answer and still be telling patients something the label does not support.

Measure all five, split them by audience, and read them as a set. In pharma the point of the exercise is not a vanity score — it is knowing, precisely and continuously, what the engines are telling your patients and prescribers, where that is wrong, and where a competitor is quietly taking the answer. That is a number you can act on.