Where AI gets its information about pharmaceutical brands — and why only 11% of it is yours
We counted every source cited by seven AI engines answering questions about ten pharmaceutical brands: 36,887 citations over three months. 89.2% pointed to pages the manufacturer does not own — and in Germany a brand's own content is cited less than half as often as in the US, in five of five brands measured in both markets.
When a patient asks ChatGPT whether a medicine is safe to take long-term, or a physician asks Perplexity how two biologics compare, the answer is assembled from somewhere. Not from nowhere, and not purely from the model's memory — these engines search, retrieve pages, and cite them. Those cited pages are the raw material of the answer your audiences read.
So we counted them. Every source cited by seven AI engines, across ten pharmaceutical brands, over three months. The result is the clearest answer we can give to a question brand teams keep asking and nobody seems to have measured: when AI talks about my medicine, whose words is it using?
The short version is that it is almost never yours.
What we measured
Ten pharmaceutical brands, one monitored workspace each, spanning oncology, immunology, neurology, respiratory and gastrointestinal medicine, across four markets. Seven AI engines, including both the developer APIs and the consumer applications people actually open. Every answer between 4 June and 25 August 2026 that carried at least one citation: 5,108 answers and 36,887 citations to 2,729 distinct domains.
Two method points matter before any number is quoted, because both change the answer.
We count references, not rows. A source cited at several points in one answer produces several records; counting those naively inflates citation volume by 19–42%. We collapse to one reference per distinct anchor within each answer and source. This is the less flattering choice — it makes every number here smaller — and it is the only one that matches what a brand team sees when they open the platform.
We ran it twice before believing it. This is a re-run of a study we published in June with roughly half the data. Restricted to the June window, the re-run reproduces that period's 23,509 citations exactly. The method was validated against a known result before any claim was made about what changed.
Finding 1: 89% of what AI cites about your brand is not yours
Across all ten brands, 89.2% of citations pointed to pages the manufacturer does not own. Brand and corporate sites together accounted for 10.8%.
What makes this worth trusting is that it did not move. In June, on the original five brands, owned share was 11.6%. Today, on those same five brands and 3,673 additional citations, it is 11.6% — identical to one decimal place. This is not a number drifting around a noisy mean. It looks like a structural property of how these engines answer questions about medicines.
Here is where the citations actually went:
| Cited entity | Share of all citations |
|---|---|
| PubMed / PMC / NIH | 10.0% |
| Brand & product sites (owned) | 7.3% |
| Corporate sites (owned) | 3.5% |
| drugs.com | 2.3% |
| ema.europa.eu | 2.0% |
| goodrx.com | 1.8% |
| accessdata.fda.gov | 1.7% |
| youtube.com | 1.6% |
| healthline.com | 1.4% |
| fda.gov | 1.2% |
The peer-reviewed literature leads, and by a wide margin. That is worth sitting with: the most-cited single entity when AI discusses a prescription medicine is the published evidence base, not the manufacturer's carefully governed product site, and not a consumer health portal either. Beneath the top ten the distribution has a long tail — 2,729 domains in total — but the head is dominated by literature, regulators and reference sites.
The share is lower again on the consumer applications. Split by surface, citations were 89.0% not-owned via the developer APIs and 92.0% not-owned on the consumer apps.
Finding 2: change the country, and your own content halves
This is the finding we did not expect, and it is the one with the most practical consequence for any team running a brand in more than one market.
Owned share in the United States is 13.7%. In Germany it is 5.7%.
| Market | Citations | Owned share |
|---|---|---|
| United States | 22,266 | 13.7% |
| United Kingdom | 2,443 | 8.9% |
| Sweden | 684 | 7.5% |
| Germany | 11,494 | 5.7% |
The obvious objection is that this is a composition artefact — that the US and German sets contain different brands, and the gap is really a brand effect wearing a country's clothes. It is not. Holding the brand constant, and with it the engines, the prompt set and the measurement window, the gap appears in five of the five brands we monitor in both markets, across three manufacturers:
| Brand (blinded) | US owned | DE owned | Gap |
|---|---|---|---|
| Brand A | 13.7% | 2.9% | +10.8 pts |
| Brand B | 14.0% | 3.9% | +10.1 pts |
| Brand C | 11.0% | 1.4% | +9.5 pts |
| Brand D | 14.0% | 10.5% | +3.5 pts |
| Brand E | 7.1% | 5.7% | +1.4 pts |
Same brand. Same engines. Same questions. Different country, and the brand's own voice drops out of the answer.
The mechanism, and it is not mysterious
In Germany the authoritative product information does not live on the manufacturer's website. It lives on third-party compendia. German citations point to Fachinfo, Gelbe Liste, the G-BA or the EMA in 13.3% of cases, against 0.8% in the United States. The mirror image holds for the American label hosts: DailyMed and the FDA account for 7.4% of US citations and 0.2% of German ones.
So the citation does not disappear in Germany. It relocates. The engine still needs a source for the prescribing information, still finds one, and still cites it — just not the brand.
The gate does not remove your content. It moves the credit.
The sharpest illustration in the data is a German brand site that redirects every request — including an identified AI crawler — to a healthcare-professional login wall. Across the entire corpus, that domain is cited zero times.
Meanwhile the Fachinformation it is gating is served openly by a third-party compendium, and that compendium is cited repeatedly in answers about the same brand.
A second German brand site in the set reaches the same place by a different route: its bare domain fails to negotiate a secure connection at all, and over plain HTTP it redirects away to a disease-awareness site. Different mechanism, identical outcome — zero citations.
This is worth stating plainly, because it inverts a common assumption. Gating professional content does not keep it out of AI answers. The information is already elsewhere — in the compendium, in the regulatory filing, in the literature. What the gate achieves is that when an AI engine explains your medicine, it does so in someone else's words, with someone else's emphasis, and attributes it to someone else.
We are deliberately not generalising this to every gated site. A third German brand site in our set returned a 403 to our crawler probe, and a 403 is not evidence of anything — it tells you a request was refused, not that a page is absent. That brand is also the strongest German performer in the set, at 10.5% owned. Whatever the rule is here, “gated equals invisible” is too crude, and we do not have the data to state the sharper version yet.
What a brand team can actually do with this
Stop treating owned content as the lever and start treating it as one input among many. If nine citations in ten point somewhere you do not control, a content strategy that consists entirely of publishing more on your own site is working on roughly a tenth of the surface area. The other nine tenths are compendia, regulators, literature and reference sites — and what those say about your medicine is measurable, correctable and, in several cases, formally contestable.
Check whether the sources AI is actually citing are correct about you. This is the part that distinguishes a regulated category from every other one. A wrong sentence on a high-authority compendium is repeated in every answer that draws on it, and it is nobody's job to notice. The remedy — a factual correction request to the publisher — is ordinary medical-affairs work, but it has to be triggered by a measurement that nobody is currently taking.
Do not assume a market strategy transfers. The US and German pictures here are different enough that a playbook built on one will underperform in the other for reasons that have nothing to do with execution quality. In Germany, influencing what AI says about your brand runs substantially through sources you do not own.
Reconsider what the login wall is buying you. There are real regulatory reasons for professional gating in some markets, and we are not arguing those away. But if the rationale has quietly become “this keeps our clinical content controlled,” the data says otherwise. The content is out. Only the attribution stayed home.
Limitations
Ten brands is ten brands. They are not a random sample of the pharmaceutical market — they are the brands we monitor, weighted toward markets where we run the most prompts, and the US and German sets carry far more citations than the UK and Swedish ones. Three of the five brands measured in both markets share a manufacturer, so the company-level independence of that comparison is weaker than the count of brands suggests. The within-brand comparison is the robust part of the market finding; the headline market averages inherit whatever composition bias the underlying set has.
The monitored prompt sets differ by brand, so cross-brand comparison of absolute owned share is not meaningful and we have not published one. What is comparable is the same brand against itself across markets, which is the comparison the second finding rests on.
Finally, this is a measurement of citations, not of influence. A cited source is one the engine retrieved and attributed.
The uncomfortable summary
For most of the past two decades, pharmaceutical digital strategy has rested on an assumption that was true: the brand website is the destination, and the job is to get people to it. In AI answers that assumption quietly stops holding. The answer is assembled before anyone clicks anything, from sources chosen by a retrieval system, and your site is about one line in ten.
That is not an argument for abandoning owned content. It is an argument for knowing what the other nine lines say — and in a category where a wrong sentence about a medicine is a safety matter rather than a marketing one, for checking whether they are right.
We rebuild this measurement continuously across the brands we monitor. If you want to see the equivalent picture for your own brand — which sources AI is citing, in which markets, and whether they describe your medicine correctly — that is what percivo.ai does. You can request a free GEO audit or request access.
Frequently asked questions
Where does AI get its information about pharmaceutical brands?
Mostly not from the brand. Across 36,887 citations made by seven AI engines answering questions about ten pharmaceutical brands, 89.2% pointed to pages the manufacturer does not own. The single largest cited entity was the biomedical literature — PubMed, PMC and NIH together accounted for 10.0% of all citations, more than any brand site, regulator or consumer health portal.
What share of AI citations point to a pharma brand's own website?
About 11%. In our measurement it was 10.8% across ten brands and 11.6% across the five brands we have measured since June — a figure that did not move across 3,673 additional citations. The share is lower still on the consumer apps most people actually use, where 92.0% of citations pointed somewhere other than the brand.
Why is a brand's own content cited less by AI in Germany than in the United States?
Because the authoritative product information sits somewhere else. In the US, 13.7% of AI citations about a brand point to the brand's own pages; in Germany, 5.7%. The mechanism is measurable: German citations go to third-party compendia and authorities — Fachinfo, Gelbe Liste, the G-BA, the EMA — 13.3% of the time, against 0.8% in the US. The citation that would land on a manufacturer's site in the US lands on a compendium in Germany.
Does gating a pharma website behind an HCP login stop AI from using the content?
No. It moves the citation to whoever serves the same content openly. One German brand site in our set redirects every request, including AI crawlers, to a healthcare-professional login. It was cited zero times across the entire corpus — while the prescribing information it gates is served publicly by a third-party compendium, which is cited. The content still reaches the answer; the brand simply is not the source of it.
