A report is a numerator with no denominator
The FDA Adverse Event Reporting System and the MHRA Yellow Card scheme collect suspicions. Someone believed a medicine might have caused something and said so. That is the entire evidentiary content of a single report: not a verified diagnosis, not an adjudicated case, and not a finding that the drug was involved. Regulators say so plainly on their own public dashboards, because the alternative reading is so tempting.
The arithmetic is the core of the problem. To state how often something happens you need cases in the numerator and exposure in the denominator, and a spontaneous system supplies neither cleanly. The numerator is a filtered, self-selected fraction of events someone chose to write down. The denominator does not exist inside the database at all, because nothing about how many people were dispensed the drug, or for how long, is captured by the act of filing a report.
The consequence is immediate for the comparison readers most want to make. Ranking two drugs by raw report count tells you which was taken by more people, discussed more loudly and marketed more recently. Semaglutide and tirzepatide accumulated reports rapidly through the early 2020s largely because dispensing volumes for incretin drugs rose steeply over the same period. A curve that rises with prescriptions is not evidence that a drug became more dangerous.
Who files a report, and what actually gets stored
Two streams feed these databases. The first is mandatory: an authorisation holder that learns of a serious and unexpected suspected reaction must submit an expedited report within a short statutory window, with everything else flowing in periodically. The second is voluntary, covering clinicians, pharmacists and, since the mid-2000s in the United Kingdom, patients reporting directly. The mandatory stream dominates the volume, which is worth holding onto: the size of a drug's record depends heavily on whether a company exists that is obliged to keep it.
What gets stored is a coded version of the clinical narrative. Free text is mapped to standard terms from the MedDRA dictionary, and the mapping is lossy in both directions. One clinical picture can be split across several preferred terms, fragmenting a signal into pieces that individually look unremarkable, while a broad term can pool genuinely different problems into one impressive count. Anyone querying a public dashboard is querying coding decisions as much as patients. The same episode can also arrive three times over, from the patient, the prescriber and the manufacturer, and deduplication is imperfect.
Under-reporting is the normal state, and it is uneven
The best single estimate of how much is missed comes from a systematic review of thirty-seven studies across twelve countries, which found a median under-reporting rate of around ninety-four percent. The figure is old enough and variable enough to treat as an order of magnitude rather than a constant, but the direction is not in dispute: most adverse events that occur never enter any national system.
Uniform under-reporting would be survivable, since relative comparisons would still hold. What happens instead is differential under-reporting. Serious events are reported more often than mild ones, and unexpected events more often than labelled ones. A clinician seeing nausea in a patient on a GLP-1 receptor agonist is seeing what the prescribing information predicts and has little reason to file; the same clinician seeing an unfamiliar symptom may well file. The database is enriched for the surprising, which is useful for detection and disastrous for estimating frequency.
Time adds another distortion. The classical description, from work on non-steroidal anti-inflammatory drugs in the 1980s, is that reporting climbs after launch and falls away as prescribers grow familiar with a product. Later attempts to reproduce that pattern have been inconsistent, so hold it as a caution rather than a law. The narrower and defensible version still matters: a drug's position in its own publicity cycle affects its report count independently of its safety.
Stimulated reporting: attention manufactures data
Notoriety bias is the pharmacovigilance term for what happens when a suspected reaction becomes news. A regulator opens a review, a broadcast segment airs, a law firm buys advertising naming a drug and a condition, and reports for that pairing step upward within weeks. The step is often visible as a near-vertical line in the monthly counts, aligned to a date in the news rather than to anything about exposure. Once a pairing has been publicised, every count for it afterwards is partly a measurement of the publicity.
The clearest modern demonstration came through the Yellow Card scheme during the COVID-19 vaccination programme, when a campaign actively encouraging direct patient reporting produced volumes far above the baseline for ordinary medicines. That was a joint product of unprecedented exposure and unprecedented awareness, not a comparative risk ranking against products nobody was urged to report on.
Peptide-adjacent signals have their own version. Online communities discuss a symptom, the discussion propagates, and a wave of reports follows it. The gastrointestinal motility reports associated with GLP-1 receptor agonists illustrate the difficulty: delayed gastric emptying is a real and expected pharmacological effect, severe cases are real, and the reporting surge nonetheless tracked media coverage and perioperative guidance closely enough that the curve cannot be read as an incidence curve. Attention works in reverse too: nobody campaigns for reports on compounds that were never approved.
How regulators turn a pile of reports into a signal
Signal detection asks a comparative question inside the database: is this drug reported with this event more often than expected, given how often each appears overall? The simplest implementations are the proportional reporting ratio and the reporting odds ratio. A widely used screening rule from the early 2000s requires a proportional reporting ratio of at least two, a chi-square of at least four, and at least three cases before a pairing warrants a human look.
Those thresholds exist because small numbers behave badly: two reports out of two produce a spectacular ratio and mean nothing. Bayesian shrinkage methods were built for this, including the empirical Bayes approach the FDA adopted for routine screening and the Bayesian confidence propagation neural network used on the global VigiBase collection. Both pull sparse estimates toward the null in proportion to how little data supports them, giving more stable rankings rather than more truthful ones.
The critical limitation is the comparator. Expected counts come from the rest of the database, so every result is relative to reporting behaviour rather than to the population. This produces masking, sometimes called competition bias: a heavily reported product dominating a class inflates the expected count for events in that class and can conceal a genuine excess in a smaller product beside it. In an area as concentrated as the incretin drugs, that is not theoretical.
Whatever emerges is a signal in the formal sense: information suggesting a possible causal association judged to warrant further evaluation. It is a hypothesis with a case count attached, and a screen calibrated to miss almost nothing will raise many that fail. Most do fail. That is the system working, not malfunctioning.
Three GLP-1 signals and what became of each
The pancreatitis signal is the cleanest worked example. A 2011 analysis of the FDA reporting database found disproportionate reporting of pancreatitis and of pancreatic and thyroid cancer for incretin-based therapies, and it received wide coverage. The FDA and the EMA then reviewed the full evidence base, including trial data, preclinical toxicology and epidemiological studies, and set out a joint assessment concluding that the available data did not support a causal association with pancreatic disease. The signal was correctly generated, properly evaluated and not confirmed.
The suicidal ideation signal followed the same arc a decade later. A small number of spontaneous reports from one national authority in 2023 prompted the European regulator's safety committee to open a formal review of GLP-1 receptor agonists. The FDA said its preliminary evaluation found no evidence of a causal relationship, and the European review concluded the evidence did not support one. Confounding by indication is severe here, since the treated population carries an elevated background rate of depression.
The third example runs the other way. Non-arteritic anterior ischaemic optic neuropathy is rare enough that spontaneous reports alone would have stayed ambiguous indefinitely. What moved the question was a retrospective cohort study published in 2024 in a specialty eye-centre population, reporting a substantially higher hazard among patients prescribed semaglutide than among matched patients on other agents. European regulators later treated it as a very rare adverse reaction and moved to reflect it in product information. The resolution came from a design with a denominator.
A fourth case never touched post-marketing reports at all. The boxed warning for thyroid C-cell tumours carried by several long-acting GLP-1 receptor agonists, including semaglutide and tirzepatide, originates in rodent carcinogenicity studies done before approval. Signals arrive from preclinical work, trials, registries and spontaneous reports, and the spontaneous stream is only one of the four.
Why grey-market peptides barely register at all
BPC-157 has no approved human indication in any major jurisdiction. Reviewing which bulk substances may be used in compounding, the FDA placed it in the category raising significant safety concerns, citing insufficient information to evaluate its risks. That status is the whole story for reporting purposes: with no approved product there is no authorisation holder, no expedited reporting obligation, and in most transactions no prescriber, pharmacy record or product identifier.
For an event to reach a national database, someone must present for care, disclose the compound, and have a clinician who recognises it and files. Every step leaks. Material bought as research-labelled powder is often not mentioned to a treating clinician, and a clinician who has never heard the name has little reason to attribute anything to it. Independent testing has also repeatedly found grey-market vial contents that do not match the label, so even a filed report may not identify the real exposure.
Compounded products sit awkwardly in between. A compounding pharmacy carries different reporting obligations than a new drug application holder, and the FDA has described receiving adverse event reports involving compounded semaglutide, including events arising from patients measuring their own doses from vials rather than from any property of the molecule. The database captures a product-and-practice combination, not the drug in isolation.
The practical rule follows. For an approved drug with high dispensing volume, an empty cell in the reporting database is weak evidence that an event is rare. For a compound sold outside the licensed supply chain, it is no evidence at all. Absence of reports is a property of the reporting pipeline before it is a property of the molecule, and for these compounds that pipeline is essentially closed.
What an individual report genuinely can establish
Spontaneous reporting is very good at one job nothing else does as cheaply: detecting rare, serious and clinically distinctive events that essentially do not occur at background in the treated population. The Yellow Card scheme exists because of exactly such an event, established in 1964 after the thalidomide disaster. Aplastic anaemia, anaphylaxis, angioedema, severe cutaneous reactions and characteristic congenital malformations have near-zero background rates, so a handful of well-documented cases can carry real weight.
At the level of a single case, certain features raise evidential value considerably. A positive dechallenge, where the event resolves after the drug is stopped, is informative; a positive rechallenge, where it recurs on restarting, is stronger. So are a plausible time to onset, the absence of a competing explanation, and a known mechanism that predicts the effect. Structured instruments such as the Naranjo algorithm and the WHO causality categories exist to make that judgement consistent across reviewers, not automatic.
The mirror weakness follows directly. Spontaneous reports are close to useless for common events in populations that already have them. Gallbladder disease, cardiovascular events, mood disorders and pancreatitis all occur at meaningful rates in people with obesity and type 2 diabetes whether or not they take anything, so the background swamps whatever the drug contributes.
That is the reasoning behind active surveillance. The FDA's Sentinel programme, built after 2007 legislation, runs pre-specified queries across large claims and health record populations where the denominator is known and an unexposed comparator exists. Regulators increasingly use spontaneous reports to raise questions and active surveillance, registries or trial safety databases to answer them.
Reading a safety dashboard without fooling yourself
A short set of habits prevents most common errors. Never compare counts between drugs without knowing something about relative exposure. Never read a rise over time as a rise in risk until you have checked what dispensing volumes did over the same window. When you see an inflection, look for a date: a regulatory announcement, a covered publication or a litigation campaign in the same month is the likeliest explanation.
Where narrative detail is available, read it rather than the coded term. A report documenting concomitant medicines, the indication, a plausible latency and a dechallenge is worth more than fifty consisting of a drug name and a symptom word. Counting treats those as equal; they are not.
Then ask whether anyone has studied the question properly. If a cohort study, registry analysis or pooled trial safety database addresses the same drug and event, that evidence supersedes the report count outright rather than sitting alongside it. The count did its job when it prompted the study.
Finally, allow a signal to stay unresolved. Much of what is flagged reaches neither confirmation nor refutation, because nobody funded the study that would settle it. The honest description is neither a confirmed harm nor a debunked scare: an open question with a known number of reports attached.