The short answer: eight questions, in this order
Before you weigh what a peptide study claims, ask eight things about it. What species was it done in, and by what route? What was the comparison group? Was allocation randomised, and was the person scoring the outcome blinded? What was actually measured, and would a person notice it? How many subjects, and how wide is the interval around the estimate? Was the protocol registered before the data existed? Who paid for it, and who else was in the room? And has anyone unaffiliated reproduced it?
The order is not decorative. Most of what circulates about research peptides is disqualified by the first two questions, and there is no point interrogating a confidence interval in a study that was never going to say anything about humans. Work down the list and stop at the first hard no. In practice that takes about ten minutes per claim, most of it spent in a methods section rather than an abstract.
Below, each question is applied to three compounds sitting at very different points on the evidence scale: semaglutide, which has been through registrational trials that counted clinical events; thymosin beta-4, which has genuine randomised human data in a tissue almost nobody buying it cares about; and BPC-157, which has an enormous following and no controlled human efficacy data at all.
Question one: what species, and by what route?
Species is the first filter because it decides whether you are reading a finding or a hypothesis. A result in a rat is a statement about rats. It earns a compound the right to be tested in people; it does not tell you what happens when a person takes it. That is not pedantry about scientific manners, it is an observation about how often the translation fails.
The best-known audit of this followed highly cited animal studies forward: roughly a third were subsequently tested in a randomised human trial at all, and only a small minority ended in an approved treatment. A systematic review comparing animal experiments with the human trials of the same interventions found agreement to be the exception. The base rate for worked in mice, therefore works in people is poor, and worst in exactly the areas where research peptides cluster: inflammation, tissue repair and neuroprotection.
Applied, the answers are immediate. Essentially the whole BPC-157 literature is rodent, much of it delivered intraperitoneally or in drinking water at exposures nobody has characterised in a human. Thymosin beta-4 does have randomised human data, but it is concentrated in topical ophthalmology and uses the intact 43-residue protein, not the injected fragment sold as TB-500 for tendons. Semaglutide's weight and cardiovascular claims rest on humans, in the tens of thousands, at the route the label describes.
Question two: what was the comparison group?
A control arm is what converts a change into an effect. Without one you are measuring the passage of time, the natural history of the condition, the rehabilitation running alongside, regression toward the mean, and expectation. Soft-tissue injuries in particular improve on their own, and people start something new when symptoms are at their worst, which loads every uncontrolled observation the same way.
The number that makes semaglutide's weight result interpretable is not the roughly 14.9 percent mean reduction reported in the 68-week phase 3 obesity trial. It is the roughly 2.4 percent the placebo arm lost over the same period, with the same lifestyle counselling and the same clinic visits. The gap between those two figures is the drug's contribution; the placebo figure is everything else the trial did to people. Quote the first number alone and you have a marketing claim rather than a result.
In preclinical repair work the question is subtler. Most animal tendon and gut models do include untreated or vehicle groups, so they clear this bar on paper. What to check instead is whether the comparison is fair: same injury severity, same timing, same handling, same operator. And when a claim rests on before-and-after testimonials, there is no comparator at all, which is why testimonial volume never accumulates into evidence no matter how large it grows.
Question three: randomised, concealed, and who scored the outcome?
Randomisation exists to make the groups comparable in the ways nobody thought to measure. Allocation concealment stops whoever enrols subjects from steering the healthier ones into the treatment arm. Blinding stops expectation from leaking into the measurement. These are three separate safeguards, and studies routinely have one without the others.
Blinding matters most where the endpoint is a judgement: a histology score, a pain scale, a wound photograph graded by eye. That is precisely the territory repair-peptide research occupies. Methodological surveys of the animal literature have repeatedly found that randomisation, blinded outcome assessment and any sample-size justification go unreported in the majority of published studies, and that studies which do not report blinding tend to report larger effects than those that do. The ARRIVE reporting guidelines exist because of this, and adherence to them remains patchy.
So when an animal paper says a peptide improved tendon healing, look for who read the slides and whether they knew which group each came from. If the paper does not say, the honest reading is that it probably was not blinded. In the human literature the equivalent check is the participant flow diagram CONSORT requires: how many were randomised, how many analysed, and whether the analysis was intention-to-treat or quietly restricted to completers.
Question four: what was measured, and would a person notice it?
There is a ladder of endpoints. At the bottom is something happening in a dish: cells migrating across a scratch, collagen expression rising in cultured fibroblasts. Above that, a biomarker or an image in a living animal. Above that, a validated symptom or function scale in humans. At the top, counted clinical events, meaning heart attacks, strokes, deaths, fractures and hospitalisations.
Semaglutide's cardiovascular indication is instructive because it sits at the top of that ladder. The outcomes trial in people with established cardiovascular disease and overweight or obesity but without diabetes enrolled 17,604 participants and counted major adverse cardiovascular events over roughly three years, reporting a hazard ratio of about 0.80 with a 95 percent confidence interval of roughly 0.72 to 0.90. That is not a biomarker nudged in the right direction, it is fewer events in people, and the regulator added the indication on that basis in 2024.
Most peptide evidence sits several rungs lower. Thymosin beta-4's mechanistic case rests substantially on actin sequestration and cell migration assays, and its human trials measured corneal staining and dry-eye symptom scores. BPC-157's case rests on histology and load-to-failure testing in rodent tendons. None of those measurements is worthless, and none of them is a person returning to sport. The distance between the two is where most overclaiming happens.
Two related traps live here. Composite endpoints bundle a hard outcome with a softer one, so a result can be driven entirely by the soft component. And a trial designates one primary endpoint for a reason: if the press release leads with something that was not it, that is the finding to doubt.
Question five: how many subjects, and how wide is the interval?
Sample size determines what a study could have detected. The interval around the estimate tells you what the data actually rule out. A typical rodent repair experiment runs six to ten animals per group, which is enough to detect a large effect and nothing else. It also means that any effect it does detect is likely to be overstated, because in a small sample only the larger random fluctuations clear the significance threshold.
This is why a p-value just under the conventional threshold in a small study is weak information while a narrow confidence interval in a large one is strong information. The semaglutide cardiovascular interval excludes no effect and also excludes an implausibly large effect, so it pins the answer down. An interval running from 0.4 to 0.99 in a few hundred people is technically significant and tells you almost nothing about the size of the benefit.
Ask the reciprocal question too: what would this study have missed? A trial with eighty participants can be immaculately conducted and still incapable of detecting a difference that would matter. No significant difference in a small study is not evidence of no difference, and collapsing those two statements together is one of the commonest errors made in both directions.
Question six: was the protocol registered before the data existed?
Prospective registration is the closest thing clinical research has to a receipt. Since 2005 the major medical journals have required registration before enrolment as a condition of publication, and United States law has required registration and results reporting for applicable trials since the FDA Amendments Act of 2007. A registry entry timestamps the primary endpoint, the analysis plan and the intended sample size before anybody has seen the results.
The value of that timestamp is easiest to see when it is missing. Projects that audit published trials against their own registry entries have found a substantial rate of outcome switching: prespecified primary outcomes quietly dropped from the paper, and new outcomes reported as though they had been planned all along. Without a registry entry none of that is detectable, and an apparent finding could be the one comparison out of twenty that happened to come out.
This question is fast to answer. Search the public registry for the compound name. Semaglutide returns a dense programme with prospectively posted endpoints and posted results. BPC-157 returns almost nothing relevant: an old oral formulation programme in inflammatory bowel disease whose efficacy results were never published, and no registered trial for any tendon, ligament or muscle indication. That absence is itself a finding, and establishing it takes under a minute.
Question seven: who paid, and who else was in the room?
Funding is not disqualifying. Nearly every large trial of an approved drug is paid for by the company that makes it, because nobody else will spend what a registrational programme costs. The Cochrane methodology review on this found that industry-sponsored drug studies report favourable results and conclusions more often than non-industry ones, which is a reason to read the methods closely rather than to discard the data.
What matters is the structure around the money. Was the trial registered in advance? Did an independent data monitoring committee see the accumulating data? Was the statistical analysis plan posted? Did the results go to a regulator with authority to audit source records? The semaglutide programme is sponsor-funded and has all of those, and that combination is what makes sponsor funding survivable rather than fatal.
The subtler version of this question is source concentration. A literature can be free of commercial funding and still depend almost entirely on one group of investigators. The great majority of the BPC-157 rodent literature originates from a single long-running research programme and its collaborators. That is not misconduct and the work may well be sound, but it means the evidence base has one point of failure, and the ordinary reassurance of many independent hands arriving at the same place has never been obtained.
Question eight: has anyone unaffiliated reproduced it?
Replication most reliably separates a durable result from one that looked good once. The large organised effort to reproduce high-profile preclinical cancer biology experiments found that many could not be repeated as described, and that the effects which did replicate were markedly smaller than the originals. There is no reason to think peptide research is exempt from that pattern.
Replication has a recognisable shape. Semaglutide's weight effect has been reproduced across separate randomised trials in different populations, alongside a diabetes programme and a cardiovascular outcomes trial, and again in independent real-world cohorts. Thymosin beta-4's ophthalmic programme shows the other outcome: replicate trials disagreed with one another on their primary endpoints, which is what an unstable effect looks like, and no marketing approval followed anywhere.
For BPC-157 the question cannot yet be asked. There is no human result to replicate, and no published independent attempt at the animal work. That is the difference between a compound that has failed replication and one that has never been subjected to it. Both are correctly described as unproven, but only the first has actually produced information.
The checklist applied, end to end
Run all eight questions across the three compounds and the picture resolves quickly. Semaglutide: human, placebo-controlled, randomised and blinded, measured on both body weight and counted cardiovascular events, tens of thousands of participants with tight intervals, prospectively registered, sponsor-funded with independent monitoring and regulatory audit, and replicated repeatedly. That is what a fully answered checklist looks like, and it is rare.
Thymosin beta-4: real randomised human data exists and is respectable, but it is topical, ophthalmic, and uses the intact protein. For the tendon and muscle uses the injected fragment is bought for, questions one through eight return nothing. BPC-157: question one ends the enquiry for every human claim made about it, and the regulatory record reflects that, since it was placed among the bulk substances judged to raise significant safety risks and so is not lawfully available through compounding in the United States.
A full set of ticks is not a blanket endorsement, and this is where the checklist earns its keep in the other direction. The same semaglutide programme that produced the cardiovascular result also reported a significant increase in diabetic retinopathy complications in its earlier diabetes outcomes trial, and the extension of the obesity trial found that participants regained roughly two-thirds of the lost weight within about a year of stopping. Good evidence describes the harms and the limits as precisely as it describes the benefit.