The short answer: a rodent result is a hypothesis, not a preview
A study showing that a peptide made a rat's cut Achilles tendon stronger at fourteen days tells you that the compound does something measurable in that model. It does not tell you that a person with six months of Achilles pain will get better. Those are different tissues, different injuries, different repair biology and different timescales, and the gap between them is where most drug candidates die.
This is not a peptide-specific complaint. It is the general failure rate of preclinical research, and it is well quantified. When researchers have gone back and tracked highly cited animal studies forward, only a minority ever reached a human randomised trial at all, and a much smaller fraction produced something a regulator approved. Whole therapeutic areas have been built on rodent results that evaporated in humans.
What follows is the specific list of mismatches for healing research, because it matters which ones apply. Some are fixable with better models. Some are structural facts about rodents. A few are not biology at all, but the way small unblinded experiments get reported and published.
How often preclinical results survive contact with humans
The most-cited attempt to measure this examined a set of highly cited animal studies and followed each one forward through the literature. Roughly a third were later tested in a human randomised trial. Around one in ten ended in a treatment that entered practice. The remainder either were never tested in people or were tested and failed.
A separate systematic review compared animal and human results head to head for a handful of interventions where both existed. The animal data agreed with the eventual human result in only about half the comparisons. That is close to a coin flip, and it was measured on interventions that had already been considered promising enough to fund a human trial.
The best-known worked example is the free-radical trapping agent NXY-059 for acute stroke. It had a large, consistent, apparently convincing animal literature. The first phase 3 trial, SAINT-I, appeared positive on its primary endpoint; the larger confirmatory SAINT-II was flatly null, and the programme ended. Later analysis of the preclinical package found the usual weaknesses: small groups, little blinding, and a tendency for the largest effects to come from the least rigorous experiments.
Rodent skin closes by contraction, and human skin mostly does not
This is the single largest artefact in the wound-healing literature and the easiest to miss. Mice and rats have a panniculus carnosus, a thin sheet of striated muscle in the subcutaneous layer. When you make a hole in rodent skin, that muscle sheet pulls the wound edges together. A large share of what gets recorded as wound closure is the skin being dragged shut, not new tissue being built.
Humans retain only vestiges of that layer, mostly in the neck and palm. Human skin is also tethered to underlying fascia rather than sliding freely. A human excisional wound therefore closes mainly by filling with granulation tissue and re-epithelialising across it. The rate-limiting steps are different, so an agent that speeds rodent closure may be doing nothing to the process that limits a human.
The field knows this, and the standard fix is the splinted excisional model, in which a silicone ring is sutured around the wound to hold the edges apart and force closure to happen by granulation and re-epithelialisation. It is a genuine improvement. But splinting is not universal, and a great deal of the peptide wound literature reports unsplinted wounds with photographs and a percentage-closure curve, which is the measurement most inflated by contraction.
Surgical transection is not tendinopathy
Nearly all rodent tendon studies create the injury with a blade. The Achilles is transected, sometimes repaired, sometimes left to heal across a gap, and healing is then measured. That is a clean acute wound in previously normal tissue, with a strong inflammatory response and an obvious repair signal.
The condition people actually want treated is usually nothing like that. Chronic tendinopathy is a degenerative process: disorganised and immature collagen, increased ground substance, neovascularisation and nerve ingrowth, and strikingly little classical inflammation. It develops over months of repetitive loading rather than in an instant. A compound that improves organisation of a healing surgical scar has not been shown to reverse a degenerative matrix, because that experiment was not run.
Two model families try to bridge this. Collagenase injection dissolves matrix chemically and produces something that looks degenerative on histology, but it is an acute chemical insult that resolves on its own, not an overuse process. Treadmill overuse models are closer to the real mechanism but are slow, variable and much less commonly used, so they appear far less often in the peptide literature.
There is also a loading problem. Human tendon rehabilitation is itself an active treatment; progressive loading has better evidence than most drugs. A caged rodent cannot be assigned a rehabilitation protocol, so animal studies measure the compound against unstructured cage activity, while any human result would have to beat structured loading.
The animals are young, male, uniform and disease-free
The typical rat in a tendon or wound study is a two- to three-month-old male of an outbred stock, housed in specific-pathogen-free conditions, fed a fixed diet, and still growing. Rats do not fully close their growth plates, so these animals are in an anabolic state throughout the experiment. That alone raises baseline repair capacity above anything a middle-aged human brings to the table.
The humans interested in healing compounds are frequently the opposite: in their forties to sixties, often with insulin resistance, vascular disease, a smoking history, or on medications that blunt repair. Each of those is a known determinant of healing rate. Aged and diabetic rodent models exist and heal noticeably worse, but they are more expensive and slower, so they show up in a small minority of published work.
Sex is a further narrowing. Preclinical pharmacology has a long-documented male bias, which is why funders introduced explicit sex-as-a-variable requirements. Where a study uses only young males, it has not established that the effect exists in females, in the aged, or in anyone with the comorbidities that make healing a clinical problem in the first place.
Weeks in a rat are not weeks in a person
Rodent healing studies usually read out at seven, fourteen or twenty-one days, because that is where rodent repair happens and because it fits a grant cycle. Human tendon and ligament remodelling runs for many months, and scar maturation in skin continues for a year or more. A three-week endpoint captures the proliferative phase and almost nothing of remodelling.
That matters because the two phases can move in opposite directions. Something that accelerates early matrix deposition can produce a bulkier, more disorganised scar that looks better at day fourteen and worse at month six. Fibrosis and rapid healing use overlapping machinery, and a short readout cannot distinguish them.
Scale compounds the problem. A rat Achilles is a couple of millimetres across; a human Achilles is a rope. Diffusion distances, vascular ingrowth, mechanical loads and the sheer volume of matrix that must be laid down differ by orders of magnitude. Processes that cross a two-millimetre defect in days do not simply take proportionally longer across a human tendon; some of them do not scale at all.
Dose does not scale linearly either. A milligram-per-kilogram figure in a rat does not convert to a human figure by multiplying body weight, because metabolic rate, clearance and body surface area all scale non-linearly. This is why reading a rodent dose as a human instruction is not conservative or aggressive, it is simply a category error.
What the readouts actually measure
Rodent tendon studies most often report a biomechanical number, usually maximum load to failure. It is objective, which is a genuine strength. But load to failure scales with cross-sectional area, so a bigger, more scar-like repair can pull higher numbers without being better tissue. Stress and stiffness normalised to cross-section tell you more, and are reported less often.
Histology scores are the other standard readout, and they are ordinal judgements made by a person looking down a microscope. Their value depends entirely on whether that person was blinded to group allocation. Where blinding is not stated, the score is an expectation as much as a measurement.
Wound studies lean on planimetry from photographs, which inherits every problem in the contraction section above, and on immunostaining for collagen types, growth factors or vessel density. Staining intensity is a legitimate mechanistic clue and a weak outcome; more vessels and more type I collagen at day ten is not the same claim as stronger, better-organised tissue at month six.
How the reporting itself inflates the signal
Before any species question, the animal literature has a measurement-quality problem. Audits of published in vivo studies have repeatedly found that a minority report randomisation, a minority report blinded outcome assessment, and very few report a sample size calculation. The ARRIVE guidelines were written in 2010 and revised in 2020 specifically to fix this, and adherence has improved slowly.
Group sizes are small, usually six to twelve animals. Small studies do not merely have wide confidence intervals; when combined with selective publication they systematically overestimate effects, because only the runs that reached significance get written up. In the animal stroke literature, where this was measured directly, publication bias was estimated to overstate efficacy by roughly a third.
Reproducibility audits in adjacent fields found the same thing from the other direction. When one industry group attempted to reproduce a set of landmark preclinical oncology findings, only a handful of the original claims held up. Nothing about tissue repair makes it immune to this.
What this means for BPC-157, TB-500 and GHK-Cu
BPC-157 is the clearest case. Its reputation rests almost entirely on rodent work, much of it Achilles or ligament transection, gastric lesion and colonic anastomosis models, reported over two decades and concentrated in a small set of affiliated groups. There is no published randomised placebo-controlled human efficacy trial. It is not FDA-approved, and in 2023 it was placed in the FDA category of bulk substances presenting significant safety risks, removing it from lawful compounding in the United States.
One feature of that literature deserves flagging on its own terms: effects are reported across a very wide range of doses and multiple routes, including intraperitoneal injection and administration in drinking water, without much dose-response structure. A flat response over orders of magnitude is unusual for a receptor-mediated drug effect and is the pattern you would also expect from unblinded scoring of noisy models.
TB-500 is a fragment marketed as a stand-in for thymosin beta-4, and thymosin beta-4 does have real human trial history, but not for the use people assume. The randomised human work was of topical formulations in eye and skin conditions run by RegeneRx, not injections for tendon or muscle repair, and that programme did not deliver a clean approval. The injectable musculoskeletal case remains preclinical.
GHK-Cu is the one with the most human data and the narrowest claims. It has a genuine in-vitro record in dermal fibroblast assays and small human cosmetic studies with skin endpoints such as wrinkle depth and appearance, typically as a topical cosmetic rather than an approved drug. That is a real body of evidence for a cosmetic claim, and it is not evidence for injected tissue repair, which is how it is often discussed.
What would actually change the picture
For any of these compounds, the evidence that would matter is not another rodent model. It is a registered, randomised, placebo-controlled human trial in a defined condition, with a validated patient-reported or functional endpoint, imaging or biomechanical confirmation, and follow-up long enough to catch remodelling rather than just early swelling. That trial has not been run for injectable BPC-157 or TB-500 in any indication.
Short of that, preclinical work can be made much more predictive: aged and comorbid animals, overuse rather than transection, splinted wounds, blinded outcome assessment, prespecified sample sizes, and independent replication in a second laboratory. Multi-centre preclinical designs, borrowed from clinical trial methodology, exist and consistently produce smaller effect sizes than single-laboratory work, which is itself informative.
The regulatory environment is moving in a related direction. The FDA Modernization Act of 2022 removed the blanket statutory requirement for animal testing before human trials, opening the door to organoids, tissue chips and computational models. Those are not obviously better predictors yet, but the change is an official acknowledgement that the animal step is a weaker filter than it was assumed to be.
Until any of that exists for these peptides, the honest description is unchanged: interesting animal data, an unvalidated model, and no human efficacy evidence.