HN Debrief

Clinical failure rates over the decades: yikes

  • Biotech
  • Healthcare
  • AI
  • Regulation
  • Economics

The post from Derek Lowe highlights a brutal fact about pharma economics: once a drug reaches human testing, roughly nine out of ten candidates still fail, and that ugly number has not improved much over decades. Lowe’s point is not that companies are careless. It is that even after years of chemistry, animal work, cell assays, and target selection, human biology keeps invalidating the bet. A candidate can look convincing before the clinic and still die on efficacy, safety, side effects, or trial endpoints. Because patents start ticking early and each late-stage failure burns huge amounts of time and capital, the few winners have to carry the whole system.

If you work around biotech, do not read a flat 90% clinical failure rate as simple industry incompetence or as proof AI will soon fix drug discovery. The practical edge is in better target validation, biomarker-driven patient selection, and tighter go or no-go decisions before expensive late-stage trials.

Discussion mood

Mostly sober and pragmatic. People did not treat the 90% failure rate as shocking so much as a reminder that human biology is still poorly modeled, late-stage trials are expensive truth tests, and easy targets have largely been harvested. The sharpest disagreements were about whether the steady rate is economically rational, driven by incentives and regulation, or simply the unavoidable cost of doing frontier medicine.

Key insights

  1. 01

    Clinical trials test finished molecules in humans

    By phase 1 through phase 3, the candidate is already the actual product candidate, not a rough draft that can be cheaply tweaked after a bad result. That makes phase 3 closer to a launch that reveals real-world rejection than to an early prototype cycle. When a late trial fails, the company often cannot patch the same molecule and rerun. It usually has to go back and find a different one entirely.

    Treat late-stage clinical risk as binary portfolio risk, not iterative product risk. If you are judging biotech timelines or budgets, assume phase 3 failure can wipe out years of work rather than trigger a quick revision.

      Attribution:
    • pama #1
    • fabian2k #1
    • ip26 #1
  2. 02

    The biggest waste is bad target validation

    Several comments sharpened the bottleneck from "drug development is hard" to "too many preclinical claims do not survive contact with humans." The expensive leak is not just chemistry execution. It is weak confidence that the biological target actually matters in patients. Better use of human biomarkers, genetics, and responder stratification was presented as the real route to improvement, because filtering bad targets earlier saves far more than optimizing compounds built on the wrong premise.

    Push diligence upstream. For any drug platform or tooling company, ask first how it improves human target validation or patient stratification, not just how it generates molecules faster.

      Attribution:
    • braiamp #1 #2
    • estearum #1
  3. 03

    A clinical failure can still reflect a good drug

    In crowded therapeutic areas, the trial only counts as a success if the drug beats the current standard of care enough to justify approval and adoption. That means a compound can be safer, easier to dose, better tolerated, or marginally more effective and still be recorded as a failure if it misses the required endpoint or competitive threshold. Comments pointed to Keytruda, anti-vascular endothelial growth factor therapies, and oral glucagon-like peptide-1 programs as markets where tiny clinical advantages can be worth billions.

    Do not read approval statistics as a clean measure of scientific usefulness. In mature markets, commercial value often comes from convenience, dosing interval, route of administration, or side-effect profile as much as from headline efficacy.

      Attribution:
    • s1artibartfast #1 #2 #3
  4. 04

    The hurdle has risen as safety knowledge improved

    Part of the flat success rate comes from a stricter modern filter, not just from bad science. One concrete example was hERG inhibition, an off-target effect linked to QT prolongation and arrhythmias that became a routine screen only after the 1990s. Older eras advanced molecules without that level of scrutiny. Today many candidates are killed for liabilities that once would have been discovered later or tolerated. The denominator got harder.

    Be careful comparing modern clinical productivity to older decades without adjusting for a tougher safety bar. A flat approval rate can hide real progress if today’s candidates are being asked to clear risks yesterday’s drugs never faced.

      Attribution:
    • refurb #1 #2
  5. 05

    AI hype collides with the human-testing bottleneck

    People who work near the field drew a clear line between useful machine learning and the stronger claim that AI will collapse drug timelines. Computational methods already help with screening and design, but commenters were blunt that they do not solve the core failure mode. A molecule that looks promising in silico still has to survive human biology. The cited AI success stories were treated as interesting but not as evidence that the clinical attrition problem has been cracked.

    Separate workflow automation from true prediction of human efficacy and safety. If an AI drug company cannot show that its models reduce clinical attrition, faster lead generation alone is not a breakthrough.

      Attribution:
    • Fomite #1
    • tchalla #1
    • tim333 #1

Against the grain

  1. 01

    Internal incentives may keep weak programs alive

    A minority view argued that the 90% number is not just about biology. It may also reflect organizational dynamics inside large pharma. Advancing a program can help careers, while killing it early rarely does. Even if no single project manager makes the call, the incentive gradient can still bias groups toward continued spending until the clinic delivers an unambiguous no.

    When evaluating a pharma partner or acquisition target, inspect governance as closely as the science. Strong kill criteria and independent review can be as valuable as better discovery tools.

      Attribution:
    • mbnielsen #1
  2. 02

    Protocol preregistration may have exposed weaker evidence

    One commenter tied the apparent drought in true breakthroughs to the 1997 shift toward registering clinical trial protocols in advance. The claim is that stricter preregistration reduced p-hacking and made it harder to rescue weak studies with post hoc endpoints, so more programs now fail honestly instead of passing on statistical gamesmanship. That does not prove innovation slowed, but it does suggest some of the old success rate may have been flattered by looser trial practice.

    Do not compare present-day approval yield with older literature as if the evidentiary rules were constant. Methodological tightening can make productivity look worse while making the output more trustworthy.

      Attribution:
    • bibouthegreat #1
  3. 03

    The stability itself may reflect regulation

    A few readers found the long-run consistency of the failure rate suspiciously neat and suggested it could partly encode regulatory policy rather than raw scientific difficulty. The point was not that biology is easy. It was that a stable number over decades can emerge from approval thresholds and trial design conventions as much as from discovery quality.

    If you want to improve success rates, look beyond lab science to endpoint design and review standards. Some of the bottleneck may be institutional rather than purely biological.

      Attribution:
    • contubernio #1

In plain english

endpoint
The specific outcome a clinical trial is designed to measure, such as survival, symptom reduction, or tumor shrinkage.
glucagon-like peptide-1
A hormone involved in blood sugar control and appetite that is targeted by a major class of diabetes and obesity drugs.
hERG inhibition
Blocking a heart-related ion channel called hERG, which can disrupt the heart’s electrical rhythm and create safety risks.
in silico
Tested or simulated on a computer rather than in a lab or in living organisms.
machine learning
A set of statistical and computational methods that learn patterns from data to make predictions or decisions.
off-target effect
An unintended action of a drug on a biological system other than the one it was designed to affect.
p-hacking
Manipulating statistical analysis or trying many outcomes until a study appears positive by chance.
Phase 1
The first stage of human clinical testing, usually focused mainly on safety, dosage, and side effects.
Phase 3
A large late-stage clinical trial that compares a drug against standard treatment or placebo to support regulatory approval.
preclinical
Research done before human trials, usually using cells, tissues, animals, and lab experiments.
preregistration
Publicly recording a study’s planned methods and endpoints before it starts to reduce bias and post hoc changes.
QT prolongation
A change in the heart’s electrical timing seen on an electrocardiogram that can raise the risk of dangerous arrhythmias.
randomized controlled trial
A study where participants are randomly assigned to treatment groups so researchers can compare outcomes fairly.
vascular endothelial growth factor
A signaling protein that promotes blood vessel growth and is a major drug target in eye disease and cancer.

Reference links

Drug discovery and clinical attrition references

AI drug discovery claims

Related reading and media

Adjacent examples and critiques