Every AI-in-drug-discovery headline eventually reaches for the same statistic: a new medicine takes ten to fifteen years and roughly $2-3 billion in fully-loaded cost to go from a research idea to an approved product, most of it spent on failures. It's a genuinely staggering number, and it's the number every AI lab implicitly promises to shrink. But "AI will compress drug discovery" is doing a lot of work in that sentence, and it's worth being precise about which part of a fifteen-year pipeline is actually shrinking -- because it is not all of it, and pretending otherwise sets up a credibility problem for the entire field.
Mapping the pipeline, stage by stage
Break the pipeline into its real stages and the picture gets much clearer.
Where AI is genuinely compressing timelines today
- Target identification and validation. Foundation models trained on multi-omics data -- genomics, transcriptomics, single-cell atlases -- can now surface and rank plausible disease targets from patterns that would have taken a research group years of hypothesis-driven investigation to uncover manually.
- Hit discovery. Generative chemistry models can propose and virtually screen candidate molecules against a target at a scale -- millions of candidates evaluated computationally -- that simply has no wet-lab equivalent. What used to require synthesizing and testing thousands of physical compounds can now be substantially pre-filtered in silico.
- Structure-based design. AlphaFold-class structure prediction and de novo design tools like RFdiffusion mean a team no longer has to wait months for an X-ray crystal structure of a target protein before starting rational design work.
- Lead optimization and early ADME/Tox triage. Predictive models for solubility, permeability, metabolic stability, and off-target liabilities let teams deprioritize weak candidates before committing wet-lab time and animal studies to them -- this is precisely the kind of iteration-cycle compression that connected preclinical evidence platforms are built to accelerate.
Where the clock does not move faster, no matter how good the model is
- Safety follow-up windows. If a regulator requires six months of observation to detect a delayed adverse effect, that is six months of biological and calendar time. No model shortens human physiology.
- Statistical power for efficacy. Detecting a real treatment effect -- and ruling out that an apparent benefit was chance -- requires a minimum number of patients observed for a minimum duration. This is a property of statistics and biological variability, not of computation.
- Patient recruitment and enrollment. Finding and enrolling eligible patients, especially for rare diseases or narrow biomarker-defined populations, remains a logistics and outreach problem that AI can assist with (site selection, eligibility screening) but not eliminate.
- Manufacturing and CMC scale-up. Moving from a bench-scale synthesis to validated, reproducible commercial manufacturing is a slow, heavily regulated process with its own multi-year timeline, largely independent of how the molecule was discovered.
- Regulatory review. The FDA's standard review clock is roughly ten to twelve months, priority review roughly six to eight -- a floor set by institutional process, not by the speed of drug discovery upstream.
So what does "decades to months" actually mean?
Read Demis Hassabis's public comments carefully and the claim is more precise than the headlines that repeat it: it applies most directly to the discovery phase -- the journey from "here is a validated target" to "here is a well-characterized candidate molecule ready for preclinical development" -- which has historically taken three to six years of iterative wet-lab work. Compressing that phase to months, or specific computational sub-steps to weeks, is a defensible and increasingly demonstrated claim. Compressing the entire discovery-to-approval pipeline to months is not what's being claimed, and conflating the two is where a lot of the public conversation goes wrong.
Faster molecule generation is real. Faster biology and faster regulatory review are a different, much harder problem -- and one that statistics and trial design, not raw compute, are what actually move.
The realistic path to a shorter total timeline
The credible version of this story isn't a single silver bullet that turns fifteen years into fifteen months. It's a stack of partial compressions: faster discovery from generative chemistry and structure prediction, fewer late-stage failures because better preclinical triage catches weak candidates earlier (the most expensive failures are the ones that happen in Phase 2 or 3, after years of investment), and shorter, smarter trials through adaptive designs, biomarker-enriched enrollment, and early-stopping rules grounded in Bayesian decision theory rather than fixed sample sizes calculated up front.
That last piece -- how trials themselves can be run smarter without cutting corners on evidence -- turns out to be one of the least-discussed and most consequential parts of this whole story. It's the subject of the next piece in this series: the statistical infrastructure quietly making faster, more adaptive clinical development possible, even as the underlying biology and regulatory review remain exactly as demanding as they should be.


