Drug discovery is described as a search for positive results: a hit, a responder, a significant biomarker, a candidate that crosses a threshold. Everything else is often compressed into “did not work.” That compression destroys information.

A negative experiment can tell us that the target is wrong, the molecule did not reach the target, the assay did not represent the disease, engagement was insufficient, the pathway adapted, the biomarker was disconnected from outcome, the patient population was heterogeneous, or toxicity prevented testing the mechanism. Those are not equivalent failures. They eliminate different regions of the hypothesis space.

Failure is data only when the causal chain is visible

Consider a Phase 2 trial with no efficacy signal. Without adequate pharmacology, the conclusion may be recorded as “target invalid.” But if tissue exposure was never established and target engagement was not measured, the trial did not cleanly test the target. It tested a bundle of assumptions: molecule, formulation, dose, distribution, engagement, disease stage, biomarker, and population.

A useful failure record therefore maps evidence across the full chain:

  • Was the disease mechanism supported in humans?
  • Did the candidate have the required properties?
  • Was relevant tissue exposure achieved?
  • Was target engagement demonstrated?
  • Did proximal biology move?
  • Did downstream disease biology move?
  • Was the population capable of responding?
  • Was the safety margin sufficient to test the hypothesis?

The failed link is the reusable knowledge.

Why organisations repeatedly rediscover the same dead ends

Negative results are scattered across electronic lab notebooks, assay databases, emails, CRO reports, regulatory documents, meeting minutes, and presentations. The final decision often survives only as a sentence: “series deprioritized” or “programme stopped.” Months later, another team sees a promising paper, builds a similar assay, and unknowingly recreates the same failure.

Staff turnover makes the loss worse. Tacit knowledge leaves with scientists. Acquisitions and reorganisations break system boundaries. Positive datasets are curated for publication and modelling; negative datasets remain inconsistently labelled and difficult to search.

Reproducibility problems are partly evidence-architecture problems

Work by Florian Prinz and colleagues at Bayer, and by Glenn Begley and Lee Ellis in oncology, drew attention to the difficulty of reproducing influential preclinical findings. The usual response is to demand better experimental standards, which is necessary. But there is also an information-system problem: protocols drift, reagents change, cell-line identity is uncertain, exclusion decisions are undocumented, and contradictory results are averaged away rather than preserved.

A knowledge base should not ask only “What is the result?” It should ask “Under exactly which conditions did this result hold, and what contradicted it?”

A failure atlas needs a controlled vocabulary

“Failed” is too coarse for machine learning and too coarse for scientific reasoning. A practical taxonomy might include:

  • Technical failure: assay performance, reagent, protocol, measurement, or sample problem.
  • Replication failure: original effect not reproduced under sufficiently matched conditions.
  • Translation failure: effect present in one model but absent in a more human-relevant system.
  • Pharmacokinetic failure: required exposure not achieved or too variable.
  • Target-engagement failure: candidate exposure occurred without sufficient target modulation.
  • Mechanistic failure: engagement occurred but predicted biology did not follow.
  • Safety failure: liability prevented testing the efficacious exposure.
  • Population failure: effect diluted because the responsive subgroup was not identified.
  • Commercial or operational stop: programme ended without falsifying the scientific hypothesis.

That final category is essential. A discontinued programme is not automatically a failed target. Portfolio strategy, funding, manufacturing, competition, or recruitment may stop development despite unresolved biology.

Negative data can train better models -- if denominators exist

Predictive models learn from labelled examples. Discovery datasets are often enriched for compounds that were synthesized, assays that were run, results that were publishable, and programmes that survived long enough to be measured. Missing failures create selection bias. A model trained only on successful chemistry can confuse “what chemists chose to make” with “what chemistry is possible.” A target model trained only on public positive associations can mistake publication frequency for causal importance.

Useful negative data requires denominators: what was attempted, under which conditions, with which detection limits, and why a result was considered negative. An unmeasured endpoint is not zero. A compound excluded because of solubility is not inactive. A trial stopped for business reasons is not evidence of no efficacy.

The next-best experiment should be chosen from the failure map

Once failures are structured, they can change prospective decisions. Before launching an experiment, a system can surface similar compounds, targets, models, and prior contradictions. After a result, it can update which causal link is weakened and recommend the experiment that most efficiently distinguishes competing explanations.

This is more valuable than a repository of reports. It is a decision memory: a versioned map of claims, evidence, uncertainty, contradictions, and outcomes. The goal is not to celebrate failure. It is to stop paying for the same lesson twice.