The Field Records Its Successes. The Knowledge Is in the Failures.
AI models in drug discovery learn almost entirely from what worked. But much of what the field needs to know is buried in its failures, and abandoned drugs may be the training data it never had.

Last month I wrote that the hard part of drug discovery isn't designing the molecule, it's understanding the disease well enough to know what to aim at. So if biology is the bottleneck, what would help us understand it better? A lot of the answer is sitting in the field's failures, and they get far less attention than they deserve.
Models Learn From What Gets Published, and Failure Rarely Does
The public data that AI models draw on in drug discovery skews heavily toward what worked. That's understandable, since the field mostly publishes its successes, but it has consequences.
The scale of the tilt is easy to underestimate. One estimate suggests that around 60 percent of negative results in drug research never reach the public record at all. Failed experiments, compounds that didn't bind, trials that went nowhere: most of these quietly stay in the drawer. A researcher reading the literature knows to correct for this, at least roughly. A model does not. It takes the record it is given as the whole picture, and that record is a catalogue of successes.
It's worth being concrete about what this does. A model trained on data that leans so heavily toward what worked learns the shape of success in detail and almost nothing about failure, which is exactly the part that would sharpen its predictions. It's a strange way to train something: like learning medicine only from the patients who recovered.
And the failure data isn't gone in any single sense. Some of it was never properly recorded, some sits unpublished in company files, and some exists but scattered across trial registries and regulatory filings, rarely in a form a model could use. However it went missing, the result is the same: it never becomes training data, even though it would help. In chemistry, a 2016 study trained a model on failed reactions from materials synthesis, the kind that never leave the lab notebook, and it predicted successful new reactions more accurately than a model trained only on published, successful ones. The failures marked the edges of what works. As one recent paper put it, a literature filtered for success can't teach a model when to abandon a hypothesis, one of the most important judgments in science.
There's a second consequence too. Because most groups train on the same public data, they tend to reach the same conclusions, with diminishing returns as everyone converges. This is the data wall. It isn't that the models are weak. It's that they've mostly run out of new things to learn from the data everyone already has.
A Failed Trial Is Not Always a Failed Drug
Even when failure data does surface, it tends to be read the wrong way. A trial that misses its endpoint gets filed under one word: failed. Development stops, the asset is shelved, and the molecule is often never tested again.
But as a clinician, this is the distinction I keep coming back to. A trial can fail without the drug being useless. Most diseases we treat as one thing are really several, biologically distinct conditions wearing the same name. When a drug is tested across a broad, mixed population, a real benefit in one subgroup can be diluted into statistical noise by everyone it didn't help. The average comes back flat. The trial fails. And a therapy that genuinely worked for a specific group of patients disappears along with it.
The question a failed trial should prompt, then, isn't only "does this drug work?" It's "did this drug fail, or did the trial fail to find the patients it worked for?" Those are completely different questions, and only one of them means the science is over.
None of this is only theoretical. Drug rescue has a long track record: sildenafil, thalidomide and sorafenib all failed a first path and went on to succeed elsewhere, some through careful reasoning rather than chance. What's changing now is the ability to do it systematically, revisiting failed assets with molecular, clinical and computational evidence together rather than waiting for a lucky observation. The shift is from rescue as accident to rescue as method.
Reading Failure Takes More Than a Model
So the field has a gap at its center. It learns from its successes, but far less from its failures, because most of those were never recorded, shared, or assembled anywhere a model could reach. And the failures that do surface are read too fast, filed as dead ends when some were only trials that found the wrong patients. Abandoned drugs sit on top of that gap. They are, in a real sense, the knowledge the field never collected.
This is the thinking behind NeoRevive, our drug rescue platform, and Clarisyn, the system underneath it. Clarisyn works across published literature and available biomedical datasets, and its computational layer doesn't just gather evidence, it analyses, predicts, and cross-checks it to reconstruct what can actually be established about a failed drug: what it did, which mechanism it acted on, who might have responded. But no model can recover what was never captured in the first place. That's where our clinical team comes in, not to clean up after the model, but as a different kind of evidence: interrogating why development really stopped, and whether the failure is biologically recoverable or genuinely final. One layer reconstructs what the record can support; the other reads what the record never held.
None of this replaces new discovery, and rescue alone won't fill the pipeline. But if the hard part of drug discovery is understanding biology, then failures are some of the most honest biology we have: experiments that already ran, in real patients, at full cost. Reading them properly means putting evidence together that was never in one place to begin with, from the literature, from computation, and from people who have seen how these diseases actually behave.
Accelerate Your Translational Research
Experience how Clarisyn transforms complex scientific problems into structured, evidence-based reports.
Get Started with Clarisyn