Cytiva's protein research strategy director Paul Belcher says AI is now designing drug candidates from scratch rather than screening existing libraries — but the industry is hitting a data wall that could stall the payoff. Bringing a new drug to market still takes 10 to 15 years, costs between $1 billion and $2.5 billion, and fails more than 90% of the time. Since the 1950s, the cost of developing new pharmaceuticals has roughly doubled every nine years, a pattern known as Eroom's Law, and AI is the industry's biggest bet on breaking it.
The clearest early win is hit identification, the stage where researchers look for molecules that bind to a disease target. Traditional workflows screen hundreds of thousands to millions of compounds using binary yes-or-no assays that produce low-fidelity data. AI now lets drug companies predict binding computationally before committing anything to a wet lab, culling weak candidates before they consume reagents or bench time.
That shift has created a second-order problem. AI generates more candidates, and better ones, but every candidate still has to be validated physically because current models cannot reliably predict kinetics or developability. Lab teams built to run high-throughput screens on simple assays now face a queue of diverse, complex molecules that each need deep characterization, and the instruments were not designed for that mix.
Key facts
- 01Since the 1950s, drug development costs have roughly doubled every nine years — a pattern known as Eroom's Law.
- 02Bringing a new drug to market now takes 10-15 years, costs $1B to $2.5B, and carries failure rates above 90%.
- 03No drug discovered primarily through AI-driven design has full FDA approval; Cytiva's Paul Belcher expects that to change within 2 to 3 years.
- 04A 2016 study by Elisabeth Bik found nearly 4% of biomedical papers contained duplicated or manipulated images — a data-integrity problem generative AI has made worse.
- 05A Stanford study found the cost of training frontier AI models has more than doubled every year since 2016, straining pharma R&D budgets.
The bigger constraint is the training data itself. Most public biomedical datasets and journal articles report only positive results, so models learn what worked without ever seeing what failed. Belcher argues that failed experiments — the compounds that did not bind, the assays that went sideways — are exactly the data that would make AI predictions reliable, but they stay buried in lab notebooks and never enter the literature.
Data integrity compounds the problem. Dutch microbiologist Elisabeth Bik found in 2016 that nearly 4% of biomedical papers contained duplicated or manipulated images, and that was before generative AI made image fabrication trivial. Belcher points to Cytiva's Image Integrity Checker, which uses secure hash algorithms drawn from blockchain to detect tampered scientific images, and says publishing houses are starting to adopt it as a standard verification step.
The endpoint Belcher describes is the autonomous lab — a lab-in-the-loop that runs prediction, testing, and optimization around the clock and feeds every result back into the model that suggested the next experiment. Better starting molecules plus more optimization cycles should mean fewer late-stage failures in clinical trials, where the money actually gets spent.
Getting there requires interoperable instruments and FAIR data — findable, accessible, interoperable, and reusable — flowing in and out of every machine in the lab. Most labs are nowhere close.
Cytiva's pitch is that the vendors owning lab hardware have to open their ecosystems or the autonomous-lab vision collapses. A single closed instrument in a workflow breaks the data loop the AI models depend on.
The economics are also getting harder. A Stanford study found the cost of training frontier AI models has more than doubled every year since 2016, and pharma R&D is already one of the most capital-intensive sectors in the economy. Belcher's view is that AI stays an advantage only as long as the cost of compute does not exceed the cost of clinical development it is meant to avoid — a ratio that is not guaranteed to hold.
No drug discovered primarily through AI-driven design has received full FDA approval yet. Belcher expects the first to clear in the next two to three years, which would be the concrete proof point the sector has been promising investors since the current AI-drug-discovery wave began.
The interesting tension in Belcher's argument is that the biggest gains from AI in pharma are being throttled not by model capability but by data hygiene and lab plumbing — the least glamorous parts of the stack. Frontier model releases get the headlines, but the drug company that first cracks integrated FAIR data across a fully-instrumented lab will likely outrun rivals with better models and messier pipelines. That is a durable competitive moat, and it is one that hardware and workflow vendors like Cytiva are better positioned to sell than any foundation-model lab.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




