Virtual screening can assess billions of compounds, but performance can fall on unfamiliar targets. Combining AI with molecular physics could help make predictions more reliable.

Finding a starting point for a new drug can require researchers to screen large numbers of compounds to identify the relatively few that interact with a target.

Virtual screening aims to reduce this experimental burden by assessing compounds computationally and prioritising candidates for laboratory testing.

AI models and large virtual chemical libraries now allow these searches to cover billions of compounds. But greater scale only helps if the predictions reliably identify molecules that perform as expected in the laboratory.

This remains a challenge when models encounter unfamiliar biology. Performance can decline when proteins and ligands differ from those represented in the training data, limiting the usefulness of models for novel targets.

Dr Garegin Papoian, Co-founder and Chief Scientific Officer at Deep Origin, is investigating why this happens and whether combining AI with molecular physics can improve the reliability of virtual screening. Papoian leads the development of the company’s hybrid AI-mechanistic drug discovery models. He previously spent 15 years at the University of Maryland, where his research focused on the physics of biomolecular systems.

Illustration showing the virtual screening workflow from a compound library and drug target through computational screening and candidate selection to experimental testing.

A typical virtual screening workflow. Virtual screening computationally assesses large compound libraries against a drug target to prioritise candidates for experimental testing.

“Virtual screening has been around for more than 40 years, but it’s estimated to be the primary hit-finding strategy in only about 1 percent of drug discovery campaigns. The biggest barrier has been reliability,” he explains.

“Historically, classical physics docking tools produced high false-positive rates and struggled to accurately predict binding free energies. While modern co-folding models aimed to fix this, many fell into a ‘memorisation trap’, performing well on familiar targets but failing when applied to genuinely novel targets.”

Virtual screening has been around for more than 40 years, but it’s estimated to be the primary hit-finding strategy in only about 1 percent of drug discovery campaigns. The biggest barrier has been reliability.

Deep Origin is now testing whether adding physics-based modelling to AI predictions can address this generalisation problem and improve performance when screening against novel targets.

The generalisation problem

Strong performance on computational benchmarks does not necessarily mean that an AI model will perform as well on targets that differ substantially from its training data.

“The existing models seem to operate in large part by memorising the training material. We call this the ‘memorisation trap’,” Papoian says.

Similarity between training and test data can contribute to the problem.

The existing models seem to operate in large part by memorising the training material. We call this the ‘memorisation trap’.

“One driver of this is data leakage in standard AI benchmarks, which leave near-identical proteins and near-twin chemical scaffolds sitting in the training set,” he explains.

“This means that while the models will work well on similar targets, they don’t work as well on novel targets. When prospective wet-lab experiments don’t pan out, trust erodes and we get a greater reliance on laboratory screening, even though it is, in theory, much slower and more expensive.”

For drug discovery teams, this creates a practical problem. A model may appear accurate during retrospective testing but provide less reliable predictions when applied to a new target, where relevant training data may be limited.

To test performance on less familiar protein–ligand complexes, Deep Origin applied stricter filters to its benchmark datasets, excluding proteins with 30 percent or greater sequence identity and ligands with more than 0.4 Tanimoto similarity to the test sets.

“The results varied significantly under this stricter test. We outperformed other virtual screening models, achieving over 50 percent accuracy on the most novel complexes, compared with below 25 percent for other models,” Papoian says.

The team also examined where model performance declined and identified the prediction of precise 3D ligand poses as a particular challenge.

We outperformed other virtual screening models, achieving over 50 percent accuracy on the most novel complexes, compared with below 25 percent for other models.

“We closely analysed the performance of leading co-folding models and discovered that while they are highly accurate in predicting protein binding pockets, their performance falls off specifically when predicting precise 3D ligand poses.”

For unfamiliar proteins and ligands, Papoian says AI-based methods can also generate predictions that are inconsistent with the physical constraints of molecular structures.

“When dealing with proteins and ligands that differ from the training data, pure AI-based methods sometimes produce physically impossible solutions, such as atoms colliding in the same space, molecular joints bent at unnatural angles or chemical rings twisted into unstable shapes.”

Combining AI with physics

Deep Origin has developed DODock, which combines AI-generated predictions with physics-based refinement. The aim is to retain the ability of AI to generate potential binding poses while using molecular forces to assess whether those predictions are physically plausible.

“Most current AI based docking tools rely almost entirely on data. A neural network predicts a 3D structure based on patterns it learned during training, much like a language model predicts the next word in a sentence,” Papoian explains.

“That approach works well when a target looks like examples the model has already seen; however, when it encounters something new, it has no way to tell whether its prediction is reliable. The result can be structures that violate basic chemistry or place the ligand in the wrong pose, with no internal mechanism for recognising either error.”

DODock adds a physics-based refinement step between generating potential ligand poses and selecting the final prediction.

“DODock takes a different approach by combining AI with physics. First, its generative model proposes a set of possible 3D poses. DOFast, a lightweight 80-parameter physics engine, then refines those candidates by calculating molecular forces, including hydrophobic and hydrogen bonding interactions, instead of relying on learned patterns alone,” he says.

Once the candidates have been physically refined, a ranking model selects the predicted pose. DODock has also been tested in a blind prediction of the binding pose of the PCSK9 inhibitor AZD0780.

Molecular structure showing the predicted and crystallographic binding poses of the PCSK9 inhibitor AZD0780 within the protein binding pocket.

Source: Deep Origin

Predicted and crystallographic binding poses of the PCSK9 inhibitor AZD0780. DODock predicted the pose with a heavy-atom RMSD of approximately 1.2 Å.

In its evaluations, Deep Origin reported that DODock achieved 80 percent pose accuracy on OpenBind, compared with 4 to 28 percent for the co-folding models evaluated. On Runs N’ Poses, which tests performance as protein–ligand complexes become less similar to the training data, DODock maintained more than 50 percent accuracy on the least-similar complexes, while the evaluated co-folding models fell to 25 percent or below.

Line graph comparing DODock with four co-folding models, showing DODock maintaining a higher success rate as similarity to the training set decreases.

Source: Deep Origin

DODock maintained a higher success rate for poses that were less similar to the training set compared with the co-folding models evaluated.

Reliable pose prediction on unfamiliar targets could help researchers identify which compounds are worth progressing to experimental testing.

“Our novel approach combines the best of both worlds: the speed and pattern-recognition power of AI with the rigour and generalisability of physics-based modelling,” Papoian says.

Testing predictions in the lab

Ultimately, the reliability of virtual screening depends on whether computational predictions hold up experimentally.

“Benchmarks are inherently retrospective. They measure how well a model performs on examples where the correct answers are already known,” says Papoian.

“The only way to really understand the reliability of an approach is to test it prospectively: have the model identify compounds, then synthesise or source them and test them in the lab.”

Deep Origin screened a virtual library of approximately 80 billion compounds against four drug targets, before selected compounds were tested experimentally.

Diagram showing prospective laboratory validation across four drug targets – CD73, IRAK4, FXI and IL-17A – with the number of compounds tested and hits identified for each.

Source: Deep Origin

All four prospective lab validation programmes returned chemically novel scaffolds. Targets were chosen to represent distinct protein classes and inhibition modalities and to range in difficulty.

For CD73, an enzyme involved in immune regulation, 56 of 183 compounds tested had a biochemical IC50 below 500 µM, giving a hit rate of 30.6 percent. Of these, 54 had an IC50 below 100 µM, including multiple single-digit micromolar compounds.

“Those are the results that matter to drug discovery teams because they show how well a model performs on real molecules in real experiments,” Papoian says.

The approach was also tested against IL-17A, a protein–protein interaction (PPI) target. PPIs can present challenges for small-molecule discovery because their interaction surfaces can differ from the defined binding pockets commonly targeted by small molecules.

“We saw a 3.1 percent hit rate, yielding six active compounds among 194 tested,” Papoian says.

Prospective testing across a wider range of targets will be needed to establish how consistently these results translate to different discovery programmes.

We saw a 3.1 percent hit rate, yielding six active compounds among 194 tested.

Tackling more complex drug targets

Papoian expects computational approaches to be applied to more complex biological interactions, including those involved in emerging drug modalities.

“We’re doing more work on PPIs and fully integrated molecular dynamics. We’re doing pioneering work in hard-to-drug modalities such as molecular glues and targeted protein degraders,” he explains.

“These are areas where the complexity of many-body interactions and other factors defy traditional virtual screening approaches.”

Molecular glues and targeted protein degraders introduce interactions between multiple molecular components, creating different modelling challenges from predicting the binding of a single ligand to a protein. Papoian also sees opportunities to apply computational approaches elsewhere in early discovery, including ADMET, toxicity and property optimisation.

We’re doing pioneering work in hard-to-drug modalities such as molecular glues and targeted protein degraders.

This extends to preclinical drug safety. Deep Origin leads the Pharmacological Research and Evaluation through Digital Integration and Clinical Trials Simulation (PREDICTS) consortium, part of an ARPA-H programme developing models to predict drug toxicity and safety profiles. The work is funded by an Other Transaction Agreement (OTA) of up to $31.7 million from ARPA-H.

However, expanding into these areas will depend on addressing the same issue that has limited virtual screening more broadly.

“The primary bottleneck remains the fundamental issue of reliability. For virtual screening to become the front-line engine in early drug discovery, computational tools must prove they are consistently predictive,” Papoian says.

“Cheaper and faster are not good enough. Industry teams will only make them the primary strategy when they can trust that computational hits translate into wet-lab success.”

Demonstrating that computational predictions translate into experimental results will be key to wider adoption.

“Ultimately, we foresee a future where we have an end-to-end computational discovery engine. But it all starts with virtual screening simply being better and more reliable,” Papoian concludes.