Insilico Medicine have announced their Drug Discovery and Development (DDD) Benchmarks as a Service (BaaS) platform offers a standardised framework for evaluating frontier AI and foundation models across medicinal chemistry, disease biology, clinical development and end-to-end drug candidate programmes.

Insilico Medicine has launched a new benchmarking platform designed to measure whether artificial intelligence models can genuinely support drug discovery rather than simply perform well on existing tests.
The clinical-stage biotechnology company said its new Drug Discovery and Development (DDD) Benchmarks as a Service (BaaS) provides a standardised framework for evaluating AI models using carefully decontaminated real-world datasets and proprietary validated programmes.
The launch comes as the pharmaceutical industry increasingly adopts both general-purpose and specialist AI models to accelerate research. However, many publicly available AI benchmarks have been criticised for relying on data that models may already have encountered during training, potentially inflating performance and failing to reflect real-world decision making.
Measuring real-world performance
The DDD Benchmark has been developed to assess how frontier AI and foundation models perform across a range of drug discovery tasks, including medicinal chemistry, chemical synthesis, disease biology, clinical development and longevity research.
The platform is available to organisations developing AI systems for drug discovery or using foundation models in pharmaceutical research. Models can be assessed through a standard chat-completions API, with Insilico comparing their outputs against expert reference baselines before providing a standardised performance scorecard.
Organisations can also receive a verified score report for internal use or publish their results on a public leaderboard to enable comparisons with other AI models.
The benchmark consists of two complementary evaluation suites:
- Drug Discovery Foundations includes more than 300 assessments based on proprietary out-of-distribution test sets and rigorously decontaminated public datasets. These evaluate a model’s capabilities across disease biology, molecular property prediction and optimisation, retrosynthesis, structure-based drug design and clinical development.
- Drug Candidate Essentials focuses on a model’s ability to support an end-to-end drug discovery programme, from identifying promising compounds through to preclinical candidate nomination. The assessments are based on Insilico’s own validated drug discovery programmes and are designed to test whether AI can make the complex sequential decisions required during pharmaceutical research.

Beyond benchmark scores
The company said the DDD Benchmark has also been designed to evaluate the growing use of AI agents capable of planning experiments, analysing scientific data and using external research tools, measuring whether these capabilities translate into effective real-world drug discovery rather than success on abstract benchmark tasks.
“The rapid progress of AI has made one question more urgent than ever: can these models actually discover drugs?” said Dr Alex Zhavoronkov, Founder and CEO of Insilico Medicine. “For more than a decade, we have built and validated generative AI across the entire drug discovery value chain, nominating 31 preclinical candidates in six years and advancing an AI-discovered and AI-designed medicine into Phase III.
”The DDD Benchmark converts that real-world experience into a rigorous, standardised yardstick that the whole field can use. As we advance toward pharmaceutical superintelligence, measuring genuine capability, not memorised answers, is how we ensure AI delivers the highest-quality, differentiated medicines and helps extend healthy, productive longevity for people everywhere.”
Building on existing AI platform
The DDD Benchmark builds on Insilico’s Pharma.AI platform and its MMAI Gym, a post-training environment for scientific AI.
Since the company was founded, it has nominated 31 preclinical drug candidates, secured more than 10 investigational new drug clearances and reduced the time needed to nominate a preclinical candidate to around 12 to 18 months, compared with the 2.5 to more than four years typically required using traditional drug discovery methods.
Its lead programme, Rentosertib (ISM001-055), is a first-in-class AI-discovered and AI-designed TNIK inhibitor that is currently in Phase III development for idiopathic pulmonary fibrosis (IPF).



No comments yet