The Antibody Developability Consortium aims to overcome longstanding limitations in antibody dataset diversity by combining proprietary and public sequences within a federated infrastructure, enabling members to train and fine-tune predictive AI models without exposing proprietary molecular data.

A new research effort is looking to tackle one of the persistent challenge of identifying candidates that may prove difficult to manufacture or develop before they reach costly later-stage testing.
The Antibody Developability Consortium has launched with a plan to assemble a dataset of 10,000 antibodies and use it to develop artificial intelligence models capable of predicting potential developability problems earlier in the drug discovery process.
Led by Ginkgo Datapoints, an offering of Ginkgo Bioworks, and Apheris, the consortium brings together AbbVie, argenx, Lundbeck and Takeda as founding members. The collaboration is designed to address gaps in existing antibody datasets, which can be small, fragmented and limited in sequence diversity.
Building a larger developability dataset
Antibody developability covers the biological and physical characteristics that can influence whether a candidate can be manufactured, formulated and ultimately advanced into a clinical product.
Identifying unfavourable properties earlier could help drug developers make better-informed decisions about which candidates to progress. However, predictive models are often trained on datasets that do not contain enough diverse antibody sequences to support robust modelling.
The consortium will seek to address this by combining proprietary antibody sequences contributed by its members with publicly available sequences supplied by Ginkgo Datapoints. The resulting dataset is intended to provide a common resource for training and benchmarking AI models.
“This consortium represents an important step forward in building predictive models for antibody developability by creating datasets that are designed for machine learning, addressing limitations associated with convenience datasets. Federated infrastructure enables participants to contribute data while keeping proprietary sequences private. These capabilities could meaningfully accelerate antibody discovery and help advance new medicines for patients.”
Athena Hadjixenofontos, Director of Data Science, Head of AI in Biotherapeutics and Genetic Medicine, AbbVie.
AI models trained without sharing proprietary sequences
Ginkgo Datapoints will oversee the scientific design of the consortium, including its approach to selecting antibody sequences, antibody production and high-throughput laboratory testing across core developability endpoints.
The company will also train a foundation antibody developability model using the resulting dataset within Apheris’ secure environment.
“We are building the largest, most standardised antibody developability dataset the industry has ever seen, along with the predictive models trained on it,” said Rich Cohen, Senior Director at Ginkgo Datapoints. “Ginkgo Datapoints brings the lab data generation scale, the diverse sequence selection expertise and the published modelling track record needed to lead this initiative.”
Apheris’ federated infrastructure will allow the foundation model to be brought into each member’s environment. Companies will be able to fine-tune models using their own proprietary molecules while retaining ownership of their sequences and assay data.
“For AI to impact developability decisions in a drug program, it has to perform on a pharma’s own molecules,” said Robin Röhm, CEO and Co-Founder of Apheris. “The Antibody Developability Consortium delivers the largest standardised antibody dataset and the foundation model trained on it. Apheris’ federated infrastructure brings that model to each member and lets them fine-tune it on their proprietary molecules inside their own environment.”
Consortium plans to expand research
Independent scientific oversight will come from Charlotte Deane, Professor of Structural Bioinformatics at the University of Oxford, and Peter Tessier, Professor of Pharmaceutical Sciences and Chemical Engineering at the University of Michigan.
The consortium’s first dataset is aiming for delivery to members by early 2027. The partners also plan to look at adding more complex antibody formats and additional properties that could help predict the development prospects of emerging drug classes.
“At argenx, collaboration is central to how we innovate. Through our Immunology Innovation Program, we combine deep disease biology expertise with antibody engineering to advance new medicines for patients with high unmet need,” said Erwin Pannecoucke, Principal Scientist Discovery at argenx. “Predictive developability models can significantly accelerate discovery and development. The Consortium’s extensive antibody dataset and federated design enable every partner to learn together and ultimately bring better medicines to patients faster.”



No comments yet