AI agents can search literature, analyse data, generate hypotheses and coordinate multistep research tasks. But not every part of drug discovery is suited to automation. In our second Discovery Toolkit, we examine which tasks could be delegated to AI, where scientists need to remain involved and how teams can decide the right level of autonomy.

In September 2026, researchers at Stanford University published the Virtual Biotech, a multi-agent AI framework designed to support therapeutic discovery and development. Led by a virtual Chief Scientific Officer, the system delegates research questions to specialised teams of AI agents working across areas including target discovery, safety assessment, modality selection and clinical development.¹

The researchers tested the system at three drug-development decision points. In one analysis, more than 37,000 AI agents reviewed and classified outcomes from 55,984 clinical trials to investigate features associated with drug-development success. In another, the agents investigated B7-H3 as a potential therapeutic target in lung cancer, integrating statistical genetics, single-cell, spatial and clinicogenomic evidence before proposing an antibody–drug conjugate (ADC) strategy.¹

Importantly, the B7-H3 analysis was conducted without web access, using a model with a January 2025 knowledge cutoff. Later that year, ifinatamab deruxtecan received FDA Breakthrough Therapy Designation for extensive-stage small cell lung cancer.² The drug, discovered by Daiichi Sankyo and jointly developed with Merck, is a B7-H3-targeted ADC, aligning with both the target and therapeutic modality proposed by the agents.¹,² However, this does not constitute experimental validation of the agents’ proposal. The Stanford team plans to test other findings from the Virtual Biotech in the laboratory.³

This is one example of AI being used for more than a single prediction or analysis.

AI agents can be given an objective, break it into tasks, use external information and computational tools, assess the results and decide what to do next. Multiple specialised agents can also work on different parts of the same problem.⁴

So which tasks should be given to AI agents and which still need a scientist?

The answer depends on the task, the quality of the available data, how easily the output can be checked and the consequences of getting it wrong.

What is an AI agent?

An AI agent is AI software that can work towards an objective through multiple steps. It can select tools, analyse results and use those results to decide what to do next.

Unlike an AI tool performing one predefined task, an AI agent can determine and carry out subsequent tasks towards a wider objective.

What makes an AI agent different?

Many AI tools used in drug discovery perform a defined task. A model might predict a molecular property, rank compounds or analyse an experimental dataset.

An AI agent can take greater control over the process. Rather than receiving an input and returning a single output, it can retrieve information, choose tools, perform analyses and use the results to determine its next action.⁴

Some systems divide these responsibilities between several agents.

Robin, a multi-agent system reported in Nature in 2026, combines specialised agents for literature searching and data analysis. The system generated therapeutic hypotheses for dry age-related macular degeneration (dAMD), proposed experiments and analysed the resulting experimental data. Human researchers carried out the laboratory experiments and provided the resulting data to Robin, which used them to generate further therapeutic hypotheses.⁵

Diagram comparing a traditional AI tool with an AI agent. Traditional AI performs a defined task before a researcher reviews the output. An AI agent works towards an objective by planning tasks and accessing tools such as scientific literature, databases, computational tools, experimental data and laboratory automation, with researcher review.

A comparison of traditional AI tools and AI agents in drug discovery. While a traditional AI model performs a defined task and returns an output for researcher review, an AI agent can plan multiple steps, select relevant tools and use new results to inform subsequent actions. Researchers remain responsible for reviewing outputs and deciding what happens next.

Where can AI agents actually help?

Drug discovery contains many tasks that could be supported by AI agents, but they do not all require the same level of autonomy.

Applications range from literature synthesis and data analysis to protocol generation, toxicity prediction, small-molecule synthesis and drug repurposing.⁴

Their role can range from gathering and organising evidence to generating hypotheses, analysing experimental results and helping determine what should be investigated next.

But being able to perform a task does not necessarily mean an agent should perform it independently. The more important question is what researchers should hand over.

What should scientists hand over?

Some tasks are easier to delegate than others.

Literature retrieval, data extraction and defined computational analyses have relatively clear inputs and outputs. Researchers can inspect the sources, check the result or repeat an analysis.

By contrast, target assessment is less straightforward. Researchers may need to weigh conflicting evidence, account for disease biology and decide whether the evidence is strong enough to justify further work. Molecule design can present similar problems when programme-specific constraints are not fully represented in the data available to an AI system.

AI methods are often tested using benchmarks, which measure how well they perform on a particular dataset or task. But performing well in these tests does not necessarily mean an AI system will improve decisions in a real drug discovery programme. A 2026 Nature Reviews Drug Discovery assessment found that evidence for clinically relevant impact from AI remains limited. The authors argue that AI should also be evaluated on whether it improves real-world decision-making.⁶

Rather than asking whether an AI agent or scientist is better overall, it is more useful to consider what each can contribute to a particular task.

AI agent, scientist or both?

What an AI agent should handle depends on the task. Agents can process information and perform defined computational work at scale, while scientists contribute experimental context, biological interpretation and judgement.

How AI agents and scientists can contribute to drug discovery tasks, from literature analysis and hypothesis generation to experimental design and decision-making.
Task AI agent Scientist Working together
Literature analysis Searches, retrieves and organises information Assesses relevance, conflicting evidence and biological context Agent organises evidence; scientist evaluates it
Data analysis Runs defined computational analyses at scale Assesses experimental context and potential artefacts Agent performs analysis; scientist interprets the results
Hypothesis generation Combines information to generate candidate hypotheses Assesses biological plausibility and relevance Agent generates hypotheses; scientist prioritises them
Target assessment Brings together evidence from multiple sources Assesses disease relevance and strength of biological rationale Agent collates evidence; scientist evaluates the case
Molecule design Uses predictive and generative tools to propose molecules Applies medicinal chemistry knowledge and programme constraints Agent proposes designs; scientists assess and refine them
Experimental design Proposes experiments or protocols Assesses biological and practical constraints Agent proposes experiments; scientist reviews them
Wet-lab execution Can direct or interact with compatible laboratory automation Handles experimental problems and unexpected observations Agent plans or analyses; automated or human laboratory executes
Unexpected findings Flags results that differ from expected patterns Investigates the result and reassesses assumptions Agent identifies the result; scientist determines its significance
Scientific decisions Organises evidence and proposes possible actions Weighs evidence, uncertainty and programme context Agent supports the decision; scientist retains oversight of consequential decisions

When does an AI output become evidence?

An AI-generated hypothesis is not experimental evidence. Agents can assess targets, propose molecules and recommend experiments, but those outputs still need to be tested.

For example, Robin generated hypotheses and proposed experiments for dAMD, while human researchers carried out the laboratory experiments. Robin then analysed the resulting data and used them to generate further therapeutic hypotheses.⁵

Not every output requires every step. A literature search used to identify papers for further reading needs different checks from an agent-generated hypothesis that could determine whether a target or compound progresses.

The level of validation should depend on how the output will be used and what happens if it is wrong.

Toolkit takeaway: The greater the consequence of an incorrect AI output, the stronger the validation it needs.

What happens when the agent chooses the next step?

Giving an agent control over the next step means an error can influence what it does next.

If an agent relies on unreliable evidence, misinterprets an analysis or chooses an unsuitable tool, that error may affect subsequent actions. A 2026 review of AI agents in drug discovery identifies data heterogeneity, system reliability, privacy and benchmarking among the challenges that need to be addressed.⁴

Researchers therefore need to be able to trace the sources, analyses and tools used by an agent, and decide where checks are needed before it can continue.

The key question is not only whether an agent can complete a task, but whether it should be allowed to act on the result without scientist review.

Toolkit takeaway: Decide where an agent must stop for review before it is allowed to act on its results.

Can AI agents run experiments too?

An AI agent can propose an experiment, but it cannot carry out the physical laboratory work itself. When connected to laboratory automation, however, the agent can direct automated equipment to execute its proposed experiments.

Coscientist, reported in Nature in 2023, demonstrated an AI system capable of planning and carrying out chemistry experiments using tools including literature searches, code execution and automated laboratory equipment.⁷

More recent systems have examined how multi-agent architectures can interact with laboratory hardware. AutoLabs, reported in Scientific Reports in 2026, uses specialised agents to turn natural-language experimental requests into hardware-ready protocols for a high-throughput liquid handler.⁸

Connecting computational and experimental steps enables a closed-loop workflow, where the results of one experiment inform what should be tested next.

Circular workflow showing five stages: design, execute, measure, analyse and select next experiment, with results feeding back into experimental design.

A closed-loop experimental workflow in which results are analysed and used to select and design subsequent experiments.

Agentic AI combined with laboratory automation could support closed-loop design–make–test–analyse workflows, where experimental results are returned to the computational system and used to determine subsequent experiments.⁹ But the reliability of the experiment remains just as important as the capabilities of the agent.

The assay needs to produce reliable measurements, the laboratory equipment needs to execute the protocol correctly and the resulting data need to be captured in a form the system can interpret. A poorly performing assay does not become useful simply because the next experiment is selected automatically.

Toolkit takeaway: Connecting an AI agent to laboratory automation does not remove the need for reliable assays, data and experimental controls.

Where should the scientist stay in the loop?

There is no single level of autonomy that will suit every discovery workflow. Teams can consider three questions before deciding how much responsibility to give an AI agent.

1. Can the result be independently checked? A clearly defined task with a verifiable output may be more suitable for automation.

2. What are the consequences of an incorrect result? If an error could lead to wasted experiments or a poor programme decision, human review becomes more important before the agent determines the next step.

3. Does the next step require scientific judgement? Conflicting evidence, unexpected biology and decisions about whether the available evidence is sufficient may require scientists to consider the wider experimental and disease context.

These questions can help teams decide whether a task is suitable for automation, collaboration between AI and scientists or scientist-led decision-making.

Toolkit takeaway: Automate based on the task, not simply because an agent can perform it. 

The examples below illustrate where automation may be appropriate, where AI can support scientists and where scientific judgement is likely to remain central.
Automate AI + scientist Scientist-led
Literature retrieval Evidence synthesis Interpreting unexpected biology
Data extraction Hypothesis generation Resolving conflicting evidence
Defined computational analyses Target assessment Assessing whether evidence is sufficient
Routine tool execution Molecule design Major go/no-go decisions
Standardised tasks with clear outputs Experimental planning Decisions with major programme consequences

These categories are not fixed. The same task may require different levels of oversight depending on the programme, available evidence and consequences of an incorrect result.

Choosing what to automate

AI agents can take responsibility for more of a research workflow than conventional AI models, but greater autonomy is not necessarily useful for every task.

Defined and independently checkable tasks may be suitable for greater automation. Tasks that combine computational analysis with biological interpretation may work better with an agent supporting a scientist. Decisions based on uncertain evidence or unexpected experimental findings require closer scientific involvement.

Discovery teams therefore need to decide which tasks can be automated reliably, where scientist oversight is needed and which decisions should remain scientist-led.

References

  1. Zhang HG, Eckmann P, Miao J, Mahon AB, Zou J. The Virtual Biotech: a multi-agent AI framework for therapeutic discovery and development. Science. Published online 17 September 2026. doi:10.1126/science.aeg6779.
  2. Merck. Ifinatamab deruxtecan granted Breakthrough Therapy Designation by U.S. FDA for patients with pretreated extensive-stage small cell lung cancer. 18 August 2025.
  3. Armitage H. Virtual biotech company puts thousands of AI scientist agents to work on drug discovery. Stanford Medicine. 17 September 2026.
  4. Huynh DL, Seal S, AIA4S Consortium, et al. AI agents in drug discovery: applications and case studies. Drug Discov Today. 2026;31(3):104650. doi:10.1016/j.drudis.2026.104650.
  5. Ghareeb AE, Chang B, Mitchener L, et al. A multi-agent system for automating scientific discovery. Nature. 2026;655(8122):497–505. doi:10.1038/s41586-026-10652-y.
  6. Bender A, Thomas MC, Scannell JW, et al. Artificial intelligence in drug discovery — what it is, where we stand and the path forward. Nat Rev Drug Discov. Published online 7 August 2026. doi:10.1038/s41573-026-01496-2.
  7. Boiko DA, MacKnight R, Kline B, Gomes G. Autonomous chemical research with large language models. Nature. 2023;624(7992):570–578. doi:10.1038/s41586-023-06792-0.
  8. Panapitiya G, Saldanha E, Job H, Hess O. AutoLabs: cognitive multi-agent systems with self-correction for autonomous chemical experimentation. Sci Rep. 2026;16(1):19554. doi:10.1038/s41598-026-45593-z.
  9. Vijayan RSK, Cross JB, Poongavanam V. Entering the agentic era of AI in drug discovery. Nat Chem Biol. Published online 19 August 2026. doi:10.1038/s41589-026-02295-x.