q-bio.QM cs.AI cs.LG

A target is not enough to design a drug — conditioning molecule generation on disease context

q-bio.QM Ali Motahharynia, Mohammadreza Ghaffarzadeh-Esfahani, Mahsa Sheikholeslami, et al. (7) Jul 2026

Computational drug design has generated molecules conditioned on target proteins or general molecular properties. This work adds disease context, designing small molecules conditioned on both a disease ontology and a target protein sequence.

Paper overview (our summary)

  • Field (arXiv category)q-bio.QM(+2)
  • AuthorsAli Motahharynia, Mohammadreza Ghaffarzadeh-Esfahani, Mahsa Sheikholeslami, et al. (7)
  • Submitted2026-07-09
  • arXiv ID2607.08404v1

Key points

  • DrugGen-2 is a generative model designing small molecules conditioned on both a disease ontology and a target protein sequence.
  • It starts from the observation that existing methods overlook how disease context bears on a target and on therapeutic outcome.
  • GPT-2 is fine-tuned on a dataset linking approved drugs to diseases and targets, in two stages ending in reinforcement learning by group relative policy optimization.
  • The reward functions optimize for chemical validity, novelty, diversity and high predicted binding affinity.
  • On five targets relevant to diabetic nephropathy it is reported to beat the baselines, with molecular docking analyses supporting the result.

1Taking AI outside AI

This series takes up eight papers not about how AI performs or how it meets society, but about the work models are doing inside the natural sciences: drug discovery, materials, chemical reaction, quantum, fluids, the brain, particle physics, soil. The fields have little in common beyond one thing, that the model is up against constraints coming from physics or from life.

2A small corner of the literature

Of the 1,900 records this site holds as of 2026-09-05, those without an article1,858Spread across 68 primary classes
Of those, papers addressing physics, chemistry, materials, life or the earth985.3%, counting each paper once
The largest primary classcs.CV at 471 (25.3%)Then cs.LG at 374 (20.1%) and cs.AI at 357 (19.2%)

Most of the literature is about AI itself: images, language, learning methods. Papers taking a natural science as their subject come to 98, or 5.3%. What that minority is doing is the subject here.

3The claim that a target is not enough

The aspectExisting computational drug designThis work, DrugGen-2
What is given as a conditionA particular target, or general molecular propertiesBoth a disease ontology and the target protein sequence
What has been overlookedHow disease context bears on target behaviour and outcome
How the model is builtGPT-2 fine-tuned on data linking approved drugs to diseases and targets
The second stageReinforcement learning by group relative policy optimization after supervised fine-tuning

For one and the same protein, which disease it is being aimed at within can change which molecule is wanted. The authors supply that as an explicit condition. The reward in the reinforcement stage was built to raise four things together: chemical validity, novelty, diversity, and predicted binding affinity.

4Measured against drugs already approved

Evaluation ran on five targets relevant to diabetic nephropathy, where the model is reported to beat existing ones on generating unique molecules, on structural similarity to approved drugs, and on predicted binding affinity. Specific comparisons with reference drugs are given: against angiotensin-converting enzyme, compounds with predicted affinities of −9.917, −9.485 and −9.367 are set beside enalapril at −8.283.

All of these are predicted rather than measured values, and distinct from experimental verification.

Why it matters

Computational drug design has long been posed as the problem of producing a molecule that binds a target, yet for one and the same target the molecule wanted can differ with the disease being aimed at. Supplying the disease ontology as an explicit condition is an attempt to handle that difference at the point of generation. Papers using AI as an instrument of natural science are a minority in these records at 5.3%, and what they turn on is less raw performance than how the structure of the subject gets folded into the conditioning.

FAQ

What changes by conditioning on a disease ontology?
The paper holds that disease context bears on how a target behaves and on therapeutic outcome, so that the molecule wanted can differ for the same target.
What is group relative policy optimization?
A reinforcement learning method, used here in the stage that follows supervised fine-tuning.
Are the reported affinities measured?
They are predicted values. Molecular docking analyses are offered in support, but that is distinct from experimental verification.

Sources (primary)

Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.

#AI research#Drug discovery#Machine learning#arXiv#Life sciences
Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.