A target is not enough to design a drug — conditioning molecule generation on disease context
Computational drug design has generated molecules conditioned on target proteins or general molecular properties. This work adds disease context, designing small molecules conditioned on both a disease ontology and a target protein sequence.
Paper overview (our summary)
- Field (arXiv category)q-bio.QM(+2)
- AuthorsAli Motahharynia, Mohammadreza Ghaffarzadeh-Esfahani, Mahsa Sheikholeslami, et al. (7)
- Submitted2026-07-09
- arXiv ID2607.08404v1
Key points
- DrugGen-2 is a generative model designing small molecules conditioned on both a disease ontology and a target protein sequence.
- It starts from the observation that existing methods overlook how disease context bears on a target and on therapeutic outcome.
- GPT-2 is fine-tuned on a dataset linking approved drugs to diseases and targets, in two stages ending in reinforcement learning by group relative policy optimization.
- The reward functions optimize for chemical validity, novelty, diversity and high predicted binding affinity.
- On five targets relevant to diabetic nephropathy it is reported to beat the baselines, with molecular docking analyses supporting the result.
1Taking AI outside AI
This series takes up eight papers not about how AI performs or how it meets society, but about the work models are doing inside the natural sciences: drug discovery, materials, chemical reaction, quantum, fluids, the brain, particle physics, soil. The fields have little in common beyond one thing, that the model is up against constraints coming from physics or from life.
2A small corner of the literature
Most of the literature is about AI itself: images, language, learning methods. Papers taking a natural science as their subject come to 98, or 5.3%. What that minority is doing is the subject here.
3The claim that a target is not enough
For one and the same protein, which disease it is being aimed at within can change which molecule is wanted. The authors supply that as an explicit condition. The reward in the reinforcement stage was built to raise four things together: chemical validity, novelty, diversity, and predicted binding affinity.
4Measured against drugs already approved
Evaluation ran on five targets relevant to diabetic nephropathy, where the model is reported to beat existing ones on generating unique molecules, on structural similarity to approved drugs, and on predicted binding affinity. Specific comparisons with reference drugs are given: against angiotensin-converting enzyme, compounds with predicted affinities of −9.917, −9.485 and −9.367 are set beside enalapril at −8.283.
All of these are predicted rather than measured values, and distinct from experimental verification.
Why it matters
Computational drug design has long been posed as the problem of producing a molecule that binds a target, yet for one and the same target the molecule wanted can differ with the disease being aimed at. Supplying the disease ontology as an explicit condition is an attempt to handle that difference at the point of generation. Papers using AI as an instrument of natural science are a minority in these records at 5.3%, and what they turn on is less raw performance than how the structure of the subject gets folded into the conditioning.
FAQ
What changes by conditioning on a disease ontology?
What is group relative policy optimization?
Are the reported affinities measured?
Sources (primary)
Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.
- arXiv abstract page (original, official)
- PDF (arXiv)
- arXiv ID: 2607.08404