Search arXiv⌕ Search

arXiv · 2312.01527

NovoMol: Recurrent Neural Network for Orally Bioavailable Drug Design and Validation on PDGFRα Receptor

Abstract

Longer timelines and lower success rates of drug candidates limit the productivity of clinical trials in the pharmaceutical industry. Promising de novo drug design techniques help solve this by exploring a broader chemical space, efficiently generating new molecules, and providing improved therapies. However, optimizing for molecular characteristics found in approved oral drugs remains a challenge, limiting de novo usage. In this work, we propose NovoMol, a novel de novo method using recurrent neural networks to mass-generate drug molecules with high oral bioavailability, increasing clinical trial time efficiency. Molecules were optimized for desirable traits and ranked using the quantitative estimate of drug-likeness (QED). Generated molecules meeting QED's oral bioavailability threshold were used to retrain the neural network, and, after five training cycles, 76% of generated molecules passed this strict threshold and 96% passed the traditionally used Lipinski's Rule of Five. The trained model was then used to generate specific drug candidates for the cancer-related PDGFRα receptor and 44% of generated candidates had better binding affinity than the current state-of-the-art drug, Imatinib (with a receptor binding affinity of -9.4 kcal/mol), and the best-generated candidate at -12.9 kcal/mol. NovoMol provides a time/cost-efficient AI-based de novo method offering promising drug candidates for clinical trials.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ishir Rao. 2023-12-03. NovoMol: Recurrent Neural Network for Orally Bioavailable Drug Design and Validation on PDGFRα Receptor. https://arxiv.org/abs/2312.01527

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

PocketVE: Stable and Property-Guided Structure-Based Drug Design with Variance-Exploding Diffusion

Protein-conditioned 3D molecule generation is a central challenge in structure-based drug design, requiring a balance between pocket compatibility, molecular properties, and physical geometry. We propose \textbf{PocketVE}, a protein-pocket-conditioned variance-exploding (VE) diffusion framework that couples stable coordinate denoising with inference-time property guidance. Specifically, PocketVE combines an EDM-style training and sampling setup for 3D denoising, classifier-free guidance for multi-property steering without external property classifiers, and adaptive protein perturbation as a training-time pocket regularizer. Evaluated on CrossDocked2020 under the GenBench3D protocol, PocketVE improves Valid$_{3\text{D}}$ from 58.6 to 80.6 and reduces strain energy from 457.4 to 127.9 relative to its TAGMol architectural baseline, while retaining competitive docking and molecular-property scores under moderate guidance. A guidance-scale study shows that moderate guidance gives a favorable balance between target-related objectives and geometric quality, whereas stronger guidance can degrade geometry and distributional fidelity. Pocket-permutation and PoseCheck diagnostics further support pocket-specific spatial compatibility with reduced steric conflicts. Overall, the results suggest that geometric stability and inference-time property guidance should be considered as coupled design objectives.

q-bio.BM↗

In Vivo Length Distributions as Mechanistic Fingerprints of Pathological Protein Aggregation

Modern imaging techniques can resolve individual pathological protein aggregates in postmortem human samples, providing detailed measurements of aggregate size distributions that are inaccessible with conventional bulk approaches. These distributions represent mechanistic fingerprints of the microscopic processes that generated the observed pathology, but extracting this mechanistic information requires a quantitative theoretical framework. Here, we develop the mathematical tools needed to interpret aggregate length distributions in living systems, where aggregate growth competes with active removal. We show that, across a class of models, the length distribution of sufficiently large aggregates approaches a geometric decay. Crucially, the decay rate is determined by the balance between aggregate elongation and removal, providing a direct quantitative readout of these competing processes from a single time point measurement. This enables mechanistic comparisons between healthy and diseased human samples without requiring longitudinal measurements of aggregate dynamics. We further analyse how additional aggregation and removal processes modify the observed length distributions. Together, these results establish the mathematical foundations and tools to use aggregate length distributions as an experimentally accessible route for inferring microscopic aggregation dynamics directly from human tissue.

q-bio.BM↗

Decoding enzyme-substrate interaction topology reveals principles underlying catalytic efficiency and mutational outcomes

The enzyme turnover number (kcat) defines catalytic efficiency and constrains quantitative models of metabolism, yet the molecular determinants governing kcat and its response to mutation remain poorly understood. Measurements are sparse and labor-intensive, and most computational approaches provide numerical predictions without explaining how enzyme-substrate interactions shape catalytic outcomes. A central challenge is therefore to identify the topological principles that determine where mutations act and how their functional outcomes are encoded within the enzyme-substrate interaction network. Here, we show that catalytic efficiency and mutational effects can be interpreted through enzyme-substrate interaction topology. We developed Interkcat, an interpretable bidirectional cross-attention framework that captures reciprocal coordination between protein residues and substrate atoms. Optimized on a unified benchmark, Interkcat achieves state-of-the-art predictive performance (R2 = 0.701). From its learned representations, we derive an Interaction Topology Score (ITS) that identifies sequence regions statistically enriched for mutation-sensitive sites without explicit structural inputs. We further demonstrate that higher-order topological features distinguish opposing mutational outcomes: lethal mutations disrupt coordinated networks, whereas activity-preserving or enhancing mutations retain sparse, globally organized coupling. These findings establish interaction topology as a unifying principle linking enzyme sequence, catalytic efficiency, and evolutionary perturbation.

q-bio.BM↗