arXiv · 2609.23073
MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs
Abstract
Recent advances in natural language processing have led to molecular Large Language Models (LLMs) with strong performance across diverse chemistry tasks. However, they still struggle to capture fine-grained structure-property relationships, particularly how small, localized modifications alter a molecule's behavior. To address this limitation, we introduce MolSC, a dataset of substituent contributions, defined as property changes induced by attaching specific substituents to molecular scaffolds. Curated from manually annotated bioactivity records, MolSC spans structural-alert liability, target-specific bioactivity, and physicochemical descriptors, and contains 181K substituent-level examples for training. We further propose MolSC-Bench, a held-out evaluation benchmark of 1,541 examples disjoint from MolSC at the scaffold, substituent, and molecule levels. Our experiments show that existing molecular LLMs and strong proprietary models such as GPT-5.2 and Gemini-3-Flash show limited reliability in substituent contribution prediction. In contrast, training on MolSC substantially improves this ability and achieves strong performance across diverse downstream molecular tasks. These results highlight substituent contribution learning as a key component of fine-grained molecular understanding.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee. 2026-09-19. MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs. https://arxiv.org/abs/2609.23073
Cite the original work for its findings. Save a collection to share your selection of sources.