Search arXiv⌕ Search

arXiv · 2507.01444

A Large Language Model for Chemistry and Retrosynthesis Predictions

Abstract

Large language models (LLM) have achieved impressive progress across a broad range of general-purpose tasks, but their effectiveness in chemistry remains limited due to scarce domain-specific datasets and the demand for precise symbolic and structural reasoning. Here we introduce ECNU-ChemGPT(name after East China Normal University), a chemistry-specialized LLM engineered for deep chemical knowledge understanding and accurate retrosynthetic route planning. Our approach is distinguished by four key strategies: structured prompt-based knowledge distillation from authoritative chemistry textbooks to construct a high-quality question-answering dataset; domain-specific prompt engineering using curated chemical keywords, combined with LLMs APIs for data derivation and knowledge distillation; large-scale fine-tuning on a meticulously cleaned and enriched Pistachio reaction dataset to enhance retrosynthesis prediction accuracy; and integration of BrainGPT, a dynamic multi-model scheduling framework that enables task-specific invocation of multiple specialized models trained for diverse chemistry-related tasks. ECNU-ChemGPT exhibits superior performance on chemistry question-answering and retrosynthetic planning benchmarks, outperforming leading general-purpose models-including Deepseek-R1, Qwen-2.5, and GPT-4o. In retrosynthesis, it achieves a Top-1 accuracy of 68.3% on the USPTO_50K dataset and successfully reconstructed 13 complete experimental pathways for real-world drug molecules from medicinal chemistry journals. These results underscore the effectiveness of domain-adapted fine-tuning combined with dynamic multi-model task scheduling, providing a scalable and robust solution for chemical knowledge question answering and retrosynthetic planning.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yueqing Zhang, Wentao Liu, Yan Zhang, Danyang Xiong, Jihang Zhai, Hao Hao, YuCheng Gu, HaiBo Yang, Shuanhu Gao, Lianrui Hu, Aimin Zhou, Xiao He. 2025-07-10. A Large Language Model for Chemistry and Retrosynthesis Predictions. https://arxiv.org/abs/2507.01444

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Task-Based Framework for Evaluating Raman Spectral Quality Measures

Raman spectral preprocessing and enhancement are often evaluated by comparing output spectra with a reference. Interpreting these comparisons requires evidence that spectral quality measures reflect downstream task performance. We present a controlled-perturbation framework for testing this relationship. Five perturbation types (baseline distortion, independent noise, correlated noise, a global wavenumber shift, and nonlinear axis warping) generate paired changes in a spectral measure (metric harm) and in downstream performance (task harm). An alignment gap (AG) quantifies how much the relationship between metric harm and task harm changes with perturbation type. Ordering concordance (OC) measures how often a metric correctly ranks two conditions by their task harm. The framework evaluates thirteen outputs (MSE, RMSE, MAE, NMSE, spectral angle, Pearson correlation, Wasserstein distance, a structure-to-noise ratio, peak precision, recall, F1, artifact ratio, and missing ratio). Three public datasets provide bacterial classification, sugar-mixture quantification, and mineral identification tasks. PCA with logistic regression, partial least squares regression, and cosine library matching supply the task outcomes. Classifiers and calibrations are fitted either to unperturbed training spectra or to each perturbed training condition, then evaluated on the same perturbed test spectra. Mineral queries are compared with an unchanged or correspondingly perturbed library. The resulting comparisons identify task-specific strengths and limitations, including cases where better ordering does not accompany a smaller AG. Removing axis perturbations and comparing spectra on a common physical grid test how these findings depend on the evaluation design. The framework provides a reproducible procedure for assessing existing measures and testing new candidates against downstream task performance.

physics.chem-ph↗

A System-Independent Metadynamics Strategy for Reactive Training Data: Application to Gas-Phase Organic Reactions

General-purpose machine-learning interatomic potentials (MLIPs) for organic reactions need to be accurate on both the minimum energy path (MEP) for static evaluation of basic properties and the broader configurational space for simulating reaction dynamics. Existing general datasets for gas-phase organic reactions rely on quasi-static relaxation that confines configurations to the MEP vicinity, so models trained on them could fail on direct molecular-dynamics trajectories; the gap is methodological, not a question of dataset size. We introduce a spatiotemporally resolved, system-independent collective variable (CV): Cartesian RMSD within randomly partitioned local domains against an expanding list of time-averaged reference geometries. The CV drives metadynamics as the main exploration engine, supplemented by structural relaxation towards transition state (TS) to augment the coverage around TS. Within a concurrent-learning workflow, this produces OpenRxn26, a dataset of 1.8~M DFT-labeled configurations covering neutral singlet unimolecular reactions in the H/C/N/O chemical space ($N_\mathrm{heavy} \leq 30$), containing reactive atomic environments underrepresented in community datasets. Trained on OpenRxn26, a DPA3 model (denoted DPA3_rxn) achieves transferable accuracy on barrier heights and reaction energies. On off-MEP reactive trajectories, DPA3_rxn is the only model in the benchmark suite to reach 1.0 kcal/mol energy accuracy compared with the labeling method, where the domain MLIP leading on static benchmarks degrades several-fold (e.g. MACE_OMol25), showing the insufficiency of quasi-static sampling and MEP-anchored benchmarks for guaranteeing dynamics reliability of MLIPs. OpenRxn26 thus provides MD-ready reactive training data for gas-phase neutral singlet organic reactions, verifying the generality and efficiency of the sampling strategy.

physics.chem-ph↗

Full-frequency GW from Cayley-transformed self-energy moments

The dynamical GW self-energy approximation is a key computational tool to provide the fundamental spectrum of electronic systems. We reformulate this approximation, representing the particle and hole parts of the GW self-energy through a highly compact set of Cayley-transformed moment constraints. The Cayley transformation maps real frequencies to the unit circle, keeping the moments bounded as their order increases, ensuring numerical stability and allowing resolution to be focused on an energy range of interest. We calculate these Cayley-transformed moments via an efficient O[N$^4$] scaling contour integration, and from them, construct a Hermitian upfolded Hamiltonian with a linearly scaling dimensionality with system size. A single-shot diagonalization of this effective Hamiltonian gives an explicit full-frequency G0W0 Green's function with manifestly real poles and non-negative spectral weights. This enables quasiparticle energies, satellite features, and their spectral weights to be obtained across the full G0W0 spectrum. Comparisons with exact G0W0 calculations and convergence across the GW100 test set and the larger Chlorophyll A molecule demonstrate substantially faster and more reliable convergence with moment order than an earlier monomial-moment approach. These Cayley moment representations therefore provide a stable, compact, and systematically improvable route to the complete spectral information of zero-temperature GW, without explicit frequency grids, plasmon-pole models and other common approximations, or analytic continuation.

physics.chem-ph↗