Search arXiv⌕ Search

arXiv · 2610.09837

Origins of Universal Machine Learning Force-Field Errors in Multicomponent Materials

Abstract

Universal machine learning force-field generalization to multicomponent environments generated by compositional design remains insufficiently assessed. We construct a benchmark of 7,599 multicomponent configurations inspired by high-entropy design, elemental substitution and anion mixing. Eleven pretrained models are evaluated against density functional theory for energies, forces and stresses, with assessment extended to elastic, vibrational and adsorption-related properties. Force errors are analysed through training-reference coverage, local geometric heterogeneity, distance directionality and elemental response. Distances to training-reference environments reveal a qualitative association between coverage differences and increasing errors, while substantial variation remains at similar distances. Higher-error groups show greater local geometric heterogeneity, although OMat24 provides broad coverage of these environments. Relative to training-reference pair medians, errors remain low near the median, rise steeply on the compression side and increase more weakly on the extension side. After matching element pairs and absolute distance deviations, compression-side force errors are 1.81-1.95 times extension-side errors. Model-predicted pairwise interaction curves show greater curvature under compression. Fitting difficulty in independent elemental systems correlates with electronic band-energy responses to atomic displacements and Fermi-level shifts, and a similar pattern is observed in multicomponent systems. In parameter-matched comparisons, spherical-harmonic representations with maximum degrees of 2 and 4 lower test force errors for 38 and 40 of 43 elements, respectively, while differences in elemental difficulty remain. These findings inform force-field selection for experimental compositional design and identify targets for training-data sampling and model representations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hongwei Du, Dingyang Lv, Baole Wei, Yu Ren, Feng Yu, Xin He, Bonan Zhu, Jiahui Liu, Yongda Huang, Yongheng Li, Jianjun Liu, Siqi Shi, Hong Wang, Ziheng Lu. 2026-10-08. Origins of Universal Machine Learning Force-Field Errors in Multicomponent Materials. https://arxiv.org/abs/2610.09837

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Predictive Inorganic Synthesis based on Machine Learning using Small Data sets: a case study of Hydrodynamic Diameter-controlled Cu Nanoparticles

Cu NPs have a broad applicability, yet their synthesis is sensitive to subtle changes in reaction parameters. This sensitivity, combined with the time- and resource-intensive nature of experimental optimization, poses a major challenge in achieving reproducible and size-controlled synthesis. While ML shows promise in materials research, its application is often limited by scarcity of large high-quality experimental data sets. This study explores ML to predict the DLS-derived hydrodynamic diameter of Cu NPs using a small data set of 25 syntheses. Latin Hypercube Sampling is used to efficiently cover the parameter space while creating the experimental data set. Ensemble regression models successfully predict hydrodynamic diameters with good predictive performance given the limited dataset. Since quantitative regression requires a unique DLS-derived hydrodynamic diameter, the regression model is restricted to mono-modal DLS distributions, while a complementary classification model identifies synthesis conditions for which quantitative prediction is applicable. Using equivalent out-of-sample validation, the ML and DoE models showed comparable generalization. The final ensemble model achieved an R2=0.74 compared to 0.60 for the DoE model, while retaining the complete synthesis parameter space, making it better suited for synthesis guidance. Additionally, classification models using both random forests and LLMs are evaluated to distinguish between large and small particles. These classification models exhibited only modest predictive performance, indicating that this small dataset is insufficient to fully exploit the capabilities of complex LLMs. Overall, this study demonstrates that carefully curated small data sets, paired with robust classical ML, can effectively support the synthesis of Cu NPs and highlights that for lab-scale studies, complex models like LLMs may offer limited benefits.

cond-mat.mtrl-sci↗

Discontinuous character of the ultrafast exciton Mott transition in monolayer WS$_2$

There are conflicting predictions and reports on the character of the exciton Mott transition (EMT) in monolayer transition metal dichalcogenides. It could be either a discontinuous or a continuous transition from the excitonic to the plasma phase, with important implications for devices such as photoswitches. To resolve the nature of the transition in monolayer WS$_2$, we study its ultrafast optical response upon resonant photoexcitation of the A exciton across a broad range of photoexcitation densities. In agreement with previously reported measurements we observe that the A exciton quenches gradually with increasing excitation density. However, a detailed lineshape analysis unveils an abrupt red shift in the transient peak positions of the A and B exciton resonances above an excitation density threshold. This is attributed to band gap renormalization arising from the formation of free charge carrier plasma, i.e., the EMT. The plasma phase decays with a time constant of 0.65 ps back into the excitonic state. The abrupt appearance of the plasma phase at the threshold density suggests that the EMT is a discontinuous and not a continuous transition. This work demonstrates how transient optical spectroscopy combined with lineshape analysis of two excitonic resonances simultaneously can be used to investigate the EMT in 2D materials.

cond-mat.mtrl-sci↗

Crystallisation kinetics of supercooled liquid palladium

In this study, we employ classical molecular dynamics (MD) simulations to investigate the crystallisation kinetics of supercooled liquid palladium and relate the results to time-resolved X-ray diffraction measurements on rapidly quenched Pd thin films. Crystal nucleation and growth rates are determined over the temperature range $700$--$1150~\mathrm{K}$ ($0.38$--$0.65 T_{\mathrm{m}}$) by analysing the evolution of the microstructure during the liquid-to-crystal transition. The self-diffusion coefficient of Pd, obtained from the atomic mean-squared displacement, follows Arrhenius behaviour over the investigated temperature range, with an activation energy of $467(6)~\mathrm{meV/atom}$, consistent with available data for supercooled liquid metals. The steady-state homogeneous nucleation rate exhibits a maximum of approximately $4 \times 10^{35}~\mathrm{m^{-3} s^{-1}}$ near $0.5 T_{\mathrm{m}}$. Crystal growth occurs at velocities of the order of metres per second, with a temperature dependence consistent with diffusion-limited Wilson-Frenkel kinetics rather than the collision-limited regime. Based on multiple statistically independent simulations, a time-temperature-transformation (TTT) diagram for crystallisation onset is constructed. The TTT curve exhibits a nose near $0.5 T_{\mathrm{m}}$ and $100~\mathrm{ps}$, corresponding to a critical cooling rate for vitrification on the order of $10^{13}~\mathrm{K s^{-1}}.$ The simulations reproduce the crystallisation onset time and temperature observed in time-resolved X-ray diffraction experiments on optically molten Pd thin films quenched at $5 \times 10^{11}~\mathrm{K s^{-1}}.$ These results indicate that homogeneous, rather than heterogeneous, nucleation governs the achievable supercooling in the experimentally studied films.

cond-mat.mtrl-sci↗