Search arXivSearch

arXiv · 2104.12457

Impact of quantum-chemical metrics on the machine learning prediction of electron density

Abstract

Machine learning (ML) algorithms have undergone an explosive development impacting every aspect of computational chemistry. To obtain reliable predictions, one needs to maintain the proper balance between the black-box nature of ML frameworks and the physics of the target properties. One of the most appealing quantum-chemical properties for regression models is the electron density, and some of us recently proposed a transferable and scalable model based on the decomposition of the density onto an atom-centered basis set. The decomposition, as well as the training of the model, is at its core a minimization of some loss function, which can be arbitrarily chosen and may lead to results of different quality. Well-studied in the context of density fitting (DF), the impact of the metric on the performance of ML models has not been analyzed yet. In this work, we compare predictions obtained using the overlap and the Coulomb-repulsion metrics for both decomposition and training. As expected, the Coulomb metric used as both the DF and ML loss functions leads to the best results for the electrostatic potential and dipole moments. The origin of this difference lies in the fact that the model is not constrained to predict densities that integrate to the exact number of electrons $N$. Since an \textit{a posteriori} correction for the number of electrons decreases the errors, we proposed a modification of the model where $N$ is included directly into the kernel function, which allowed to lower the errors on the test and out-of-sample sets.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ksenia R. Briling, Alberto Fabrizio, Clemence Corminboeuf. 2021-11-15. Impact of quantum-chemical metrics on the machine learning prediction of electron density. https://doi.org/10.1063/5.0055393

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

SPIBER: Reconstructing Free Energy Landscapes from Short, Unconverged Trajectories with Generative Flow Networks

Molecular systems have many degrees of freedom, but their metastable behavior can often be described by a few collective variables. Identifying these variables and estimating free energies along them from limited simulation data remains a challenging, important problem. Separate short trajectories may sample different metastable states without capturing transitions or establishing their relative equilibrium populations. For unbiased trajectories generated with the same Hamiltonian at a single temperature, alternate methods based on histogram reweighting cannot correct this imbalance. Here we present SPIBER, which combines the State Predictive Information Bottleneck (SPIB) with Generative Flow Networks (GFlowNets). SPIB uses deep learning to approximate slow degrees of freedom through a past-future information bottleneck, retaining information needed to predict future metastable states. We show that this compression limits conditional entropy variations in populated regions, allowing conditional mean potential energies, which are much easier to calculate, to be used to approximate free energy differences. Given sufficient local sampling to estimate these energies, they define the target distribution for GFlowNets, energy-based generative samplers that sample according to estimated thermodynamic stability rather than observed populations. For a particle in a radial double-well potential, for alanine dipeptide, and for the nine-residue peptide AIB9, SPIBER recovers free energy differences between sampled metastable states to within one thermal energy unit of reference values. The method combines collective-variable learning and free energy estimation in up to four latent dimensions, without requiring converged state populations or additional molecular dynamics simulations.

physics.chem-ph

Intrinsic Matching Frustration in Fluctuating Finite Systems

We formulate intrinsic matching frustration (IMF), a fluctuation-induced, kinetics-independent reduction in the mean capacity permitted by a prescribed matching rule. For complementary one-to-one matching, the instantaneous capacity is set by the minority population, so fluctuations produce a nonzero mean deficit even when the two populations are balanced on average. At finite size, this deficit depends on the full distribution of the population difference and is determined by its variance alone only in the Gaussian limit. Compartmentalization hides matching capacity by preventing cancellation between local imbalances of opposite sign. Fusion releases this hidden capacity monotonically under coarse graining, producing a measurable recovery of product yield following local reaction to completion.

physics.chem-ph

Phonon chirality as an additive control of CISS: a symmetry-protected law

Chirality-induced spin selectivity (CISS) is usually associated with molecular handedness. The possible contribution of chiral phonons is less established. We study a helical tight-binding model in which local phonon angular momentum modulates spin-dependent nearest-neighbor hopping. Fewest-switches surface hopping calculations give the transmitted spin polarization $\mathrm{SP}=aC+b\mathrm{PH}$. Here $C$ is the molecular chirality and $\mathrm{PH}$ is the phonon chirality. A mirror symmetry reverses $C$, $\mathrm{PH}$, and $\mathrm{SP}$ simultaneously. This symmetry excludes both a chirality-independent offset and a $C\cdot\mathrm{PH}$ term. The phonon contribution can therefore enhance, cancel, or reverse the molecular CISS signal.

physics.chem-ph