Search arXivSearch

arXiv subjects

Luca Ghiringhelli

Publications and source records attributed to Luca Ghiringhelli.

5 recordsLinked to original sources

Perspective: Towards sustainable exploration of chemical spaces with machine learning

Artificial intelligence is transforming molecular and materials science, but its growing computational and data demands raise critical sustainability challenges. In this Perspective, we examine resource considerations across the AI-driven discovery pipeline--from quantum-mechanical (QM) data generation and model training to automated, self-driving research workflows--building on discussions from the ``SusML workshop: Towards sustainable exploration of chemical spaces with machine learning'' held in Dresden, Germany. In this context, the availability of large quantum datasets has enabled rigorous benchmarking and rapid methodological progress, while also incurring substantial energy and infrastructure costs. We highlight emerging strategies to enhance efficiency, including general-purpose machine learning (ML) models, multi-fidelity approaches, model distillation, and active learning. Moreover, incorporating physics-based constraints within hierarchical workflows, where fast ML surrogates are applied broadly and high-accuracy QM methods are used selectively, can further optimize resource use without compromising reliability. Equally important is bridging the gap between idealized computational predictions and real-world conditions by accounting for synthesizability and multi-objective design criteria, which is essential for practical impact. Finally, we argue that sustainable progress will rely on open data and models, reusable workflows, and domain-specific AI systems that maximize scientific value per unit of computation, enabling efficient and responsible discovery of technological materials and therapeutics.

cs.LG

Latent Point Collapse on a Low Dimensional Embedding in Deep Neural Network Classifiers

The configuration of latent representations plays a critical role in determining the performance of deep neural network classifiers. In particular, the emergence of well-separated class embeddings in the latent space has been shown to improve both generalization and robustness. In this paper, we propose a method to induce the collapse of latent representations belonging to the same class into a single point, which enhances class separability in the latent space while enforcing Lipschitz continuity in the network. We demonstrate that this phenomenon, which we call \textit{latent point collapse}, is achieved by adding a strong $L_2$ penalty on the penultimate-layer representations and is the result of a push-pull tension developed with the cross-entropy loss function. In addition, we show the practical utility of applying this compressing loss term to the latent representations of a low-dimensional linear penultimate layer. The proposed approach is straightforward to implement and yields substantial improvements in discriminative feature embeddings, along with remarkable gains in robustness to input perturbations.

cs.LG

Extrapolation to complete basis-set limit in density-functional theory by quantile random-forest models

The numerical precision of density-functional-theory (DFT) calculations depends on a variety of computational parameters, one of the most critical being the basis-set size. The ultimate precision is reached with an infinitely large basis set, i.e., in the limit of a complete basis set (CBS). Our aim in this work is to find a machine-learning model that extrapolates finite basis-size calculations to the CBS limit. We start with a data set of 63 binary solids investigated with two all-electron DFT codes, exciting and FHI-aims, which employ very different types of basis sets. A quantile-random-forest model is used to estimate the total-energy correction with respect to a fully converged calculation as a function of the basis-set size. The random-forest model achieves a symmetric mean absolute percentage error of lower than 25% for both codes and outperforms previous approaches in the literature. Our approach also provides prediction intervals, which quantify the uncertainty of the models' predictions.

physics.comp-ph

Discovering dynamic laws from observations: the case of self-propelled, interacting colloids

Active matter spans a wide range of time and length scales, from groups of cells and synthetic self-propelled particles to schools of fish, flocks of birds, or even human crowds. The theoretical framework describing these systems has shown tremendous success at finding universal phenomenology. However, further progress is often burdened by the difficulty of determining the forces that control the dynamics of the individual elements within each system. Accessing this local information is key to understanding the physics dominating the system and to create the models that can explain the observed collective phenomena. In this work, we present a machine-learning model, a graph neural network, that uses the collective movement of the system to learn the active and two-body forces controlling the individual dynamics of the particles. We verify our approach using numerical simulations of active brownian particles, considering different interaction potentials and levels of activity. Finally, we apply our model to experiments of electrophoretic Janus particles, extracting the active and two-body forces that control the dynamics of the colloids. Due to this, we can uncover the physics dominating the behavior of the system. We extract an active force that depends on the electric field and also area fraction. We also discover a dependence of the two-body interaction with the electric field that leads us to propose that the dominant force between these colloids is a screened electrostatic interaction with a constant length scale. We expect that this methodology can open a new avenue for the study and modeling of experimental systems of active particles.

cond-mat.soft

Test set for materials science and engineering with user-friendly graphic tools for error analysis: Systematic benchmark of the numerical and intrinsic errors in state-of-the-art electronic-structure approximations

Understanding the applicability and limitations of electronic-structure methods needs careful and efficient comparison with accurate reference data. Knowledge of the quality and errors of electronic-structure calculations is crucial to advanced method development, high-throughput computations, and data analyses. In this paper, we present a test set for computational materials science and engineering (MSE), that aims to provide accurate and easily accessible crystal properties for a hierarchy of exchange-correlation approximations, ranging from the well-established mean-field approximations to the state-of-the-art methods of many-body perturbation theory. We consider cohesive energy, lattice constant and bulk modulus as representatives for the first- and second-row elements and their binaries with cubic crystal structures and various bonding characters. A strong effort is made to push the borders of numerical accuracy for cohesive properties as calculated using the local-density approximation (LDA), several generalized gradient approximations (GGAs), meta-GGAs and hybrids in \textit{all-electron} resolution, and the second-order M\o{}ller-Plesset perturbation theory (MP2) and the random-phase approximation (RPA) with frozen-core approximation based on \textit{all-electron} Hartree-Fock, PBE and/or PBE0 references. This results in over 10,000 calculations, which record a comprehensive convergence test with respect to numerical parameters for a wide range of electronic structure methods within the numerical atom-centered orbital framework. As an indispensable part of the MSE test set, a web site is established \href{http://mse.fhi-berlin.mpg.de}{\texttt{http://mse.fhi-berlin.mpg.de}}. This not only allows for easy access to all reference data but also provides user-friendly graphical tools for post-processing error analysis.

cond-mat.mtrl-sci