Search arXiv⌕ Search

arXiv subjects

Raffaello Potestio

Publications and source records attributed to Raffaello Potestio.

At least 19 recordsLinked to original sources

Coarse-Graining Hidden Representations: Unsupervised Neuron Selection via Mapping Entropy

Overparameterized neural networks carry far more hidden units than a task nominally requires, raising the question of which neurons are essential and whether that distinction is legible in the representation itself, without labels or gradients. We cast neuron selection as the problem of coarse-graining the hidden layer by retaining a subset of its neurons, and score each putative selection by the mapping entropy (ME). This quantity measures the loss of discriminatory power inherent in discarding part of the network neurons, and the selection that minimises the ME is taken as particularly informative. This criterion is fully unsupervised, in that it depends only on hidden-activation statistics. In teacher-student networks, ME optimisation recovers the minimal teacher-consistent representation and retains extra units in proportion to the hidden layer's residual variability; in a non-linear Gaussian process task, it selects coherent functional-class mappings whose preferred class shifts across training. On this task and on translation-augmented MNIST, ME-selected subnetworks outperform random subsets of equal size, most clearly under strong compression - linking configurational distinguishability to predictive performance.

cs.LG↗

The bliss of dimensionality: how an unsupervised criterion identifies optimal low-resolution representations of high-dimensional datasets

Selecting the optimal resolution for discretizing high-dimensional data is a central problem in physics and data analysis, particularly in unsupervised settings where the underlying distribution is unknown. The Relevance-Resolution (Res-Rel) framework addresses this issue through an information-theoretic trade-off between descriptive detail and statistical reliability. Here we provide a systematic validation of this approach by comparing its characteristic optima--maximum relevance and the -1 slope (information-theoretic) point--with the discretization that minimizes the Kullback-Leibler divergence from a known or physically motivated ground truth distribution. Across unstructured and structured synthetic datasets, Gaussian clones of MNIST, and molecular dynamics simulations of the alanine dipeptide, we find that as the dimensionality or informative content increases the KL-optimal discretization consistently lies within the Res-Rel optimality region. Furthermore, in high-dimensional regimes the -1 slope criterion closely matches the KL divergence minimum. These results establish the quantitative consistency of unsupervised information-theoretic selection with distribution-based optimality.

cond-mat.stat-mech↗

NET4EXA: Pioneering the Future of Interconnects for Supercomputing and AI

NET4EXA aims to develop a next-generation high-performance interconnect for HPC and AI systems, addressing the increasing demands of large-scale infrastructures, such as those required for training Large Language Models. Building upon the proven BXI (Bull eXascale Interconnect) European technology used in TOP15 supercomputers, NET4EXA will deliver the new BXI release, BXIv3, a complete hardware and software interconnect solution, including switch and network interface components. The project will integrate a fully functional pilot system at TRL 8, ready for deployment into upcoming exascale and post-exascale systems from 2025 onward. Leveraging prior research from European initiatives like RED-SEA, the previous achievements of consortium partners and over 20 years of expertise from BULL, NET4EXA also lays the groundwork for the future generation of BXI, BXIv4, providing analysis and preliminary design. The project will use a hybrid development and co-design approach, combining commercial switch technology with custom IP and FPGA-based NICs. Performances of NET4EXA BXIv3 interconnect will be evaluated using a broad portfolio of benchmarks, scientific scalable applications, and AI workloads.

cs.NI↗

Artificial life of an active droplets system: a quantitative lifecycle analysis

The study of synthetic active matter systems holds the promise for designing smart materials and devices with emergent characteristics akin to those of living organisms, eventually opening the doors to the realization of artificial life. Such an investigation, however, is challenged by the difficulty inherent in identifying the relationship between the features of the elementary constituents and the emergent properties of the whole; to this end, a key step consists in the accurate quantification of the system's observed behavior. Here, we report the study of 50 self-propelled oil droplets floating on the surface of an aqueous solution. 25 droplets are stained with a red dye, and the other 25 are stained blue: the colorants affect the droplets' interfacial tension properties differently, consequently influencing their collective dynamics. Droplet trajectories extending for up to 5 hours are extracted from video recordings with a tracking pipeline developed ad hoc. The structural and dynamical analysis of the system reveals a ``life-to-death'' cycle unfolding in qualitatively distinct stages, showcasing a complex interplay between individual droplet mobility and collective organization. The tools developed and the results obtained in our work pave the way to the in silico modelling as well as the experimental design of synthetic active matter systems displaying life-like and programmable behavior.

cond-mat.soft↗

Determining the optimal structural resolution of proteins through an information-theoretic analysis of their conformational ensemble

The choice of structural resolution is a fundamental aspect of protein modelling, determining the balance between descriptive power and interpretability. Although atomistic simulations provide maximal detail, much of this information is redundant to understand the relevant large-scale motions and conformational states. Here, we introduce an unsupervised, information-theoretic framework that determines the minimal number of atoms required to retain a maximally informative description of the configurational space sampled by a protein. This framework quantifies the informativeness of coarse-grained representations obtained by systematically decimating atomic degrees of freedom and evaluating the resulting clustering of sampled conformations. Application to molecular dynamics trajectories of dynamically diverse proteins shows that the optimal number of retained atoms scales linearly with system size, averaging about four heavy atoms per residue--remarkably consistent with the resolution of well-established coarse-grained models, such as MARTINI and SIRAH. Furthermore, the analysis shows that the optimal retained atoms number depends not only on molecular size but also on the extent of conformational exploration, decreasing for systems dominated by collective motions. The proposed method establishes a general criterion to identify the minimal structural detail that preserves the essential configurational information, thereby offering a new viewpoint on the structure-dynamics-function relationship in proteins and guiding the construction of parsimonious yet informative multiscale models.

q-bio.BM↗

Low-resolution descriptions of model neural activity reveal hidden features and underlying system properties

The analysis of complex systems such as neural networks is made particularly difficult by the overwhelming number of their interacting components. In the absence of prior knowledge, identifying a small but informative subset of network nodes on which the analysis should focus is a rather challenging task. In this work, we address this problem in the context of a Hopfield model, which is observed through the lenses of low-resolution representations, or decimation mappings, consisting of subgroups of its neurons. The optimal, most informative mappings of the network are defined through a recently developed methodology, the mapping entropy optimisation workflow (MEOW), which performs an unsupervised analysis of the states sampled by the network and identifies those subgroups of spins whose configuration distribution is closest to that of the full, high-resolution model. Which neurons are retained in an optimal mapping is found to critically depend on the properties of the interaction matrix of the network and the level of detail employed to describe the system; by these means, it is thus possible to extract quantitative insight about the underlying properties of the high-resolution model through the analysis of its optimal low-resolution representations. These results show a tight and potentially fruitful relation between the level of detail at which the network is inspected and the type and amount of information that can be gathered from it, and showcase the MEOW approach as a practical, enabling tool for the study of complex systems.

cond-mat.dis-nn↗

Density of states in neural networks: an in-depth exploration of learning in parameter space

Learning in neural networks critically hinges on the intricate geometry of the loss landscape associated with a given task. Traditionally, most research has focused on finding specific weight configurations that minimize the loss. In this work, born from the cross-fertilization of machine learning and theoretical soft matter physics, we introduce a novel, computationally efficient approach to examine the weight space across all loss values. Employing the Wang-Landau enhanced sampling algorithm, we explore the neural network density of states - the number of network parameter configurations that produce a given loss value - and analyze how it depends on specific features of the training set. Using both real-world and synthetic data, we quantitatively elucidate the relation between data structure and network density of states across different sizes and depths of binary-state networks.

cond-mat.stat-mech↗

A multi-scale analysis of the CzrA transcription repressor highlights the allosteric changes induced by metal ion binding

Allosteric regulation is a widespread strategy employed by several proteins to transduce chemical signals and perform biological functions. Metal sensor proteins are exemplary in this respect, e.g., in that they selectively bind and unbind DNA depending on the state of a distal ion coordination site. In this work, we carry out an investigation of the structural and mechanical properties of the CzrA transcription repressor through the analysis of microsecond-long molecular dynamics (MD) trajectories; the latter are processed through the mapping entropy optimisation workflow (MEOW), a recently developed information-theoretical method that highlights, in an unsupervised manner, residues of particular mechanical, functional, and biological importance. This approach allows us to unveil how differences in the properties of the molecule are controlled by the state of the zinc coordination site, with particular attention to the DNA binding region. These changes correlate with a redistribution of the conformational variability of the residues throughout the molecule, in spite of an overall consistency of its architecture in the two (ion-bound and free) coordination states. The results of this work corroborate previous studies, provide novel insight into the fine details of the mechanics of CzrA, and showcase the MEOW approach as a novel instrument for the study of allosteric regulation and other processes in proteins through the analysis of plain MD simulations.

q-bio.BM↗

Molecular dynamics characterization of the free and encapsidated RNA2 of CCMV with the oxRNA model

The cowpea chlorotic mottle virus (CCMV) has emerged as an exemplary model system to assess the balance between electrostatic and topological features of ssRNA viruses, specifically in the context of the viral self-assembly process. Yet, in spite of its biophysical significance, little structural data of the RNA content of the CCMV virion is currently available. Here, the conformational dynamics of the RNA2 fragment of CCMV was assessed via coarse-grained molecular dynamics simulations, employing the oxRNA2 model. The behavior of RNA2 has been characterized both as a freely-folding molecule and within a mean-field depiction of a CCMV-like capsid. For the latter, a multi-scale approach was employed, to derive a radial potential profile of the viral cavity, from atomistic structures of the CCMV capsid in solution. The conformational ensembles of the encapsidated RNA2 were significantly altered with respect to the freely-folding counterparts, as shown by the emergence of long-range motifs and pseudoknots in the former case. Finally, the role of the N-terminal tails of the CCMV subunits (and ionic shells thereof) is highlighted as a critical feature in the construction of a proper electrostatic model of the CCMV capsid.

cond-mat.soft↗

EXCOGITO, an extensible coarse-graining toolbox for the investigation of biomolecules by means of low-resolution representation

Bottom-up coarse-grained (CG) models proved to be essential to complement and sometimes even replace all-atom representations of soft matter systems and biological macromolecules. The development of low-resolution models takes the moves from the reduction of the degrees of freedom employed, that is, the definition of a mapping between a system's high-resolution description and its simplified counterpart. Even in the absence of an explicit parametrisation and simulation of a CG model, the observation of the atomistic system in simpler terms can be informative: this idea is leveraged by the mapping entropy, a measure of the information loss inherent to the process of coarsening. Mapping entropy lies at the heart of the extensible coarse-graining toolbox, or EXCOGITO, developed to perform a number of operations and analyses on molecular systems pivoting around the properties of mappings. EXCOGITO can process an all-atom trajectory to compute the mapping entropy, identify the mapping that minimizes it, and establish quantitative relations between a low-resolution representation and the geometrical, structural, and energetic features of the system. Here, the software, which is available free of charge under an open-source licence, is presented and showcased to introduce potential users to its capabilities and usage. Published on the J. Chem. Inf. Model. on June 11, 2024. DOI: https://doi.org/10.1021/acs.jcim.4c00490

cond-mat.soft↗

Emergent circulation patterns from anonymized mobility data: Clustering Italy in the time of Covid

Using anonymized mobility data from Facebook users and publicly available information on the Italian population, we model the circulation of people in Italy before and during the early phase of the SARS-CoV-2 pandemic (COVID-19). We perform a spatial and temporal clustering of the movement network at the level of fluxes across provinces on a daily basis. The resulting partition in time successfully identifies the first two lockdowns without any prior information. Similarly, the spatial clustering returns 11 to 23 clusters depending on the period ("standard" mobility vs. lockdown) using the greedy modularity communities clustering method, and 16 to 30 clusters using the critical variable selection method. Fascinatingly, the spatial clusters obtained with both methods are strongly reminiscent of the 11 regions into which emperor Augustus had divided Italy according to Pliny the Elder. This work introduces and validates a data analysis pipeline that enables us: i) to assess the reliability of data obtained from a partial and potentially biased sample of the population in performing estimates of population mobility nationwide; ii) to identify areas of a Country with well-defined mobility patterns, and iii) to distinguish different patterns from one another, resolve them in time and find their optimal spatial extent. The proposed method is generic and can be applied to other countries, with different geographical scales, and also to similar networks (e.g. biological networks). The results can thus represent a relevant step forward in the development of methods and strategies for the containment of future epidemic phenomena.

physics.soc-ph↗

Fast, accurate, and system-specific variable-resolution modelling of proteins

In recent years, a few multiple-resolution modelling strategies have been proposed, in which functionally relevant parts of a biomolecule are described with atomistic resolution, while the remainder of the system is concurrently treated using a coarse-grained model. In most cases, the parametrisation of the latter requires lengthy reference all-atom simulations and/or the usage of off-shelf coarse-grained force fields, whose interactions have to be refined to fit the specific system under examination. Here, we overcome these limitations through a novel multi-resolution modelling scheme for proteins, dubbed coarse-grained anisotropic network model for variable resolution simulations, or CANVAS. This scheme enables the user-defined modulation of the resolution level throughout the system structure; a fast parametrisation of the potential without the necessity of reference simulations; and the straightforward usage of the model on the most commonly used molecular dynamics platforms. The method is presented and validated on two case studies, the enzyme adenylate kinase and the therapeutic antibody pembrolizumab, by comparing results obtained with the CANVAS model against fully atomistic simulations. The modelling software, implemented in python, is made freely available for the community on a collaborative github repository.

cond-mat.soft↗

Coarse-grained Mori-Zwanzig dynamics in a time-non-local stationary-action framework

Coarse-grained (CG) models are simplified representations of soft matter systems that are commonly employed to overcome size and time limitations in computational studies. Many approaches have been developed to construct and parametrise such effective models for a variety of systems of natural as well as artificial origin. However, while extremely accurate in reproducing the stationary and equilibrium observables obtained with more detailed representations, CG models generally fail to preserve the original time scales of the reference system, and hence its dynamical properties. In order to improve our understanding of the impact of coarse-graining on the model system dynamics, we here formulate the Mori-Zwanzig generalised Langevin equations (GLEs) of motion of a CG model in terms of a time non-local stationary-action principle. The latter is employed in combination with a data-driven optimisation strategy to determine the parameters of the GLE. We apply this approach to a system of water molecules in standard thermodynamical conditions, showing that it can substantially improve the dynamical features of the corresponding CG model.

cond-mat.stat-mech↗

Making sense of complex systems through resolution, relevance, and mapping entropy

Complex systems are characterised by a tight, nontrivial interplay of their constituents, which gives rise to a multi-scale spectrum of emergent properties. In this scenario, it is practically and conceptually difficult to identify those degrees of freedom that mostly determine the behaviour of the system and separate them from less prominent players. Here, we tackle this problem making use of three measures of statistical information: resolution, relevance, and mapping entropy. We address the links existing among them, taking the moves from the established relation between resolution and relevance and further developing novel connections between resolution and mapping entropy; by these means we can identify, in a quantitative manner, the number and selection of degrees of freedom of the system that preserve the largest information content about the generative process that underlies an empirical dataset. The method, which is implemented in a freely available software, is fully general, as it is shown through the application to three very diverse systems, namely a toy model of independent binary spins, a coarse-grained representation of the financial stock market, and a fully atomistic simulation of a protein.

cond-mat.stat-mech↗

Information-theoretical measures identify accurate low-resolution representations of protein configurational space

A steadily growing computational power is employed to perform molecular dynamics simulations of biological macromolecules, which represents at the same time an immense opportunity and a formidable challenge. In fact, large amounts of data are produced, from which useful, synthetic, and intelligible information has to be extracted to make the crucial step from knowing to understanding. Here we tackled the problem of coarsening the conformational space sampled by proteins in the course of molecular dynamics simulations. We applied different schemes to cluster the frames of a dataset of protein simulations; we then employed an information-theoretical framework, based on the notion of resolution and relevance, to gauge how well the various clustering methods accomplish this simplification of the configurational space. Our approach allowed us to identify the level of resolution that optimally balances simplicity and informativeness; furthermore, we found that the most physically accurate clustering procedures are those that induce an ultrametric structure of the low-resolution space, consistently with the hypothesis that the protein conformational landscape has a self-similar organisation. The proposed strategy is general and its applicability extends beyond that of computational biophysics, making it a valuable tool to extract useful information from large datasets.

cond-mat.soft↗

A mapping space Odyssey: characterising the statistical and metric properties of reduced representations of macromolecules

Simplified representations of macromolecules help in rationalising and understanding the outcome of atomistic simulations, and serve to the construction of effective, coarse-grained models. The number and distribution of coarse-grained sites bears a strict relation with the amount of information conveyed by the representation and the accuracy of the associated effective model; in this work, we investigate this relationship from the very basics: specifically, we propose a rigorous notion of scalar product among mappings, which implies a distance and a metric space of simplified representations. Making use of a Wang-Landau enhanced sampling algorithm, we exhaustively explore the space of mappings, quantifying their qualitative features in terms of their squared norm and relating them with thermodynamical properties of the underlying macromolecule. A one-to-one correspondence with an interacting lattice gas on a finite volume leads to the emergence of discontinuous phase transitions in mapping space that mark the boundaries between qualitatively different representations of the same molecule.

cond-mat.soft↗

Accelerating the identification of informative reduced representations of proteins with deep learning for graphs

The limits of molecular dynamics (MD) simulations of macromolecules are steadily pushed forward by the relentless developments of computer architectures and algorithms. This explosion in the number and extent (in size and time) of MD trajectories induces the need of automated and transferable methods to rationalise the raw data and make quantitative sense out of them. Recently, an algorithmic approach was developed by some of us to identify the subset of a protein's atoms, or mapping, that enables the most informative description of it. This method relies on the computation, for a given reduced representation, of the associated mapping entropy, that is, a measure of the information loss due to the simplification. Albeit relatively straightforward, this calculation can be time consuming. Here, we describe the implementation of a deep learning approach aimed at accelerating the calculation of the mapping entropy. The method relies on deep graph networks, which provide extreme flexibility in the input format. We show that deep graph networks are accurate and remarkably efficient, with a speedup factor as large as $10^5$ with respect to the algorithmic computation of the mapping entropy. Applications of this method, which entails a great potential in the study of biomolecules when used to reconstruct its mapping entropy landscape, reach much farther than this, being the scheme easily transferable to the computation of arbitrary functions of a molecule's structure.

physics.comp-ph↗

An information theory-based approach for optimal model reduction of biomolecules

In the theoretical modelling of a physical system a crucial step consists in the identification of those degrees of freedom that enable a synthetic, yet informative representation of it. While in some cases this selection can be carried out on the basis of intuition and experience, a straightforward discrimination of the important features from the negligible ones is difficult for many complex systems, most notably heteropolymers and large biomolecules. We here present a thermodynamics-based theoretical framework to gauge the effectiveness of a given simplified representation by measuring its information content. We employ this method to identify those reduced descriptions of proteins, in terms of a subset of their atoms, that retain the largest amount of information from the original model; we show that these highly informative representations share common features that are intrinsically related to the biological properties of the proteins under examination, thereby establishing a bridge between protein structure, energetics, and function.

cond-mat.stat-mech↗