Search arXivSearch

arXiv subjects

E. Calore

Publications and source records attributed to E. Calore.

18 recordsLinked to original sources

gCAMB: A GPU-accelerated Boltzmann solver for next-generation cosmological surveys

Inferring cosmological parameters from Cosmic Microwave Background (CMB) data requires repeated and computationally expensive calculations of theoretical angular power spectra using Boltzmann solvers like CAMB. This creates a significant bottleneck, particularly for non-standard cosmological models and the high-accuracy demands of future surveys. While emulators based on deep neural networks can accelerate this process by several orders of magnitude, they first require large, pre-computed training datasets, which are costly to generate and model-specific. To address this challenge, we introduce gCAMB, a version of the CAMB code ported to GPUs, which preserves all the features of the original CPU-only code. By offloading the most computationally intensive modules to the GPU, gCAMB significantly accelerates the generation of power spectra, saving massive computational time, halving the power consumption in high-accuracy settings and, among other purposes, facilitating the creation of extensive training sets needed for robust cosmological analyses. We make the gCAMB software available to the community at https://github.com/lstorchi/CAMB/tree/gpuport.

astro-ph.CO

Quantifying memory in spin glasses

Rejuvenation and memory, long considered the distinguishing features of spin glasses, have recently been proven to result from the growth of multiple length scales. This insight, enabled by simulations on the Janus~II supercomputer, has opened the door to a quantitative analysis. We combine numerical simulations with comparable experiments to introduce two coefficients that quantify memory. A third coefficient has been recently presented by Freedberg et al. We show that these coefficients are physically equivalent by studying their temperature and waiting-time dependence.

cond-mat.dis-nn

Multifractality in spin glasses

We unveil the multifractal behavior of Ising spin glasses in their low-temperature phase. Using the Janus II custom-built supercomputer, the spin-glass correlation function is studied locally. Dramatic fluctuations are found when pairs of sites at the same distance are compared. The scaling of these fluctuations, as the spin-glass coherence length grows with time, is characterized through the computation of the singularity spectrum and its corresponding Legendre transform. A comparatively small number of site pairs controls the average correlation that governs the response to a magnetic field. We explain how this scenario of dramatic fluctuations (at length scales smaller than the coherence length) can be reconciled with the smooth, self-averaging behavior that has long been considered to describe spin-glass dynamics.

cond-mat.dis-nn

On the superposition principle and non-linear response in spin glasses

The extended principle of superposition has been a touchstone of spin glass dynamics for almost thirty years. The Uppsala group has demonstrated its validity for the metallic spin glass, CuMn, for magnetic fields $H$ up to 10 Oe at the reduced temperature $T_\mathrm{r}=T/T_\mathrm{g} = 0.95$, where $T_\mathrm{g}$ is the spin glass condensation temperature. For $H > 10$ Oe, they observe a departure from linear response which they ascribe to the development of non-linear dynamics. The thrust of this paper is to develop a microscopic origin for this behavior by focusing on the time development of the spin glass correlation length, $\xi(t,t_\mathrm{w};H)$. Here, $t$ is the time after $H$ changes, and $t_\mathrm{w}$ is the time from the quench for $T>T_\mathrm{g}$ to the working temperature $T$ until $H$ changes. We connect the growth of $\xi(t,t_\mathrm{w};H)$ to the barrier heights $\Delta(t_\mathrm{w})$ that set the dynamics. The effect of $H$ on the magnitude of $\Delta(t_\mathrm{w})$ is responsible for affecting differently the two dynamical protocols associated with turning $H$ off (TRM, or thermoremanent magnetization) or on (ZFC, or zero field-cooled magnetization). In this paper, we display the difference between the zero-field cooled $\xi_{\text {ZFC}}(t,t_\mathrm{w};H)$ and the thermoremanent magnetization $\xi_{\text {TRM}}(t,t_\mathrm{w};H)$ correlation lengths as $H$ increases, both experimentally and through numerical simulations, corresponding to the violation of the extended principle of superposition in line with the finding of the Uppsala Group.

cond-mat.dis-nn

Memory and rejuvenation in spin glasses: aging systems are ruled by more than one length scale

Memory and rejuvenation effects in the magnetic response of off-equilibrium spin glasses have been widely regarded as the doorway into the experimental exploration of ultrametricity and temperature chaos (maybe the most exotic features in glassy free-energy landscapes). Unfortunately, despite more than twenty years of theoretical efforts following the experimental discovery of memory and rejuvenation, these effects have thus far been impossible to simulate reliably. Yet, three recent developments convinced us to accept this challenge: first, the custom-built Janus II supercomputer makes it possible to carry out "numerical experiments" in which the very same quantities that can be measured in single crystals of CuMn are computed from the simulation, allowing for parallel analysis of the simulation/experiment data. Second, Janus II simulations have taught us how numerical and experimental length scales should be compared. Third, we have recently understood how temperature chaos materializes in aging dynamics. All three aspects have proved crucial for reliably reproducing rejuvenation and memory effects on the computer. Our analysis shows that (at least) three different length scales play a key role in aging dynamics, while essentially all theoretical analyses of the aging dynamics emphasize the presence and the crucial role of a single glassy correlation length.

cond-mat.dis-nn

Spin-glass dynamics in the presence of a magnetic field: exploration of microscopic properties

The synergy between experiment, theory, and simulations enables a microscopic analysis of spin-glass dynamics in a magnetic field in the vicinity of and below the spin-glass transition temperature $T_\mathrm{g}$. The spin-glass correlation length, $\xi(t,t_\mathrm{w};T)$, is analysed both in experiments and in simulations in terms of the waiting time $t_\mathrm{w}$ after the spin glass has been cooled down to a stabilised measuring temperature $T<T_\mathrm{g}$ and of the time $t$ after the magnetic field is changed. This correlation length is extracted experimentally for a CuMn 6 at. % single crystal, as well as for simulations on the Janus II special-purpose supercomputer, the latter with time and length scales comparable to experiment. The non-linear magnetic susceptibility is reported from experiment and simulations, using $\xi(t,t_\mathrm{w};T)$ as the scaling variable. Previous experiments are reanalysed, and disagreements about the nature of the Zeeman energy are resolved. The growth of the spin-glass magnetisation in zero-field magnetisation experiments, $M_\mathrm{ZFC}(t,t_\mathrm{w};T)$, is measured from simulations, verifying the scaling relationships in the dynamical or non-equilibrium regime. Our preliminary search for the de Almeida-Thouless line in $D=3$ is discussed.

cond-mat.dis-nn

Temperature chaos is present in off-equilibrium spin-glass dynamics

We find a dynamic effect in the non-equilibrium dynamics of a spin glass that closely parallels equilibrium temperature chaos. This effect, that we name dynamic temperature chaos, is spatially heterogeneous to a large degree. The key controlling quantity is the time-growing spin-glass coherence length. Our detailed characterization of dynamic temperature chaos paves the way for the analysis of recent and forthcoming experiments. This work has been made possible thanks to the most massive simulation to date of non-equilibrium dynamics, carried out on the Janus~II custom-built supercomputer.

cond-mat.dis-nn

Scaling law describes the spin-glass response in theory, experiments and simulations

The correlation length $\xi$, a key quantity in glassy dynamics, can now be precisely measured for spin glasses both in experiments and in simulations. However, known analysis methods lead to discrepancies either for large external fields or close to the glass temperature. We solve this problem by introducing a scaling law that takes into account both the magnetic field and the time-dependent spin-glass correlation length. The scaling law is successfully tested against experimental measurements in a CuMn single crystal and against large-scale simulations on the Janus II dedicated computer.

cond-mat.stat-mech

The Mpemba effect in spin glasses is a persistent memory effect

The Mpemba effect occurs when a hot system cools faster than an initially colder one, when both are refrigerated in the same thermal reservoir. Using the custom built supercomputer Janus II, we study the Mpemba effect in spin glasses and show that it is a non-equilibrium process, governed by the coherence length \xi of the system. The effect occurs when the bath temperature lies in the glassy phase, but it is not necessary for the thermal protocol to cross the critical temperature. In fact, the Mpemba effect follows from a strong relationship between the internal energy and \xi that turns out to be a sure-tell sign of being in the glassy phase. Thus, the Mpemba effect presents itself as an intriguing new avenue for the experimental study of the coherence length in supercooled liquids and other glass formers.

cond-mat.dis-nn

Energy-efficiency evaluation of Intel KNL for HPC workloads

Energy consumption is increasingly becoming a limiting factor to the design of faster large-scale parallel systems, and development of energy-efficient and energy-aware applications is today a relevant issue for HPC code-developer communities. In this work we focus on energy performance of the Knights Landing (KNL) Xeon Phi, the latest many-core architecture processor introduced by Intel into the HPC market. We take into account the 64-core Xeon Phi 7230, and analyze its energy performance using both the on-chip MCDRAM and the regular DDR4 system memory as main storage for the application data-domain. As a benchmark application we use a Lattice Boltzmann code heavily optimized for this architecture and implemented using different memory data layouts to store its lattice. We assessthen the energy consumption using different memory data-layouts, kind of memory (DDR4 or MCDRAM) and number of threads per core.

cs.DC

Aging rate of spin glasses from simulations matches experiments

Experiments on spin glasses can now make precise measurements of the exponent $z(T)$ governing the growth of glassy domains, while our computational capabilities allow us to make quantitative predictions for experimental scales. However, experimental and numerical values for $z(T)$ have differed. We use new simulations on the Janus II computer to resolve this discrepancy, finding a time-dependent $z(T, t_w)$, which leads to the experimental value through mild extrapolations. Furthermore, theoretical insight is gained by studying a crossover between the $T = T_c$ and $T = 0$ fixed points.

cond-mat.dis-nn

Matching microscopic and macroscopic responses in glasses

We first reproduce on the Janus and Janus II computers a milestone experiment that measures the spin-glass coherence length through the lowering of free-energy barriers induced by the Zeeman effect. Secondly we determine the scaling behavior that allows a quantitative analysis of a new experiment reported in the companion Letter [S. Guchhait and R. Orbach, Phys. Rev. Lett. 118, 157203 (2017)]. The value of the coherence length estimated through the analysis of microscopic correlation functions turns out to be quantitatively consistent with its measurement through macroscopic response functions. Further, non-linear susceptibilities, recently measured in glass-forming liquids, scale as powers of the same microscopic length.

cond-mat.dis-nn

Optimization of Lattice Boltzmann Simulations on Heterogeneous Computers

High-performance computing systems are more and more often based on accelerators. Computing applications targeting those systems often follow a host-driven approach in which hosts offload almost all compute-intensive sections of the code onto accelerators; this approach only marginally exploits the computational resources available on the host CPUs, limiting performance and energy efficiency. The obvious step forward is to run compute-intensive kernels in a concurrent and balanced way on both hosts and accelerators. In this paper we consider exactly this problem for a class of applications based on Lattice Boltzmann Methods, widely used in computational fluid-dynamics. Our goal is to develop just one program, portable and able to run efficiently on several different combinations of hosts and accelerators. To reach this goal, we define common data layouts enabling the code to exploit efficiently the different parallel and vector options of the various accelerators, and matching the possibly different requirements of the compute-bound and memory-bound kernels of the application. We also define models and metrics that predict the best partitioning of workloads among host and accelerator, and the optimally achievable overall performance level. We test the performance of our codes and their scaling properties using as testbeds HPC clusters incorporating different accelerators: Intel Xeon-Phi many-core processors, NVIDIA GPUs and AMD GPUs.

cs.DC

Performance and Portability of Accelerated Lattice Boltzmann Applications with OpenACC

An increasingly large number of HPC systems rely on heterogeneous architectures combining traditional multi-core CPUs with power efficient accelerators. Designing efficient applications for these systems has been troublesome in the past as accelerators could usually be programmed using specific programming languages threatening maintainability, portability and correctness. Several new programming environments try to tackle this problem. Among them, OpenACC offers a high-level approach based on compiler directive clauses to mark regions of existing C, C++ or Fortran codes to run on accelerators. This approach directly addresses code portability, leaving to compilers the support of each different accelerator, but one has to carefully assess the relative costs of portable approaches versus computing efficiency. In this paper we address precisely this issue, using as a test-bench a massively parallel Lattice Boltzmann algorithm. We first describe our multi-node implementation and optimization of the algorithm, using OpenACC and MPI. We then benchmark the code on a variety of processors, including traditional CPUs and GPUs, and make accurate performance comparisons with other GPU implementations of the same algorithm using CUDA and OpenCL. We also asses the performance impact associated to portable programming, and the actual portability and performance-portability of OpenACC-based applications across several state-of-the- art architectures.

cs.DC

Massively parallel lattice-Boltzmann codes on large GPU clusters

This paper describes a massively parallel code for a state-of-the art thermal lattice- Boltzmann method. Our code has been carefully optimized for performance on one GPU and to have a good scaling behavior extending to a large number of GPUs. Versions of this code have been already used for large-scale studies of convective turbulence. GPUs are becoming increasingly popular in HPC applications, as they are able to deliver higher performance than traditional processors. Writing efficient programs for large clusters is not an easy task as codes must adapt to increasingly parallel architectures, and the overheads of node-to-node communications must be properly handled. We describe the structure of our code, discussing several key design choices that were guided by theoretical models of performance and experimental benchmarks. We present an extensive set of performance measurements and identify the corresponding main bot- tlenecks; finally we compare the results of our GPU code with those measured on other currently available high performance processors. Our results are a production-grade code able to deliver a sustained performance of several tens of Tflops as well as a design and op- timization methodology that can be used for the development of other high performance applications for computational physics.

cs.DC

A statics-dynamics equivalence through the fluctuation-dissipation ratio provides a window into the spin-glass phase from nonequilibrium measurements

The unifying feature of glass formers (such as polymers, supercooled liquids, colloids, granulars, spin glasses, superconductors, ...) is a sluggish dynamics at low temperatures. Indeed, their dynamics is so slow that thermal equilibrium is never reached in macroscopic samples: in analogy with living beings, glasses are said to age. Here, we show how to relate experimentally relevant quantities with the experimentally unreachable low-temperature equilibrium phase. We have performed a very accurate computation of the non-equilibrium fluctuation-dissipation ratio for the three-dimensional Edwards-Anderson Ising spin glass, by means of large-scale simulations on the special-purpose computers Janus and Janus II. This ratio (computed for finite times on very large, effectively infinite, systems) is compared with the equilibrium probability distribution of the spin overlap for finite sizes. The resulting quantitative statics-dynamics dictionary, based on observables that can be measured with current experimental methods, could allow the experimental exploration of important features of the spin-glass phase without uncontrollable extrapolations to infinite times or system sizes.

cond-mat.dis-nn

High-spin structure in $^{40}$K

High-spin states of $^{40}$K have been populated in the fusion-evaporation reaction $^{12}$C($^{30}$Si,np)$^{40}$K and studied by means of $\gamma$-ray spectroscopy techniques using one AGATA triple cluster detector, at INFN - Laboratori Nazionali di Legnaro. Several new states with excitation energy up to 8 MeV and spin up to $10^-$ have been discovered. These new states are discussed in terms of J=3 and T=0 neutron-proton hole pairs. Shell-model calculations in a large model space have shown a good agreement with the experimental data for most of the energy levels. The evolution of the structure of this nucleus is here studied as a function of excitation energy and angular momentum.

nucl-ex

AGATA - Advanced Gamma Tracking Array

The Advanced GAmma Tracking Array (AGATA) is a European project to develop and operate the next generation gamma-ray spectrometer. AGATA is based on the technique of gamma-ray energy tracking in electrically segmented high-purity germanium crystals. This technique requires the accurate determination of the energy, time and position of every interaction as a gamma ray deposits its energy within the detector volume. Reconstruction of the full interaction path results in a detector with very high efficiency and excellent spectral response. The realization of gamma-ray tracking and AGATA is a result of many technical advances. These include the development of encapsulated highly-segmented germanium detectors assembled in a triple cluster detector cryostat, an electronics system with fast digital sampling and a data acquisition system to process the data at a high rate. The full characterization of the crystals was measured and compared with detector-response simulations. This enabled pulse-shape analysis algorithms, to extract energy, time and position, to be employed. In addition, tracking algorithms for event reconstruction were developed. The first phase of AGATA is now complete and operational in its first physics campaign. In the future AGATA will be moved between laboratories in Europe and operated in a series of campaigns to take advantage of the different beams and facilities available to maximize its science output. The paper reviews all the achievements made in the AGATA project including all the necessary infrastructure to operate and support the spectrometer.

physics.ins-det