Search arXivSearch

arXiv · 2604.06035

cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations

Abstract

We present cuRAMSES, a suite of advanced domain decomposition strategies and algorithmic optimizations for the ramses adaptive mesh refinement (AMR) code, designed to overcome the communication, memory, and solver bottlenecks inherent in massive cosmological simulations. The central innovation is a recursive k-section domain decomposition that replaces the traditional Hilbert curve ordering with a hierarchical spatial partitioning. This approach substitutes global all-to-all communications with neighbour-only point-to-point communications. By maintaining a constant number of communication partners regardless of the total rank count, it significantly improves strong scaling at high concurrency. To address critical memory constraints at scale, we introduce a Morton-key hash table for octree-neighbour lookup alongside on-demand array allocation, drastically reducing the per-rank memory footprint. Furthermore, a novel spatial hash-binning algorithm in box-type local domains accelerates supernova and AGN feedback routines by over two orders of magnitude (an about 260 times speedup). For hybrid architectures, an automatic CPU/GPU dispatch model with GPU-resident mesh data is implemented and benchmarked. The multigrid Poisson solver achieves a 1.7 times GPU speedup on H100 and A100 GPUs, although the Godunov solver is currently PCIe-bandwidth-limited. The net improvement is about 20 per cent on current PCIe-connected hardware, and a performance model predicts about 2 times on tightly coupled architectures such as the NVIDIA GH200. Additionally, a variable-Nrank restart capability enables flexible I/O workflows. Extensive diagnostics verify that all modifications preserve mass, momentum, and energy conservation, matching the reference Hilbert-ordering run to within 0.5 per cent in the total energy diagnostic.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Juhan Kim. 2026-06-02. cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations. https://arxiv.org/abs/2604.06035

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Small hosts, big appetites: unveiling rapid and early low-mass black hole growth in cosmological zoom-in simulations of dwarf galaxies

Dwarf galaxies are ideal laboratories to probe the interplay between galaxy formation and the growth of black holes (BHs) in the early Universe. Mounting observational evidence reveals the presence of BHs in low-mass galaxies across cosmic time, with $\textit{JWST}$ uncovering a likely population of $\textit{overmassive}$ BHs at $2 \lesssim z \lesssim 11$. Simulations struggle to reproduce this high-redshift regime, motivating revisions to models of BH accretion and feedback from active galactic nuclei (AGN). To address this, we present high-resolution cosmological zoom-in simulations of a dwarf galaxy based on FABLE physics, introducing novel sink-based BH accretion models and relaxing the fiducial assumption of strong supernova feedback. BHs accrete more efficiently in the sink-based runs compared to the `traditional' Bondi-based counterparts, with AGN feedback leading to early, rapid quenching maintained by fast, hot and metal-enriched outflows. These outflows pollute the outer circumgalactic medium, yielding flat metallicity gradients down to $z=0$. We further assess the performance of two widely used virial estimators and find significant departures from the true dynamical mass, especially during the high-redshift dwarf assembly. Since our galaxy is dark-matter-dominated at all times and radii, BH growth, tied to the baryon cycle, shows no clear correlation with global dynamical properties. Efficient AGN feedback is produced by overmassive BHs relative to extrapolated local $M_\bullet - M_\star$ relations, raising the possibility that dormant, overmassive BHs in local quenched dwarfs and those probed by $\textit{JWST}$ may reflect a common mode of early and rapid BH growth in low-mass galaxies.

astro-ph.GA

RUBIES: The Evolution of the Ionization Parameter from 0 < z < 9

The dimensionless ionization parameter, U=q/c, where q is the ratio of the local ionizing photon flux to the local hydrogen density, is a key metric to parameterize nebular conditions. Prior to JWST, the rest-frame optical emission lines and their ratios which trace the ionization parameter (e.g., O32=[OIII]/[OII]) were inaccessible at high redshifts. Here we quantify, for the first time, the evolution of the ionization parameter in galaxies across the last 13 billion years of cosmic time by comparing JWST/NIRSpec PRISM and G395M spectroscopy of 434 galaxies at 3<z<9 from the RUBIES survey with z<3 samples from SDSS, LEGA-C, and KBSS. We leverage a large suite of photoionization models to infer U from [OIII] and [OII]. We find that U increases with redshift and specific star formation rate (sSFR), and decreases with stellar mass. Crucially, and in contrast to previous linear best-fit calibrations, our inference results in a systematic uncertainty in logU of ~0.3 dex at zero measurement uncertainty due to the wide range of models that predict the same O32 ratio without informative priors. We compare to SPHINX20 and LUMEN simulations and find that the simulated galaxies exhibit higher O32 ratios at fixed redshift and stellar mass compared to RUBIES observations. Finally, we combine the predictive power of observed and inferred quantities with multivariate relations to estimate U from redshift, stellar mass, and sSFR for use where O32 is not available. We find that U increases at fixed stellar mass and sSFR by a factor of ~4 from z=2 to z=6, demonstrating that the redshift evolution encapsulates physics beyond that traced by stellar mass and sSFR alone. Finally, we show that a toy model with the first order assumption that HII region volume is proportional to galaxy volume can explain the excess redshift dependence of logU as being consistent with observed evolution in galaxy sizes.

astro-ph.GA

An extreme ram-pressure stripping event in a protocluster at redshift 4.3

In the nearby Universe, the environment plays a crucial role in suppressing star formation in dense regions. In particular, ram-pressure stripping (RPS) is a major mechanism for removing gas from galaxies in clusters, occurring when galaxies travel through a dense hot atmosphere and leave trailing gaseous wakes. By depleting the cold gas reservoir, RPS can drive outside-in quenching and is therefore thought to be an important route for transforming cluster galaxies. At earlier times, however, the hot atmosphere in protoclusters is expected to be immature, so environmental effects are commonly assumed to be dominated by gravitational interactions. Here we report ALMA and JWST observations of SPT2349$-$56-C26 (hereafter C26), a massive galaxy experiencing an extreme RPS event in the SPT2349$-$56 protocluster at $z\,{=}\,4.3$. More than half of the [CII]-traced cold gas lies outside its stellar body, with the emission peak offset by 6 kpc. These observations show that RPS can remove most of the cold gas from massive galaxies in dense protocluster cores as early as $z\,{=}\,4.3$, providing a direct hydrodynamic pathway for environmental quenching at $z\,{>}\,4$.

astro-ph.GA