Search arXivSearch

arXiv · 0907.3390

GAMER: a GPU-Accelerated Adaptive Mesh Refinement Code for Astrophysics

Abstract

We present the newly developed code, GAMER (GPU-accelerated Adaptive MEsh Refinement code), which has adopted a novel approach to improve the performance of adaptive mesh refinement (AMR) astrophysical simulations by a large factor with the use of the graphic processing unit (GPU). The AMR implementation is based on a hierarchy of grid patches with an oct-tree data structure. We adopt a three-dimensional relaxing TVD scheme for the hydrodynamic solver, and a multi-level relaxation scheme for the Poisson solver. Both solvers have been implemented in GPU, by which hundreds of patches can be advanced in parallel. The computational overhead associated with the data transfer between CPU and GPU is carefully reduced by utilizing the capability of asynchronous memory copies in GPU, and the computing time of the ghost-zone values for each patch is made to diminish by overlapping it with the GPU computations. We demonstrate the accuracy of the code by performing several standard test problems in astrophysics. GAMER is a parallel code that can be run in a multi-GPU cluster system. We measure the performance of the code by performing purely-baryonic cosmological simulations in different hardware implementations, in which detailed timing analyses provide comparison between the computations with and without GPU(s) acceleration. Maximum speed-up factors of 12.19 and 10.47 are demonstrated using 1 GPU with 4096^3 effective resolution and 16 GPUs with 8192^3 effective resolution, respectively.

Explore related subjects

Keep this discovery

BibTeXRIS

Hsi-Yu Schive, Yu-Chih Tsai, Tzihong Chiueh. 2009-12-24. GAMER: a GPU-Accelerated Adaptive Mesh Refinement Code for Astrophysics. https://doi.org/10.1088/0067-0049/186/2/457

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The EDD Radio Astronomy Backend Framework

Modern digital radio astronomy receivers produce increasingly wide-bandwidth, high bit-rate data streams that necessitate the development of flexible, scalable, and maintainable backend processing and recording systems. Historically, such backend instrumentation has been tightly coupled to telescope observing modes, limiting reuse between observatories and science cases. We present the Effelsberg Direct Digitisation (EDD) backend framework, a software-defined architecture for constructing real-time radio astronomy backends on commodity off-the-shelf computing infrastructure. We describe its design, implementation, supported observing modes, and operational deployments. EDD separates a common core framework from plugin-provided observing capabilities. The core provides orchestration, telescope interfaces, pipeline lifecycle management, monitoring, and deployment tooling, while plugins implement processing pipelines for specific observing modes. The framework is designed to support both single-dish and interferometric instruments through site-specific configuration and plugin selection. EDD currently supports spectroscopy and spectropolarimetry, pulsar timing and searching, baseband recording, very long baseline interferometry, correlation, and beamforming. Operational deployments include the Effelsberg 100-m telescope, the SKA-MPI prototype dish, the Thai National Radio Telescope, and the ARGOS interferometric prototype array. By separating common services, observing-mode plugins, and site-specific configuration, it allows backend capabilities to be deployed across heterogeneous telescope environments and provides a community resource for broadband radio astronomy instrumentation.

astro-ph.IM

Bayesian Superiority in On/Off analysis

We present a detailed comparison of Bayesian criteria with three non-informative priors - flat, Jeffreys, and scale-invariant - for testing a signal against an unknown background and compare them with the classical frequentist Li-Ma approach in the On/Off problem. We perform Monte Carlo simulations for various background levels and evaluate the Li-Ma and Bayesian criteria by their Type I error rates. We then simulate a nonzero signal and compare the criteria in terms of Type II error rates. We find that the Bayesian criterion with the Jeffreys prior yields lower Type I and Type II error rates than the Li-Ma criterion. In addition, we show that the Bayesian criteria are more robust than the Li-Ma criterion when the background distribution is overdispersed relative to the Poisson distribution.

astro-ph.IM

An RFSoC-based Backend and Timing System for the Balloon-borne Very Long Baseline Interferometry Experiment

We present the design and performance characterization of the digital backend and precision-timing system for the Balloon-borne Very Long Baseline Interferometry Experiment (BVEX), a pathfinder for high-frequency stratospheric VLBI at 22 GHz. The backend uses one of the four 14-bit analog-to-digital converter inputs on an AMD-Xilinx RFSoC 4x2. Although the converters support sampling rates up to 5 GSPS, the flight configuration digitizes the 2-4 GHz intermediate frequency at 4.096 GSPS. CASPER firmware provides both a high-resolution spectrometer for pointing and receiver verification, and a VLBI acquisition chain with two-bit requantization that records at a rate of about 8.2 Gbps. The timestamped data packets are sent over 100 Gigabit Ethernet (GbE) to a 16 TB NVMe array in a storage computer that draws approximately 70-80 W. The timing chain uses a Rakon oven-controlled crystal oscillator as a timing reference while a time-interval counter measures its drift relative to a GPS reference with approximately 60 ps resolution. This is the first deployment of an RFSoC-based VLBI backend and precision-timing system on a stratospheric balloon. Ground tests validated the backend, spectrometer, and timing chain. The August 2025 CSA STRATOS flight ended before reaching the target float altitude because of a balloon failure, and as a result no science observations were obtained. For the planned 2027 reflight, we are developing a conduction-cooled data storage computer with 24 TB of NVMe capacity and a direct data path from the 100 GbE interface to the NVMe array.

astro-ph.IM