Search arXivSearch

arXiv · 1701.00426

Design and optimization of a portable LQCD Monte Carlo code using OpenACC

Abstract

The present panorama of HPC architectures is extremely heterogeneous, ranging from traditional multi-core CPU processors, supporting a wide class of applications but delivering moderate computing performance, to many-core GPUs, exploiting aggressive data-parallelism and delivering higher performances for streaming computing applications. In this scenario, code portability (and performance portability) become necessary for easy maintainability of applications; this is very relevant in scientific computing where code changes are very frequent, making it tedious and prone to error to keep different code versions aligned. In this work we present the design and optimization of a state-of-the-art production-level LQCD Monte Carlo application, using the directive-based OpenACC programming model. OpenACC abstracts parallel programming to a descriptive level, relieving programmers from specifying how codes should be mapped onto the target architecture. We describe the implementation of a code fully written in OpenACC, and show that we are able to target several different architectures, including state-of-the-art traditional CPUs and GPUs, with the same code. We also measure performance, evaluating the computing efficiency of our OpenACC code on several architectures, comparing with GPU-specific implementations and showing that a good level of performance-portability can be reached.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Claudio Bonati, Enrico Calore, Simone Coscetti, Massimo D'Elia, Michele Mesiti, Francesco Negro, Sebastiano Fabio Schifano, Giorgio Silvi, Raffaele Tripiccione. 2017-01-02. Design and optimization of a portable LQCD Monte Carlo code using OpenACC. https://doi.org/10.1142/s0129183117500632

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Symplectic lattice gauge theories in the Grid framework: domain wall fermions and continuum extrapolations

We report the results of the first numerical lattice study using domain-wall fermions in the Sp(4) gauge theory coupled to two flavours of (Dirac) fermions, transforming in the fundamental representation of the gauge group. This theory plays a prominent role in the literature on extensions of the Standard Model with composite dynamics. It provides a short-distance completion for a class of composite Higgs models, or, alternatively, of dark matter models based on the strongly interacting massive particle paradigm. We adopt the Möbius formulation of domain-wall fermions (MDWF), implemented within the Grid software environment. We report the results of extensive tests of the algorithm implementation, and of the optimisation of the choices of algorithmic parameters appearing in the MDWF action. We then measure masses and decay constants of the lightest flavoured mesons in ensembles with moderately large fermion masses and several choices of lattice coupling, and perform an extrapolation to the continuum. We compare our results for the physical observables to published measurements obtained in the same field theory, but derived on the lattice by employing Wilson fermions. We demonstrate that, with the deployment of moderate computational resources, the MDWF formulation can yield order-of-magnitude gains in the approach to the continuum limit, in regions of physical parameter space relevant to phenomenological applications of this theory.

hep-lat

Numerical Investigations of Phase Transitions in Lattice Field Theories

The study of phase transitions plays an important role in understanding qualitative changes in the behaviour of physical systems at criticality. Despite decades of progress, there is still a strong demand for high-precision numerical tools capable of resolving subtle critical phenomena. Motivated by this need, in this thesis, we present two complementary numerical investigations of phase transitions in lattice systems. The first uses GPU-accelerated higher-order tensor renormalization group (HOTRG) techniques to study the two-dimensional generalized XY model, characterizing its ferromagnetic, nematic, and paramagnetic phases and mapping their phase boundaries using thermodynamic observables in the thermodynamic limit. The second develops and benchmarks a configurational temperature estimator, constructed from gradients and Hessians of the Euclidean lattice action, in compact U(1) lattice gauge theories. On one hand, tensor network methods capture rich phase structures when truncation and finite-bond effects are adequately controlled. On the other hand, the configurational temperature estimator provides an independent, low-overhead means of validating thermal sampling across different algorithms and models, and can also be used as a runtime diagnostic to identify sampling pathologies before large-scale production runs.

hep-lat

Computability of GPDs near $x=\pmξ$ in Lattice QCD

In lattice QCD computations of generalized parton distributions (GPDs), the large momentum expansion generally requires all hard scales, $2|x\pmξ|P^z$ and $2|1\pm x|P^z$, to be much larger than $Λ_{\rm QCD}$. We show that this condition can be relaxed for $2|x\pmξ|P^z$ at large $ξ$, making the important $x\sim\pmξ$ regions accessible to lattice calculations and considerably expanding the region of computability. Revisiting previous lattice results with complete one-loop matching, we obtain the expected partonic threshold behavior---GPDs continuous at $x=\pmξ$ but with discontinuous derivatives---which has not previously been observed on the lattice. We thus obtain, for the first time, important prediction for GPDs in the distribution-amplitude-like region, which smoothly connects the quark and antiquark PDF-like behaviors.

hep-lat