Search arXiv⌕ Search

arXiv · 2610.08238

GPU Acceleration of Awkward Arrays: Using Python cuda.compute

Abstract

Awkward Array is a widely used library in high-energy physics (HEP) for representing and manipulating nested, variable-length data in Python. Previous CHEP contributions have explored GPU acceleration for Awkward Array, demonstrating the feasibility and performance benefits of CUDA-based backend while also identifying limitations related to irregular data access, fine-grained kernel launches, and composability of operations. In this contribution, we present recent developments that build directly on these earlier efforts by introducing a CUDA execution model for Awkward Array based on the Python CUDA Core Compute Libraries (CCCL). Using CCCL, we eliminate the need for custom CUDA kernels and can instead use a high-level Python interface. The CCCL-based approach also enables fusion of multiple Awkward operations into a reduced number of CUDA kernels, addressing kernel launch overhead observed in earlier GPU implementations. Lazy execution allows expression graphs to be constructed and optimized prior to kernel generation, improving performance for analysis workflows involving jagged arrays, combinatorial operations, and reductions. In contrast to earlier approaches, this design also emphasizes extensibility, allowing user-defined Python code to be incorporated into GPU execution paths with minimal boilerplate and without breaking existing analysis semantics. We present performance studies that demonstrate improvements over previously reported eager GPU execution strategies for representative HEP analysis patterns. These developments extend the GPU capabilities of Awkward Array toward a more composable and sustainable backend, aligned with the needs of Python-based analysis at the HL-LHC and beyond.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Maksym Naumchyk, Ianna Osborne. 2026-10-06. GPU Acceleration of Awkward Arrays: Using Python cuda.compute. https://arxiv.org/abs/2610.08238

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

VeriNC: Finding Design Risks of In-Network Computing Systems

The emergence of programmable switches has brought in-network computing (INC) into the spotlight in recent years. By offloading computation directly onto the data transmission process, INC improves network utilization, reduces latency to sub-RTT levels, saves link bandwidth, and maintains throughput. However, INC disrupts the transparency of traditional networks, forcing developers to consider network exceptions like packet loss and out-of-order. If not properly handled, these exceptions can lead to violations of application properties, such as cache consistency and lock exclusion. Usual testing cannot exhaustively cover these exceptions, raising doubts about the correctness of INC systems and hindering their deployment in the industry. This paper presents VeriNC, the first general-purpose tool for verifying INC systems. VeriNC provides a high-level specification language and saves developers 67.2% lines of code on average. To help better understand the behavior of the system, VeriNC offers configurable network environments. VeriNC enables developers to express INC-specific correctness properties. VeriNC translates developer-specified systems into state transition representations, performs model checking to detect potential design risks, and reports violation traces to developers. We propose optimizations for INC-specific scenarios to address the challenge of state space explosion. We modeled INC systems across four application domains and identified design risks with VeriNC in seconds. VeriNC has also been adopted to guide the design of a new INC protocol. Based on our verification experience, we summarize lessons that help develop a correct INC protocol. We further reproduce them in real systems to confirm the validity of our verification result.

cs.DC↗

TorchGWAS 1.0: GPU-accelerated GWAS at scale

Imaging, molecular, and machine-learning workflows can generate thousands of quantitative phenotypes in a single cohort, creating substantial computational and output bottlenecks when testing traits individually. TorchGWAS is a GPU-accelerated framework that uses batched operations for high-throughput, covariate-adjusted linear association testing across large panels of quantitative phenotypes. Across 500,036 allele-harmonized tests, TorchGWAS t statistics agreed with PLINK 2.0. On an NVIDIA H100 80-GB GPU with a 48-core Intel Xeon Gold 6442Y host and measured disk read and write rates of 5.98 and 1.49 GB/s, respectively, median end-to-end times for 4.57 billion associations (8,931,083 variants by 512 phenotypes in 35,365 samples) were 28.46 s for BED, 29.13 s for hard-call PGEN, 51.48 s for BGEN, and 58.95 s for dosage PGEN, including writing 36.7 GB of binary summary statistics. TorchGWAS provides an efficient Python-based framework for parallel fixed-effect association screening at biobank scale.TorchGWAS is implemented in Python and distributed as a documented source repository at https://github.com/ZhiGroup/TorchGWAS.

cs.DC↗

AI-Assisted Computational Reproducibility on the FABRIC Testbed

Computational reproducibility remains difficult despite being central to scientific research. In this paper, we show how the international FABRIC testbed, combined with a large language model (LLM) coding agent through LoomAI, can simplify reproducing published experiments across multiple domains. We reproduced three case studies on FABRIC, covering BBR-family congestion-control evaluations, LAMMPS molecular dynamics scaling benchmarks on a CPU-only MPI cluster, and stress protein homeostasis genomics pipelines. Rather than focusing only on matching numerical outputs, we evaluate whether the reproduced experiments support the same scientific conclusions as the original studies. The AI assistant was effective in setting up the environment, adapting code, and debugging, but struggled with the analysis stages that lacked clearly defined workflows, which required human guidance to establish execution order and data dependencies. Across the case studies, the AI-assisted workflow reduced reproduction effort by roughly 4--6 times. We conclude with practical recommendations for improving AI-assisted reproducibility on research testbeds.

cs.DC↗