Search arXivSearch

arXiv · 2312.06838

Optimizing High Throughput Inference on Graph Neural Networks at Shared Computing Facilities with the NVIDIA Triton Inference Server

Abstract

With machine learning applications now spanning a variety of computational tasks, multi-user shared computing facilities are devoting a rapidly increasing proportion of their resources to such algorithms. Graph neural networks (GNNs), for example, have provided astounding improvements in extracting complex signatures from data and are now widely used in a variety of applications, such as particle jet classification in high energy physics (HEP). However, GNNs also come with an enormous computational penalty that requires the use of GPUs to maintain reasonable throughput. At shared computing facilities, such as those used by physicists at Fermi National Accelerator Laboratory (Fermilab), methodical resource allocation and high throughput at the many-user scale are key to ensuring that resources are being used as efficiently as possible. These facilities, however, primarily provide CPU-only nodes, which proves detrimental to time-to-insight and computational throughput for workflows that include machine learning inference. In this work, we describe how a shared computing facility can use the NVIDIA Triton Inference Server to optimize its resource allocation and computing structure, recovering high throughput while scaling out to multiple users by massively parallelizing their machine learning inference. To demonstrate the effectiveness of this system in a realistic multi-user environment, we use the Fermilab Elastic Analysis Facility augmented with the Triton Inference Server to provide scalable and high throughput access to a HEP-specific GNN and report on the outcome.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Claire Savard, Nicholas Manganelli, Burt Holzman, Lindsey Gray, Alexx Perloff, Kevin Pedro, Kevin Stenson, Keith Ulmer. 2023-12-11. Optimizing High Throughput Inference on Graph Neural Networks at Shared Computing Facilities with the NVIDIA Triton Inference Server. https://arxiv.org/abs/2312.06838

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Search for dark matter in a signature with a four-prong large-radius jet in proton-proton collisions at $\sqrt{s}$ = 13 TeV

A search for a pair of nonprompt dark matter (DM) candidates produced in association with an initial-state radiation jet, in a signature containing a four-prong large-radius jet, is presented. The signal model contains a heavy vector or axial-vector mediator, which produces long-lived dark-sector particles that decay to a stable DM particle and a light boson, which decays to quarks. The analysis is based on data collected in the years 2016$-$2018 with the CMS detector at the LHC in proton-proton collisions at $\sqrt{s}$ = 13 TeV, corresponding to an integrated luminosity of 138 fb$^{-1}$. Signal candidates feature large-radius jets, which are identified using a jet substructure tagger based on a graph neural network. The large-radius jet aims to reconstruct the decay of light DM mediators into four quarks, which are produced in association with two stable DM particles. The standard model background contributions are estimated from data using dedicated control regions. The missing transverse momentum spectrum is probed for a potential signal over the expected background. No significant excess over the standard model expectation is observed. Upper limits at 95% confidence level are set on the signal strength as functions of either the mediator mass or the relevant coupling. This is the first search for a pair of nonprompt DM candidates in the Lorentz-boosted topology, characterized by a large-radius jet and large missing transverse momentum.

hep-ex

Charge-dependent atmospheric muon flux at 17 GV geomagnetic cutoff with the mini-ICAL detector

The Iron CALorimeter (ICAL) detector at the India-Based Neutrino Observatory (INO) was conceived as an underground experiment designed to measure atmospheric neutrino oscillation parameters. As part of the R\&D programme, a scaled prototype (mini-ICAL), 85\,ton, approximately 1/600$^{\mathrm{th}}$ the mass of the full detector, was constructed at the IICHEP Transit Campus, Madurai (altitude 150\,m; latitude 9.9372$^\circ$\,N; longitude 78.013$^\circ$\,E; geomagnetic latitude 1.44$^\circ$\,N; vertical cutoff rigidity 17\,GV) and operated between 2018 and 2022. The prototype enabled measurements of charge-dependent cosmic muon spectra in the vicinity of the geomagnetic equator and provided an important validation of detector performance, reconstruction algorithms, and simulation frameworks for the ICAL experiment. Differential fluxes of $μ^{-}$ and $μ^{+}$ were measured over the momentum range $\sim$\,1--5\,GeV/c. The obtained momentum spectra are systematically lower than those reported at sites with smaller geomagnetic cutoff rigidities, consistent with the suppression of low- and intermediate-rigidity primary cosmic rays at the 17\,GV cutoff. The measurements are compared with predictions from different hadronic interaction models available in CORSIKA simulations.

hep-ex

Laboratory constraints on peV-scale mass splitting between ordinary and sterile neutron states

Sterile states of matter, represented by a parallel ``mirror'' sector, may contribute to the observed dark matter in the Universe. We investigated the parameter space of neutron $(n)$ to mirror-neutron $(n')$ oscillations, in the case where the two states are not necessarily mass-degenerate, taking into account interactions in the mirror sector. By tuning the magnitude of an applied magnetic-field in the range $5~μ\mathrm{T} < B < 360~μ\mathrm{T}$ to corresponding resonance conditions for finite mass splitting, we derive exclusion limits for the $n-n'$ oscillation time constant reaching about $20~\text{s}$ over the mass-difference range $0.3 - 22~\text{peV}$. In parts of this parameter range, our limits exceed the model-dependent neutron-star-cooling bound, providing the first experimental constraints in this scenario that are more stringent than this astrophysical estimate.

hep-ex