Search arXiv⌕ Search

arXiv · 2610.06784

How to scale your HEP ML models: A recipe for robust architecture comparisons at scale

Abstract

Much of the recent progress in machine learning domains such as language models has come from scaling laws that predict performance as a function of training effort. In high-energy physics (HEP) similar behavior has now been observed. To aid further study, we present a systematic procedure to derive robust scaling laws and compare design choices on the relevant budget axes for HEP tasks. We first validate the full scaling trajectory on toy problems and then apply the procedure to multi-task transformers on the ~11 billion-jet ATLAS JetSet2 dataset, in both the compute- and data-constrained regimes. For the latter, we predict, to the best of our knowledge for the first time, the jointly optimal model size, training horizon, learning rate and batch size under early stopping. At compute-optimal scaling, we recover a near-equal $\sqrt{C}$ dependence of model and dataset size, and find that auxiliary objectives lower the primary jet-classification loss at equal compute budget. Expanding the inputs toward lower-level data systematically lowers the loss while leaving the scaling exponent nearly unchanged. The onset of the power-law regime is itself set by scale: below a threshold in dataset size the loss carries little information about high-compute scaling, underscoring the value of large, high-quality full-simulation datasets as a foundation for scaling studies and the development of foundation models in HEP.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Matthias Vigl, Nikita Pond, Jackson Barr, Alexander Froch, Dan Guest, Nicole Hartman, Michael Kagan, Lukas Heinrich. 2026-10-05. How to scale your HEP ML models: A recipe for robust architecture comparisons at scale. https://arxiv.org/abs/2610.06784

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Off-shell Higgs boson measurements: Yukawa couplings, self-coupling, compositeness, and width

Measurements of Higgs boson production in the off-shell region are presented, using the four-lepton decay channel. Data from proton-proton collisions at the CERN LHC, collected by the CMS experiment and corresponding to an integrated luminosity of 138 fb$^{-1}$ at a center-of-mass energy of 13 TeV, are utilized. The first direct test of composite Higgs boson models is performed, with a lower limit on the compositeness scale $Λ_\mathrm{H}$ set at 870 GeV at the 95% confidence level. Tests of gluon-fusion production within the standard model effective field theory framework are performed. The first constraint on the Higgs boson self-coupling in the off-shell region is obtained. By combining on- and off-shell measurements, the analysis sets the tightest constraints to date on light-quark Yukawa couplings, while relaxing assumptions such as the bound on the Higgs boson coupling to vector bosons, thereby providing more model-independent results. Constraints on the Higgs boson width are provided while accounting for a range of beyond-the-standard-model effects, including both light and heavy particles in the gluon-fusion production loop, as well as modified couplings to vector bosons. A combined analysis of the H$\to$ZZ and PH$\to$WW channels is performed to improve sensitivity to the Higgs boson width, yielding $Γ_\mathrm{H}$ = 5.1$^{+2.0}_{-1.8}$ MeV. The scenario of no off-shell Higgs boson production is excluded at a confidence level exceeding 5 standard deviations.

hep-ex↗

The mathematical theory of photomultiplier tube calibration

In this technical note we describe the main features of the mathematical theory of photomultiplier tube (PMT) calibration. Attention was paid to explain the various arguments and concepts in a simple and pedagogical manner that everybody understands. The basic operational principles of a PMT are discussed from a theoretical standpoint. The essential steps of its function (photoconversion, focusing, multiplication, etc.) are laid down together with the mathematical schemes necessary to model the charge output of a PMT when illuminated by a faint poissonian light source. In case of important omissions we direct the reader to some of the standard references. The most common numerical methods, used to calculate the charge amplification function $S_R(x)$, are also presented. As an example, we plot $S_R(x)$ utilizing a gamma function model for the single photoelectron (SPE) response and showcase its main characteristics. The basic techniques for gain calibration are also introduced, and we probed their precision using toy Monte Carlo data. Additionally, data from a Hamamatsu R1408 PMT were analyzed. Finally, we show how the presence of soft charge component can affect the results of gain determination. We conclude this report with some general comments regarding \emph{in situ} calibration of large-scale detectors. We hope that this document can serve as a reference for students and young researchers that want to learn more about the theory of PMT calibration.

hep-ex↗

Measurement of the fragmentation properties of jets containing $Υ$(nS) mesons in proton-proton collisions at $\sqrt{s}$ = 13 TeV

A measurement of the fragmentation properties of jets containing $Υ$(nS) mesons using proton-proton collision data at $\sqrt{s}$ = 13 TeV, corresponding to an integrated luminosity of 138 fb$^{-1}$, is presented. The $Υ$(nS) mesons associated with jets are reconstructed via their decays producing pairs of oppositely charged muons. The longitudinal and transverse projections of the momenta of the $Υ$(nS) mesons along the momenta of the corresponding jets are studied. The results are compared with Monte Carlo predictions, including recent developments in the modelling of quarkonia production in parton showers. The description of the data by these Monte Carlo predictions is found to be unsatisfactory, leaving room for improvements in the modelling of such processes.

hep-ex↗