Search arXivSearch

arXiv · 2605.09881

Dissecting Jet-Tagger Through Mechanistic Interpretability

Abstract

Mechanistic interpretability seeks to reverse engineer a trained neural network by identifying the minimal subset of internal components. We perform a mechanistic interpretability analysis of the Particle Transformer architecture, trained on the Top Quark Tagging reference dataset, with the goal of identifying the computational circuit responsible for jet classification and characterizing the physical content of its internal representations. Combining zero ablation, path patching with two complementary on-manifold corruption strategies and linear probing of the residual stream, we identify a sparse six-head circuit that recovers the great majority of the full model performance while admitting a clean source-relay-readout interpretation. In this circuit, a single early layer head serves as the primary causal source, a cluster of middle-layer heads acts as relays selectively attending to hard pairwise substructure and a single late-layer head reads out the aggregated signal. Linear probes show that the residual stream is preferentially aligned with the energy correlator basis over the $N$-subjettiness basis. Within the energy correlator basis, the model preferentially encodes 2-prong substructure observables over the 3-prong observables. A per-layer trained probe further reveals that the apparent single step commitment of the model to a classification decision in the first class attention block is in fact a basis rotation, with the discriminating signal already saturating in the particle attention stack. These results demonstrate that mechanistic interpretability methods developed for natural language models can be used for jet physics classifiers and indicate that gradient descent may rediscover physically meaningful aspects of jet tagging without supervision.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Saurabh Rai, Sanmay Ganguly. 2026-05-11. Dissecting Jet-Tagger Through Mechanistic Interpretability. https://arxiv.org/abs/2605.09881

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Exploring the Singlino-dominated Thermal Neutralino Dark Matter in the $Z_3$ invariant NMSSM

We examine the parameter space of the Next to Minimal Supersymmetric Standard Model (NMSSM) with Singlino-dominated neutralino $\widetildeχ_1^0$ as the lightest supersymmetric particle (LSP). Our study focuses on identifying the regions within this parameter space that produce a thermal relic abundance of $\widetildeχ_1^0$ smaller than the observed cold dark matter relic density while remaining consistent with constraints from LEP measurements, low-energy experiments, Higgs measurements, LHC data, and dark matter direct detection experiments. We identify the dominant annihilation modes of the LSP neutralino across varying LSP mass ranges $\sim \mathcal{O}(1)-\mathcal{O}(10^{3})~$GeV. Furthermore, we conduct a benchmark study to assess the production rates of triple-boson final states emerging from direct electroweakino pair production at the LHC. Drawing insights from these findings, we perform a detailed collider analysis to explore the future potential of probing the triple-boson final states involving a light Higgs boson at the high-luminosity LHC (HL-LHC).

hep-ph

Unveiling the Collins-Soper kernel in inclusive DIS at threshold

We revisit the factorization of inclusive deep inelastic scattering (DIS) near the kinematic threshold in terms of collinear, off-light-cone operators. At threshold, particle production develops around two opposite near-light-cone directions in close analogy with transverse-momentum-dependent semi-inclusive DIS. The Collins-Soper kernel then emerges as the universal function governing the rapidity evolution of the relevant parton correlators in both cases. Our new framework also clarifies outstanding issues related to soft radiation and rapidity divergences at threshold.

hep-ph

Novel Light Dark Matter Detection with Quantum Parity Detector Using Qubit Arrays

We present the design and the sensitivity reach of the Qubit-based Light Dark Matter detection experiment. We propose the novel two-chip design to reduce signal dissipation, with quantum parity measurement to enhance single-phonon detection sensitivity. We demonstrate the performance of the detector with full phonon and quasiparticle simulations. The experiment is projected to detect $\gtrsim 30$ meV energy deposition with nearly $100\%$ efficiency and high energy resolution. The sensitivity to $m_χ\gtrsim 0.01$ MeV dark matter scattering cross section is expected to be advanced by orders of magnitude for both light and heavy mediators, and similar improvements will be achieved for axion and dark photon absorption in the $0.04$-$0.2$ eV mass range.

hep-ph