Search arXivSearch

arXiv · 2609.03599

Reduced latent leakage does not reliably predict lower likelihood bias in collider inference

Abstract

Reusable collider representations can be evaluated through downstream discrimination and probes of retained information, but neither quantity directly tests the behaviour of score templates in a profiled likelihood. We test a specific prediction in a controlled two-channel routing protocol: if reduced physics-label readability in a nuisance branch indicates a more inference-robust representation, it should accompany a smaller profiled signal-strength bias under fixed unmodelled shifts. In a public Compact Muon Solenoid $H\rightarrow ZZ\rightarrow4\ell$ workflow, a downstream split of fixed EveNet embeddings preserves signal/background area under the receiver operating characteristic curve ($0.9894\pm0.0004$) while reducing nuisance-branch physics readability from $0.961\pm0.013$ to $0.593\pm0.030$. Probe-sensitivity and effective-rank controls exclude a failed readout and branch collapse. In a separate top quark jet-tagging workflow, the leakage reduction recurs with preserved task performance. Across two development event shards, however, its Spearman association with maximum absolute profiled bias is $0.036$, and three of six material leakage-improving transitions do not reduce that bias. A one-shot preregistered confirmation on an independently accessed shard produces material leakage reductions in all three paired seeds, while the maximum absolute bias increases in two. Thus, within the tested protocol, latent readability is a useful routing diagnostic but not a likelihood-robustness certificate. The result supports a practical validation rule: claims about inference robustness require a prespecified likelihood-facing stress test and held-out confirmation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tong Pan. 2026-09-03. Reduced latent leakage does not reliably predict lower likelihood bias in collider inference. https://arxiv.org/abs/2609.03599

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Online local learning for generative thermodynamic computing

Generative thermodynamic computers turn thermal noise into structured data through Langevin dynamics. We train these systems with a local update at each integration step. The reverse-path Onsager-Machlup objective yields a coupling gradient that is a symmetric sum of local residual-state correlations. We apply this gradient immediately rather than accumulating it over a full trajectory. In digital simulations using MNIST prototypes, online and trajectory-batch training reach similar validation losses on fixed noising paths. Models trained online release less heat on average in all five independently seeded pairs, with both models' parameters held fixed during sampling. Auxiliary classifier and nearest-prototype measures change modestly, while pairwise diversity decreases. The response to noise depends strongly on where the errors enter: independent zero-mean errors in the formed updates produce little heat change over a finite range of noise amplitudes, whereas residual offset and temporal correlation have much larger effects. Storing trained couplings requires substantially less precision than resolving deterministic updates during training. Together, these results establish a local online training method and show how update timing, noise structure, and precision affect generative thermodynamic computing.

physics.data-an

Information-Loss Location Estimation and Curvature-Scale Analysis of Physical Measurements: A Gaussian, Cauchy, and Logistic Neutron-Lifetime Benchmark

We develop and implement a two-step location-scale analysis for repeated physical measurements when the underlying distribution is not known. For a chosen substitute distribution, stationarity of population Kullback-Leibler information loss at fixed scale motivates the location score, which we apply to the observed measurements through empirical cross-entropy. Steiner's curvature rule then defines a companion scale separately. We place Gaussian, Cauchy, and logistic substitutes in one common construction and implement them reproducibly on a physics benchmark. For the logistic substitute, we pair the established bounded location score with a scale defined by the curvature rule rather than by a fitted tuning constant or a scale-likelihood equation. For Gaussian measurements with different quoted uncertainties, an experiment-level curvature convention and a stated variance mapping give an exact algebraic connection to a standard residual-based variance estimate for a weighted fit. Applied to a published 21-measurement neutron-lifetime compilation, the inverse-variance weighted, most frequent value, and logistic locations are 878.689, 881.164, and 882.835 s, respectively. These values are benchmark results of the compared constructions and are not proposed as a new recommended neutron lifetime.

physics.data-an

Finite Asimov Sample Construction in Unbinned Neural Simulation-Based Inference

Expected sensitivity calculations in frequentist analysis using neural simulation-based inference often require large simulated samples to approximate an unbinned Asimov dataset. We construct a finite weighted reference sample, for analyses based on learned density ratios, for which the generating parameters globally maximize the weighted likelihood. A toy example with five-dimensional observable space based loosely on high-energy physics models shows that an Asimov dataset consisting of as few as 256 weighted reference events closely reproduce the expected test statistic scans obtained with two million simulated reference events. The resulting test statistic distributions are also shown to agree with independent simulator-based pseudo-experiments. We note that this agreement with the simulator depends on the accuracy of the learned ratios and must be validated explicitly.

physics.data-an