Search arXivSearch

arXiv · 2608.12392

Learned proposals in trans-dimensional inference are optimal at equilibrium, not during assembly

Abstract

Inferring the dimension of a model - the number of components needed to explain data - jointly with the parameters is a pervasive problem, from counting sources in an image to mixture modeling, and reversible-jump Markov chain Monte Carlo solves it exactly but mixes slowly. Learned proposals are well established at fixed dimension, but whether they can accelerate the dimension-changing moves themselves has remained largely untested. We show that the answer has a structural origin: the optimal proposal for the dimension-changing birth move is a different object in different phases of the run. While the fit is being assembled it must match the current residual - a state-dependent quantity no state-independent network can represent - but at equilibrium it degenerates to the posterior's single-component marginal, which is exactly the distribution an adaptive normalizing flow learns from the sampler's own history. A learned state-independent birth proposal is therefore useless in one phase and optimal in the other. Controlled experiments confirm the attribution: applied with an exact Metropolis--Hastings correction that leaves the target invariant for any network, the learned births leave acceptance rates unchanged yet accelerate model-order mixing - in a ten-seed benchmark they meet a pre-specified stopping rule in six of ten runs, typically several times sooner, where a strong hand-tuned baseline meets it in one (one-sided p=0.03) - and an isolation experiment shows the same flow deployed within-model buys nothing. Making no domain-specific assumptions, the same sampler counts sources in a noisy image and reconstructs signals across scientific domains, including gravitational waves from ground- and space-based detectors and a scalp EEG recording. We release the method as HyperWave, an open-source package.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Argyro Sasli, Nikolaos Karnesis, Minas Karamanis, Michael L. Katz, Dimitrios Kourtesis, Michael W. Coughlin, Vuk Mandic, Nikolaos Stergioulas. 2026-08-04. Learned proposals in trans-dimensional inference are optimal at equilibrium, not during assembly. https://arxiv.org/abs/2608.12392

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

When Should Team KPIs Be Absolute or Relative for Match-Outcome Prediction?

In rugby union and association football, team key performance indicators (KPIs) can be represented in absolute terms or relative to the opponent. Relativisation sometimes improves match-outcome prediction and sometimes harms it. There has been no general account of when each occurs. The answer depends on how much the two teams differ in variability and how strongly their KPI values rise and fall together. We combine these properties into the Paired Efficiency Factor (PEF), which generalises Fisher's paired-efficiency result to the unequal-variance conditions typical of competitive sport. The PEF also connects a KPI's statistical efficiency to how much information its relative form carries about the outcome. Combined KPIs can interact in complex ways, so we analyse each indicator on its own. Across 86 team KPIs from professional rugby union and association football, the PEF places every metric in one of four regimes. Anti-correlation is common, especially for high-volume competitive counts, and absolute measures are then usually preferable. Relativisation can still improve prediction when a noisier difference carries more outcome information. An idealised simulation and one representative KPI from each regime confirm both signs under team-blocked cross-validation. The same pairing geometry appears in healthcare, genomics, finance, and manufacturing. The PEF turns an ad hoc feature-engineering choice into a transparent, data-informed diagnostic for when to relativise performance metrics and when not to.

physics.data-an

Comparison of Image Processing Models in Quark Gluon Jet Classification

Quark-gluon discrimination provides a useful test case for studying how different machine-learning architectures learn the spatial structure of QCD radiation. In this work, we compare convolutional neural network (CNN), Vision Transformers (ViT), and hierarchical Swin Transformers using the same three-channel jet-image representation, consisting of charged-particle momentum, neutral-particle momentum, and charged-particle multiplicity from PYTHIA 8 jets. We study their performance for different training-set sizes and fine-tuning configurations, with particular attention to the role of local and global information in the jet images. CNN and Swin models consistently perform better than ViT in the cases studied. Since both CNN and Swin retain a strong local component in their architectures, this suggests that local jet substructure plays an important role in quark-gluon discrimination. The performance of the hierarchical Swin model also suggests that combining local features over larger spatial scales is useful. Block-wise fine-tuning improves the performance of the Transformer models, although the improvement becomes smaller and the training less stable as more blocks are unfrozen. We also find that self-supervised Momentum Contrast (MoCo) pretraining improves the model initialization, particularly when the amount of labeled training data is limited. Based on these observations, we developed a smaller Swin model adopted to the jet-image representation used in this study. It achieves comparable performance with substantially fewer parameters. The results show that it is important to adapt the model architecture and training procedure to the specific input characteristics of High Energy Physics (HEP) data when applying vision models in HEP.

physics.data-an

Online local learning for generative thermodynamic computing

Generative thermodynamic computers turn thermal noise into structured data through Langevin dynamics. We train these systems with a local update at each integration step. The reverse-path Onsager-Machlup objective yields a coupling gradient that is a symmetric sum of local residual-state correlations. We apply this gradient immediately rather than accumulating it over a full trajectory. In digital simulations using MNIST prototypes, online and trajectory-batch training reach similar validation losses on fixed noising paths. Models trained online release less heat on average in all five independently seeded pairs, with both models' parameters held fixed during sampling. Auxiliary classifier and nearest-prototype measures change modestly, while pairwise diversity decreases. The response to noise depends strongly on where the errors enter: independent zero-mean errors in the formed updates produce little heat change over a finite range of noise amplitudes, whereas residual offset and temporal correlation have much larger effects. Storing trained couplings requires substantially less precision than resolving deterministic updates during training. Together, these results establish a local online training method and show how update timing, noise structure, and precision affect generative thermodynamic computing.

physics.data-an