Search arXivSearch

SEARCH · Search arXiv

Results for “nucl-th”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

117 recordsLinked to original sources

Inclusive electron-nucleus cross section models from domain adaptation

We apply transfer learning (TL) to construct data-driven models of inclusive electron-nucleus cross sections. Starting from an ensemble of deep neural networks pretrained on \(^{12}\)C data, we fine-tune the models separately for \(^{3}\)He, \(^{6}\)Li, \(^{16}\)O, \(^{27}\)Al, \(^{40}\)Ca, and \(^{56}\)Fe. The resulting models improve for all targets, marginally so for oxygen, where the carbon baseline is already adequate, although their predictive robustness depends on the amount, coverage, and precision of the available target data. We systematically study how model performance depends on the number of fine-tuned layers, on the fraction and selection of the training data, and on the overlap between the source and target kinematic domains. The layer-wise analysis shows that oxygen requires only shallow adaptation, whereas helium, calcium, and iron require substantially deeper fine-tuning. Lithium represents the least robust case because of its limited dataset, while aluminum demonstrates a strong sensitivity to a small subset of highly constraining measurements. For selected kinematic configurations outside the coverage of the carbon training data, the adapted models remain consistent with the measurements within their estimated uncertainties. Finally, we compare the resulting predictions with those of the phenomenological F1F2 model.

hep-ph

Bridging Ab Initio Symmetries and Global Nuclear Masses with Interpretable Neural Networks

Ab initio theory establishes Wigner's SU(4) and Elliott's SU(3) as dominant symmetries of the nuclear force in light and intermediate-mass nuclei. Previous work shows the relevance of the former symmetry for nuclear binding, whether the latter organizes binding remains elusive. We probe whether both these symmetries organize nuclear masses, aiming at physical insights through interpretable models and predictive capability. From the SU(3) and SU(4) Casimirs we build 3 neural-network models. Two are conventional, a feature-informed NN (FINN) and a Gaussian variant (GINN) with predictive spread, while Wigner-informed network (WINN) is a new design constraining the mass formula to be linear in the operators, learning their (N,Z)-dependent couplings, so that the model is intrinsically explainable. All are trained on AME2016 subtracted by the liquid drop model at 4 data fractions and validated on nuclei new to AME2020, with extrapolation benchmarked against HFB-26 and r-process. The Casimir features carry binding information far beyond the bulk, and SHAP analysis suggests the quadratic SU(4) Casimir as the leading contributor to the residual binding. The WINN yields the best performance, reaching a 0.412 MeV validation error and competitive with state-of-the-art models, and importantly, when trained on the sparsest dataset it outperforms the other NNs trained on the densest. Off the known chart its masses track HFB-26 as closely as WS3 and reproduce the solar abundance peaks of a neutron-star-merger simulation. The WINN's coupling fields reveal an enhanced even-SU(4) contribution toward the neutron dripline, hinting at restoration of Wigner's symmetry. The SU(4) and SU(3) structures reach beyond individual nuclei to organize binding, and embedding symmetry-preserving operators directly in a domain-informed interpretable architecture yields a physically transparent model less hungry for data.

nucl-th

How Architecture and Training Affect TPC Representations Across Experiments

Deep-learning efforts have increasingly shifted toward foundation model approaches. In experimental physics, this allows models and learned representations to be reused beyond the experiments in which they were developed. This work evaluates the reusability of representations across experiments and detector systems using probes on frozen encoders. These probes reveal task-relevant structure before downstream adaptation, complementing fine-tuning. Together with random-weight controls, they distinguish contributions from architecture and encoder training that downstream performance alone cannot resolve. Time projection chamber (TPC) data provide a useful testbed because events from TPC systems can be represented as variable-length sparse tensors, while detector geometries, event topologies, and scientific tasks can differ substantially. We investigate whether fixed-dimensional TPC event representations can be reused across classification tasks, experiments, and detector systems. Sparse ResNet and PointNet-style encoders produce 512-dimensional embeddings for four datasets from the GADGET II TPC and AT-TPC. Randomly initialized encoders isolate the contribution from architecture before supervised training. We then train each encoder on a classification task, freeze its parameters, and train a linear or nonlinear probe for each downstream task. We find that this architecture-induced structure remains useful across experiments and detector systems. The randomly initialized PointNet-style representation is highly informative on several tasks. The two architectures organize their embedding spaces differently, but neither exhibits a large, systematic loss of utility cross-detector. These results show that architecture is a major source of task-relevant structure in TPC embeddings and should be treated explicitly when assessing representation learning and developing reusable detector models.

cs.LG

A simple derivation of the Kalman filter

In this lecture note, we present a concise and self-contained derivation of the discrete-time Kalman filter equations that requires only a basic understanding of least squares estimation. The treatment is designed to minimize mathematical overhead while preserving both rigor and generality.

math.OC

Jointly Satisfying Pareto Optimality and Justified Representation is NP-Hard in Approval-Based Multiwinner Voting

An open problem in approval-based multiwinner voting concerns whether we can efficiently compute committees that satisfy both justified representation and Pareto optimality. We answer this question negatively by proving that, on the domain of all profiles, outputting a committee satisfying both axioms is NP-hard. An initial proof was found by ChatGPT Astra. This was then verified and rewritten by the author.

cs.GT

On the Sequential Test and Distributed Detection

We present a simple definition of stopping time and its role in the formulation of sequential tests for both centralized and distributed detection, providing a straightforward procedure for obtaining optimal decision rules. Upper bounds for optimal stopping time are derived and numerically shown to possess certain qualitative features expected of the optimal stopping time. The results are extended to any distributed detection network in the form of an acyclic directed graph.

cs.IT

On the Complexity of Bayesian Signal Processing

We develop a computational framework for Bayesian decision-making. We show that as long as no action is optimal in every state, Bayes-optimal choice is intractable. This hardness need not arise from large action, state, or signal spaces, nor from a complicated represented utility function: extracting enough information from a hard-to-interpret signal to act optimally can itself be computationally hard. We also characterize tractability across approximation notions and identify their sources of difficulty. Under the probably approximately correct criterion, sample-based Bayesian learning is tractable if and only if the signal support is bounded. Our results provide justifications for bounded rationality, costly Bayesian inference, and sample-based Bayesian learning.

econ.TH

A 1.283 Price-of-Anarchy Bound for the Repeated Virtual First-Price Auction

We study the repeated allocation of a single indivisible resource among $n$ strategic players. Each player $i$ has a privately known value distribution $D_i$, and values are drawn independently across players and periods. The goal is to find fair and efficient mechanisms. We apply the repeated first-price auction with equal initial endowments of virtual money. We show that each player can asymptotically secure the same fair-floor guarantee $f(D_i)$ as in Csóka 2026; consequently, the mechanism is $1.283$-optimal. This provides a simpler and more robust alternative mechanism for this special case and may also help derive sharper upper bounds on the price of anarchy.

cs.GT

Sharp mean-field analysis of permutation mixtures and permutation-invariant decisions

We develop sharp bounds on the statistical distance between high-dimensional permutation mixtures and their i.i.d. counterparts. Our approach establishes a new geometric link between the spectrum of a complex channel overlap matrix and the information geometry of the channel, yielding tight dimension-independent bounds that close gaps left by previous work. Within this geometric framework, we also derive dimension-dependent bounds that uncover phase transitions in dimensionality for Gaussian and Poisson families. Applied to compound decision problems, this refined control of permutation mixtures enables sharper mean-field analyses of permutation-invariant decision rules, yielding strong non-asymptotic equivalence results between two notions of compound regret in Gaussian and Poisson models.

math.ST

A Complete Characterization of Tensorizable $f$-divergences

Csiszar's formulation of the $f$-divergence introduced a vast family of functionals for quantifying dissimilarity between probability distributions. However, many applications in statistics and information theory rely only on a few $f$-divergences, such as the Kullback-Leibler divergence, the $χ^2$-divergence, and the squared Hellinger distance. These divergences are especially useful because they admit simple compositional formulas under product measures, a property sometimes referred to as tensorization. In this work, we refine a formalism of tensorization previously introduced in the literature. Then, we show that any possible tensorization formula has a multi-affine form characterized by a single parameter, and identify all tensorizable $f$-divergences under our adopted notion of tensorization.

cs.IT

Course design in the age of AI

I develop a model of learning-by-doing and course design, and use it to study the impacts of artificial intelligence (AI). A myopic student faces a sequence of tasks that he can work on or delegate to AI. Work requires costly effort but builds skill; delegation requires no effort but builds no skill. A teacher designs the task sequence ("course") to maximize the student's skill development, given his choices to work or delegate. Without AI, the teacher makes earlier tasks more effort-intensive and later tasks more skill-intensive. With AI, the teacher must redesign early tasks to induce effort, leading to less skill development. If AI complements effort, then improvements in AI quality make high-skill students learn faster but low-skill students learn slower.

econ.TH

Accelerated High-Accuracy Sampling from a Warm Start via the Proximal Bouncy Particle Sampler

We study the problem of sampling from $μ(\mathrm{d}x)\propto e^{-V(x)}\,\mathrm{d}x$ on $\mathbb{R}^d$, where $V$ is $α$-strongly convex and $β$-smooth, and write $κ:=β/α$. We design and analyze the Proximal Bouncy Particle Sampler (Proximal BPS), a new sampler that combines ideas from the proximal sampler and the bouncy particle sampler. From a warm start initialization with $ O(1) $ Rényi divergence w.r.t. $μ$, Proximal BPS returns a sample whose law is $\varepsilon$-close to $μ$ in total variation distance using $\widetilde O(\sqrtκ\,d^{1/4} \,\mathrm{polylog}(1/\varepsilon))$ gradient queries in expectation.

math.ST

Tensor network representations of discrete maximum entropy distributions via mean polytopes

We present tensor network representations for discrete maximum entropy distributions under expectation constraints. To this end, we introduce Computation-Activation Networks (CompActNets), a tensor network architecture that subsumes exponential families. By leveraging the geometry of the convex polytope of realizable expectation vectors, we represent any maximum entropy distribution in the same architecture. We exploit the fact that proper faces of this polytope correspond to the boundary closure of exponential families, which restricts the distribution's support. We then derive explicit representations for the support within the CompActNet architecture. The proposed framework suggests tensor network ranks as complexity measures for faces. Finally, a case study on Boolean statistics links the geometry of 0/1-polytopes directly to propositional formulas.

math.ST

The Silent Distortion of Relative Prices

This paper presents a production network model of inflationary dynamics in which inflation can have near-zero correlation with the size of price change yet generate significant distortions in relative prices. New money enters unevenly across firms and percolates through buyer-seller links, thereby displacing relative prices even when every price is free to adjust and every market clears. The magnitude of price distortion depends critically on the spectral gap of the production network, which sets the rate at which the economy converges to equilibrium. We quantify the mechanism on a reconstructed production network with the universe of firms in the United States. At 1% inflation, the typical firm's relative-price distortion is of the same magnitude as the inflation rate. The size of price change remains essentially uncorrelated with the magnitude of relative price distortion.

econ.TH

From the Social Choice Problem to a Collusion-Proof Tendering Mechanism for Dynamic Stochastic Projects

The VCG family and the AGV mechanism are two classical approaches to efficient implementation in the static social choice problem. In 2024, Csóka et al. showed that AGV has critical weaknesses. In contrast, the transferable-utility Guaranteed Utility Mechanism (TU-GUM) retains all the standard desirable properties of AGV while adding further ones, including collusion-proofness, because it implements efficiency in Guaranteed Utility Equilibrium. TU-GUM also applies to a more general dynamic setting with multiple extensions. Moreover, TU-GUM is a special case of an even more general and robust mechanism that combines contingent first-price tendering with the coordinated execution of dynamic stochastic multi-agent projects through a surprisingly simple rule. This paper summarizes and connects existing results from a different perspective, with some minor new observations.

econ.TH

Mechanism Design for Alignment and Control

We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unknown. We want such agents to act on our behalf so mechanisms must incentivize both honesty and obedience. A one-sided imitation structure---capabilities can be concealed but not counterfeited---yields a revelation principle, a characterization of implementable policies via nested cyclical monotonicity, and conditions under which eliciting higher-order beliefs can discipline multiple agents. We apply our framework to stylized examples of (i) sandbagging in which a more capable agent pretends to be less capable; (ii) an alignment--interpretability trade-off, where the two are substitutes in the instrument but complements in value; (iii) discipline via peer scoring; (iv) coupling rewards to induce competition among multiple agents; and (v) scalable oversight and reward shaping.

econ.TH

Bayesian Adversarial Privacy

Theoretical and applied research into privacy encompasses an incredibly broad swathe of differing approaches, emphases and aims. This work introduces a novel quantitative notion of privacy that is both contextual and specific. Building on and extending ideas from statistical disclosure control and differential privacy, our aim is to model the implications of a disclosure decision in an adversarial setting. Our definition relies on concepts inherent to standard Bayesian decision theory, while departing from them in several important respects. In particular, (i) inference about the data itself becomes meaningful and (ii) the party controlling the release of sensitive information should make disclosure decisions from the prior viewpoint, rather than conditional on the data, which is a feature shared with Bayesian design. Illuminating toy examples are exploited towards highlighting the specificities of the method.

math.ST

Generalization Error Curves for Analytic Spectral Algorithms under Power-law Decay

The generalization error curve of certain kernel regression method aims at determining the exact order of generalization error with various source condition, noise level and choice of the regularization parameter rather than the minimax rate. In this work, under mild assumptions, we rigorously provide a full characterization of the generalization error curves of the kernel gradient descent method (and a large class of analytic spectral algorithms) in kernel regression. Consequently, we could sharpen the near inconsistency of kernel interpolation and clarify the saturation effects of kernel regression algorithms with higher qualification, etc. Thanks to the neural tangent kernel theory, these results greatly improve our understanding of the generalization behavior of training the wide neural networks. A novel technical contribution, the analytic functional argument, might be of independent interest.

cs.LG