Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Nonlinear scalarization and thermodynamics in Einstein-Phantom scalar-Gauss-Bonnet theory

We study nonlinear scalarization of Schwarzschild black-holes for a phantom scalar coupled to the Gauss--Bonnet invariant. Three polynomial couplings satisfying $ζ''(0)=0$ are introduced to investigate its nonlinear coupling dependence. Since the phantom factor $ε$ disappears in the probe scalar equation, its nonlinear evolution is the same as in the canonical scalar theory. We construct nodeless probe solutions and fully backreacted two branches of lower and primary. At equal mass and equal temperature, respectively, the lower branches have higher Wald entropy and lower Helmholtz free energy than primary branch, but their heat capacities are negative. This implies that lower branch favors than primary branch. Finally, we show that there is no nonlinerly stable scalar phase for a phantom kinetic theory.

gr-qc↗

Trust the Critic More

Standard language model RL algorithms credit every token of a long rollout with the same advantage determined by the terminal reward. Actor-critic methods can provide finer-grained credit assignment, but learned critics are generally considered too inaccurate to trust when training LLMs with RL. In recent works, even when a critic is present, it is used only for baseline estimation, so every trajectory must be rolled out to its terminal reward. We introduce Actor-Critic with Action Chunking (AC2) that removes the need to roll every trajectory to completion. AC2 instead assigns credit to action chunks: short continuations of prefixes of past trajectories. A learned critic scores the state reached at the end of each action chunk, allowing the policy to update without observing a terminal reward. We make critic-based credit assignment reliable through three design choices. First, we introduce local readiness which uses critic-based updates on a problem only when the critic is sufficiently accurate on that particular problem. Second, when available, we provide the critic with a reference solution from a previous successful rollout. Third, we assign credit over action chunks of 10k tokens rather than individual tokens, giving the critic a more meaningful portion of the trajectory to evaluate. We train Qwen3-4B on FineProofs-RL using AC2 and evaluate on IMO-ProofBench. AC2 exceeds GRPO's peak validation score of 18.5% using 2.5x fewer decoding FLOPs. This gain comes from two sources, (1) AC2 requires 25% fewer training steps to reach this score, and (2) each step generates fewer tokens because the policy does not need to continue every trajectory to completion. Conceptually, we demonstrate that we can remove the need to roll out every trajectory to completion, opening up a large previously unexplored design space for LLM RL algorithms.

cs.LG↗

RW-Flow: One-Step Generation on Compact Manifolds via Wasserstein Gradient Flows

Manifold-valued data, and consequently the distributions they induce, are prevalent across many domains, ranging from the locations of geospatial events, such as earthquakes, to biomolecular torsion angles that encode information about three-dimensional structure. While diffusion and flow-based generative models have been successfully extended to compact manifolds, sampling typically requires tens or hundreds of sequential network evaluations. We introduce RW-Flow, a theoretically grounded framework for learning one-step generative models on compact manifolds via Wasserstein gradient flows. The main challenge is identifiability: driving the velocity field to zero should guarantee that the model distribution matches the target distribution. We establish a necessary and sufficient condition for identifiability on compact, connected Riemannian manifolds. We specifically show that, for a symmetric, Lipschitz-continuous cost function, the velocity field induced by the Sinkhorn divergence is identifiable if and only if the associated Gibbs kernel is nondegenerate. This characterization provides a general principle for designing identifiable costs on compact manifolds. It also reveals that the squared geodesic distance, the natural manifold analogue of the squared Euclidean distance, does not always guarantee identifiability. Across benchmarks involving geospatial events, protein side chain torsion angles, RNA backbone torsion angles, and general manifolds discretized as triangular meshes, RW-Flow outperforms existing one-step methods in nearly all settings under fair comparison conditions.

cs.LG↗

Bounds on adiabatic path geometry from the width class of the gap profile

In an adiabatic evolution, the wave function follows the path of ground states $|Φ_0(s)\rangle$ of a Hamiltonian $H(s)$ as $s$ is tuned from $0$ to $1$. The geometry of this adiabatic path is described by quantities such as its length $L=\int_0^1 \|\partial_s |Φ_0(s)\rangle\|\,ds$ and its total curvature $K$. This geometry affects the evolution time $T$. For instance, $T$ is at least of order $L/Δ_*$, where $Δ_*$ is the minimum energy gap over $s\in[0,1]$. The traditional approach gives $L=O(Δ_*^{-1/2})$ and $K=O(Δ_*^{-1})$, which are loose in many cases. We improve these bounds by using $μ_<(γ)$, the width of the region in $s$ on which the gap is below $γ$. The gap profile is in width class $p$ when $μ_<(γ)$ falls at least as fast as $γ^{1/p}$. A typical avoided crossing is in width class $p=1$, which gives $L=O(\sqrt{\log Δ_*^{-1}})$ and $K=O(\log Δ_*^{-1})$. A profile in width class $p>1$ gives power laws with exponents $(p-1)/(2p)$ for $L$ and $(p-1)/p$ for $K$. The width class also bounds the evolution time. At $p=1$ we have $T=O(Δ_*^{-2})$ with a schedule that advances $s$ at a constant rate, and $T=O(Δ_*^{-1})$ up to a polylogarithmic correction with a schedule that traverses the path at a constant geometric speed. The bounds on $L$ hold for any twice continuously differentiable $H(s)$, and the bounds on $K$ for affine $H(s)$. We also prove the tightness of scaling of $L$ in $Δ_*$ for every integer $p\ge1$, and the scaling of $K$ at $p=1$. We demonstrate the bounds on the adiabatic Grover search, the XXZ spin chain, and molecular electronic Hamiltonians.

quant-ph↗

The Margolis-Rhodes Monoid of a Graph

We investigate the structural, combinatorial, ideal-theoretic and Krohn-Rhodes complexity of the Margolis-Rhodes monoid, MR(G), of a finite simple graph G, viewed topologically as a 1-dimensional simplicial complex. Alongside the full monoid, we examine some subsemigroups including St(G), defined by the condition that the full inverse image is an edge or the empty set and Inj(G), the monoid of all partial 1-1 continuous functions. We provide explicit combinatorial enumerations and struture for paths and cycles. We compute Green's relations showing in particular that the partial order of regular J-classes is isomorphic to the poset of induced subgraphs of G. Finally, we apply these structural invariants to Krohn-Rhodes complexity theory. It is known that the Margolis-Rhodes monoid has complexity at most 2 and 1 for St(G) and Inj(G). We show that for cycles the complexity of its Margolis-Rhodes monoid is 2 if and only if the cycle is of length at least 4. For paths, we prove that the complexity of its Margolis-Rhodes monoid is 2 if the path length is at least 13.

math.GR↗

Logical Operator Decomposition for Distance Analysis of Bivariate Bicycle Codes

Bivariate bicycle (BB) quantum codes are a prominent finite-length family of quantum low-density parity-check codes, but their minimum distance is usually established numerically rather than read from the defining polynomials. We study the $Z$-logical quotient $K/S$ over $\mathbb F_2[x,y]/(x^\ell-1,y^m-1)$ and show that it fits into a short exact sequence with an annihilator quotient as kernel and a colon quotient as cokernel. The sequence gives an explicit logical basis, a dimension formula, and a componentwise distance identity $d_Z=\min(d_{\mathcal A},d_{\mathcal C})$. Using the Frobenius structure of the finite group algebra, we prove $r_{\mathcal A}=r_{\mathcal C}=k/2$ for every BB code, including repeated-root cases. The algebraic component of a logical class is distinct from the support shape of its lightest representatives: an annihilator class can have a lighter two-block representative, and a colon class can have a one-sided minimum. For lower bounds we show that every proper subset of a minimum-weight logical operator has nonzero syndrome, and that this property persists inside the colon component but not inside the annihilator component. A translation-anchored cluster search built on it proves the distances $4,6,10,10,12,18$ of the six standard BB codes of lengths $18$ to $288$ and enumerates every minimum-weight logical operator. The resulting census shows that $[[108,8,10]]$ is the only one of the six whose distance is attained in a single component, with $d_{\mathcal C}=10$ and $d_{\mathcal A}=12$.

quant-ph↗

Measurement of the $^{39}$Ar specific activity in atmospheric argon with DArT at the Canfranc Underground Laboratory

We report the measurement of the $^{39}$Ar specific activity of atmospheric argon using DArT, a low-background single-phase liquid-argon detector read out by cryogenic silicon photomultipliers, at the Canfranc Underground Laboratory (LSC) in Spain. DArT was first filled with atmospheric argon and then with a sample of underground argon of known radioactivity from the DarkSide-50 experiment. The underground argon serves as a background reference in this measurement, making the result robust because we do not rely on background simulations. We measure the $^{39}$Ar specific activity with both a cut-and-count method and a binned maximum-likelihood fit, yielding consistent results. We obtain $a_{\text{AAr}}= 0.955 \pm 0.008~\text{Bq}/\text{kg}$. This result agrees with previous measurements from other experiments and provides the most precise determination of the $^{39}$Ar specific activity in atmospheric argon to date. It also validates DArT as a key component of the DArTInArDM experiment at LSC, which aims to measure the $^{39}$Ar activity of the argon extracted from deep underground wells in Colorado (USA) for the Darkside-20k and LEGEND-1000 experiments at Laboratori Nazionali del Gran Sasso in Italy.

physics.ins-det↗

Recommendation Systems for Exploratory Data Tasks

A large class of data-centric tasks is exploratory, where users iteratively steer workflows, refining subjective goals as new insights emerge. These Exploratory Data Tasks (EDTs) are performed by millions of users with varying levels of expertise to understand unfamiliar data, discover trends, and identify evidence that informs critical decision-making. However, a key challenge in EDTs is the enormous space of possible actions that one can take at each step: users struggle to choose among thousands of joins, transformations, and aggregations, causing "exploration paralysis". Because EDT workflows are interconnected, each choice impacts subsequent exploration, and suboptimal choices can lead to inefficiency, missed insights, confirmation bias, and incomplete coverage. This calls for intelligent recommendations that efficiently guide users toward optimal EDT actions. We envision recommendation as a core capability of data systems, proactively guiding users toward promising actions and thereby lowering the barrier to exploratory data tasks. EDT recommendation is challenging because the action space is combinatorial and actions are data-dependent, which require costly materialization. Moreover, recommendation often involves bundles or sequences of actions across interdependent tasks, requiring coordination across tasks. In this paper, we present our vision of EDT recommendation systems along two axes: single-task vs. multi-task settings and single-action vs. multi-action recommendations. We outline a research agenda that progresses from recommending individual EDT actions to constrained bundles and sequences of actions, and ultimately to coordinated recommendations across interconnected EDTs. We identify research directions for incorporating various contexts (user, data, task, and ecosystem), addressing efficiency challenges, and coordinating across tasks.

cs.DB↗

Quantum Well Resonant Tunneling Diode Probe of Correlated States in Twisted Bilayer MoS$_2$

Moiré superlattices formed in transition metal dichalcogenides (TMDs) offer a versatile platform for exploring emergent quantum phases arising from strong electronic correlations. In this work, we develop a new experimental platform, the quantum well resonant tunneling diode (QWRTD), to probe the electronic landscape of $\approx 57^\circ$ twisted bilayer MoS$_2$ (tMoS$_2$). By measuring the differential conductance ($dI/dV_{\text{Probe}}$) as a function of filling factor $ν$ and displacement field $D$, we observe a robust integer correlated insulating state at $ν= 1$ that persists across the measured displacement field range. Furthermore, we identify a displacement-field-induced fractional insulating state at $ν= 3/4$ for $D < -75$~mV/nm. Temperature- and magnetic-field-dependent measurements characterize this $ν= 3/4$ state as a potential generalized Wigner crystal. These correlated states are also observed in another device with a similar twist angle ($\approx 56.5^\circ$). Our results provide direct evidence of correlated states in near-AB-stacked tMoS$_2$ and establish QWRTD as a powerful experimental tool for investigating strong correlations and topology in van der Waals heterostructures.

cond-mat.str-el↗

Early Planet Formation in Embedded Disks (eDisk). XXV. Inclination-Induced Minor-Axis Brightness Asymmetries Reveal Limited Dust Settling in Embedded Protostellar Disks

How and when dust settles in young protostellar disks is a key open question for the dust concentration needed to form planetesimals and, ultimately, planets. However, directly measuring the vertical dust distribution in embedded (Class 0/I) systems remains challenging. We show that brightness asymmetry along the minor axis of highly inclined disks provides a simple, powerful geometric diagnostic of vertical dust structure. Using radiative transfer modeling with RADMC-3D, we generate synthetic continuum images showing that the observed asymmetry arises naturally from disk inclination, optical depth, and dust scale height. We apply this framework to nine Class 0 and I disks from the ALMA Large Program, Early Planet Formation in Embedded Disks (eDisk), using Markov Chain Monte Carlo (MCMC) fitting. Outflow observations independently validate the inferred near- and far-side geometries: all eight sources with useful outflow constraints agree with the orientations predicted by the dust continuum modeling. Our results indicate that most embedded disks show no strong evidence of significant dust settling, with dust scale heights comparable to the gas scale height. Given that the literature indicates Class II disks tend to be well settled, our results reinforce the notion that significant dust settling occurs during the Class I phase, when deeply embedded Class 0 disks transition to their more revealed Class II counterparts. Intriguingly, the timing of dust settling appears to broadly coincide with the development of widespread dust substructures, suggesting that gravitationally driven vertical dust concentration may have triggered substructure formation.

astro-ph.EP↗

CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL

During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigating reward hacking during training remain challenging, which is limited by a lack of testbeds that reproduce hacking and reliably identify it. We introduce CATCH, a controllable testbed for studying reward hacking in coding RL. CATCH deliberately exposes environmental loopholes and provides execution-based gold labels by comparing success under a vulnerable evaluator with task correctness under an independent audit. It also can control the model's initial hacking tendency through supervised fine-tuning data mixtures and the difficulty of earning rewards through reward designing, enabling systematic comparisons of hacking dynamics and interventions. Experiments show that CATCH can produce diverse RL training trajectories with clear reward hacking, and analyses demonstrate that both initial models and reward difficulties shape the emergence of reward hacking. We further evaluate the effectiveness of different reward hacking detection and mitigation methods. A key finding is that a chain-of-thought monitor initially suppresses hacking, but this protection erodes as the policy model learn to mislead the monitor with code comments. This highlights the need to evaluate hacking mitigations throughout training with CATCH. The source code and resources are publicly released at https://github.com/THUAIS-Lab/CATCH.

cs.CL↗

Growing an Agent/Prover Interface: Evolutionary Tool Design for Cost-Efficient Theorem Proving in Rocq and Lean

Recent achievements in AI-assisted mathematics require intensive interaction of agents with proof assistants to generate machine-checked proof certificates. Agents interact with proof assistants such as Rocq or Lean through an interface that controls what the agent receives from the prover and the cost of these interactions. Today, these interfaces are adapted from tools designed for humans and not optimized for agents. We propose an evolutionary method where a frontier model incrementally proposes new features and only keeps the ones that improve the overall performance of smaller models. We demonstrate the effectiveness of our method by growing, on a curated set of mathematical problems, \rme, a new MCP server for the Rocq prover. On the held-out \texttt{test} split of miniF2F-Rocq, an agent equipped with \rme outperforms both the baseline that only exposes the Rocq compiler and an established MCP server, across four models from two families, in success rate, cost per solve, and time per solve. Although evolved for Rocq, the resulting server transfers to Lean, improving cost and time per solve on a subset of PutnamBench. We release \rme and its port to Lean.

cs.AI↗

Complete total-transmission modes of Kerr black holes

We construct the complete spectrum of gravitational total-transmission modes (TTMs) of Kerr black holes and find four globally continuous families, $n_\infty=1,2,3,4$, each with two complex-conjugate frequency branches. Compared with the three-family classification of Cook and Lu, our global continuation resolves the mirror-related sectors of their $n=2$ family into separate families without introducing additional symmetry-unrelated roots. This organization distinguishes complex conjugation at fixed $m$ from the mirror symmetry connecting the $m$ and $-m$ spectra. The $n_\infty=3$ family approaches the Schwarzschild algebraically special frequency, whereas the $n_\infty=1,2,$ and $4$ families diverge as $ω\propto a^{-4/3}$ along lower-half-plane directions $-150^\circ$, $-90^\circ$, and $-30^\circ$, respectively. High-precision data up to $\ell=32$ confirm the Cook--Lu small-spin asymptotics and reveal how the divergent branches are embedded in the global four-family spectrum, with a family-dependent relation between the spherical-limit angular labels and the large-$|aω|$ ordering. For axisymmetric perturbations, the $n_\infty=2$ and $n_\infty=3$ branches coalesce at genuine exceptional points, where the scattering spectral function develops a second-order zero and the TTM excitation factors are strongly enhanced, offering a different interpretation of the structure previously described by Cook and Lu as an overtone-multiplet splitting. Finally, we identify a previously unreported anomalous proximity between the $n_\infty=3$ TTM family and an unconventional Kerr quasinormal-mode sequence, with frequency separations reaching order $10^{-8}$ in the representative case studied.

gr-qc↗

RiboUnmix: Learning Shared Translational Dynamics from Biased and Noisy Ribo-seq Measurements

Ribosome profiling (Ribo-seq) measures ribosome distributions along mRNAs, but observed occupancy profiles also contain experiment-specific distortions and stochastic variability. Consequently, models that accurately predict measured profiles may reproduce technical effects rather than recover the underlying biology. We ask whether jointly modeling datasets collected under different experimental conditions can reveal shared, sequence-dependent patterns of ribosome occupancy. We introduce RiboUnmix, a probabilistic multi-dataset framework in which each expected measured profile is represented as a shared sequence-dependent signal modulated by a dataset-specific multiplicative factor. A negative-binomial observation model captures variability across replicates. We evaluate RiboUnmix on a controlled synthetic benchmark combining programmed translation kinetics, ribosome traffic, stochastic count sampling, and sequence-dependent experimental distortions. Because the underlying kinetics and distortions are known, recovery of the shared profile and dataset-specific effects can be assessed separately. Both inferred components correlate strongly with their targets, demonstrating that RiboUnmix can disentangle shared kinetic patterns from experimental effects. Across four organism-specific real-data benchmarks, RiboUnmix outperforms sequence-to-profile baselines in predicting measured profiles. Models trained independently on subsets of 114 HEK-derived datasets recover concordant shared profiles for held-out transcripts, and experiments varying the number and composition of training datasets show that the learned representation remains stable. RiboUnmix thus converts variation across experiments into evidence for reproducible sequence-dependent patterns of ribosome occupancy, supporting biological hypothesis generation from diverse Ribo-seq datasets.

cs.LG↗

Phase-sensitive avalanche quantum sensing of sub-shot-noise fields

Avalanche-based detection of small perturbations is commonplace in precision measurement devices from Geiger counters to single-photon avalanche detectors. Here, we expand this principle to sensing of light waves below the quantum shot noise limit, demonstrating numerically how atomic clouds trapped in photonic cavities can exhibit non-perturbative sensitivity to changes in the cavity field when the cavity mode is tuned to a parity-prohibited transition. We find that, when pumped by a continuous-wave laser, the atomic cloud creates emissions which amplify the initially sub-shot-noise fluctuation by two orders of magnitude. Crucially, this amplification is broadly tunable with respect to frequency and independent of atomic energy structure. Even more strikingly, our amplification mechanism preserves the phase imparted on the atoms by the original ultra-weak light wave, marking a qualitative improvement over conventional protocols. Our finding has implications for both quantum state characterization and detection of weak classical signals.

quant-ph↗

A depolarizing choir sings in Gaussian harmony

We study the noise threshold for positive quantum capacity for the qubit depolarizing channel. We explore analytically the action of the qubit depolarizing channel on the symmetric subspaces of the input qubits, in the limit of asymptotically many uses of the channel. We observe the emergence of a bosonic Gaussian channel. Furthermore, the codes previously developed for the depolarizing channel can be translated to codes for the emergent Gaussian channel, and it is easier to further optimize these codes for the simpler emergent Gaussian channel. Translating these codes back to the depolarizing channel leads to extremely good input states for the coherent information of the depolarizing channel producing new lower bounds on the noise threshold for positive capacity. In addition to improved lower bounds on the threshold, this newly found link between depolarizing noise and Gaussian channels offers a novel perspective contributing to our understanding of these symmetric codes.

quant-ph↗

Safety of Latent Communication in Multi-Agent Systems

Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole. Code: https://github.com/Muhammad-Huzaifaa/latent-safety

cs.AI↗

TopTimeNet: Topologically-assisted time-series classification model

Distinguishing periodic from chaotic dynamics in a time series is a fundamental challenge in both physics and engineering. Yet, end-to-end learned architectures must discover both a representation and a decision boundary from data, at substantial cost. We introduce TopTimeNet, which decouples these tasks: a fixed, non-learned stage extracts a $42$-dimensional geometric and topological descriptor from Takens delay embeddings and persistent homology, and a lightweight learnable stage performs classification. On a benchmark of $49$ nonlinear dynamical systems, a $1{,}638$-parameter configuration matches the mean accuracy of one with $33\times$ more trainable parameters. Additionally, this approach delivers mean accuracy comparable to convolutional neural networks and surpasses the average performance of converged Transformer models, while requiring three to four orders of magnitude fewer trainable parameters. Robustness also depends sharply on where noise is introduced: TopTimeNet degrades gracefully under perturbations to its precomputed features, but degrades sharply when noise is introduced into the raw signal and the full feature-extraction pipeline is recomputed, showing that robustness to perturbations of the precomputed features does not imply robustness of the complete raw-signal-to-prediction pipeline. These results show that decoupling fixed geometric and topological feature construction from a lightweight discriminative stage can achieve comparable classification accuracy with substantially fewer trainable parameters.

cs.LG↗