Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Two-Fidelity Best-Action Identification for Stochastic Minimax Tree

We study fixed-confidence best-action identification (BAI) in stochastic minimax trees. This problem is increasingly relevant in modern AI planning, where deep minimax search and Monte Carlo Tree Search (MCTS) with language model long rollouts face a fundamental tradeoff: heuristic evaluations are cheap but biased, while accurate rollouts are reliable but prohibitively expensive. We propose 2FFS, a two-fidelity tree-search algorithm that brings multi-fidelity flat bandit ideas into trees. The algorithm combines minimax-style fast expansion with MCTS-style stochastic sampling, adaptively deciding when to exploit cheap biased evaluations and when to invoke expensive accurate evaluations for local certification. We prove fixed-confidence correctness, establish finite stopping for exact identification, and give a polynomial-depth cost upper bound for general-depth trees. Across numerical stochastic-tree experiments, 2FFS uses substantially fewer samples and computational operations comparing to existing BAI-MCTS baseline.

cs.LG↗

How Optimality Structures Sparse Dictionaries: Theory for Interpreting SAE Representations

Sparse Autoencoders (SAEs) have found success parsing neural network representations into interpretable concepts, providing a basis for understanding and control. However, what exactly SAEs extract and, hence, the scientific conclusions we can draw from them are not obvious. In short, if your SAE behaves strangely, does that reflect interesting neural network behaviour or an SAE-imposed distortion? Towards answering this, we use dictionary learning identifiability results to derive constraints that optimal dictionary learning features must satisfy. For example, an optimal feature will never turn on only while another is active. We use these conditions to explain various SAE oddities - hierarchical splitting & absorption, which features can be left in the residuals, dense antipodal features, and infinite feature splitting - simply as properties imposed by the dictionary learning objective. Finally, these constraints are diagnostic: real SAEs pass when measured on the dataset on which they were trained, but increasingly fail as the test dataset becomes more `distant'. In sum, we hope to provide theoretical tools to explain puzzling SAE patterns, allowing more principled inferences about internal model behaviour.

q-bio.NC↗

Poking Around in the Dark: Why a Shared Understanding of Components Matters

By listing the components included in an application, Software Bills of Materials (SBOMs) are intended to support the timely identification of vulnerable components and ensure the security of the software supply chain. However, we question the underlying assumption that there is agreement on the components to be listed in an SBOM and that current technology is sufficient to secure the software supply chain. First, we propose a ground-up analysis of Component Inclusion Mechanisms (CIM) in the software's development lifecycle. Then we systematically analyze the four popular SBOM generation tools, cdxgen, syft, trivy, ORT, and the Microsoft sbom-tool, to understand how they define and identify relevant components. Finally, we assess these using a ground truth across the programming languages Python, Java, Go, PHP, Rust, and C. While today's tools are a step toward identifying components, our results show that no tool covers all identified CIMs and that common gaps exist across tools. We demonstrate that, under the current vague definitions and tooling, SBOMs exhibit ambiguity and blind spots in component inclusion. Thus, a security-grade SBOM is not achievable with the evaluated tools, necessitating further progress to ensure software supply chain security. We need to go back to the drawing board to clarify which components should be included in an SBOM and revise SBOM generators accordingly. Without a shared understanding of what a component is, any effort to secure software supply chains with SBOMs will fail.

cs.SE↗

The Radio-FIR Correlation in the Context of Deep Radio Source Counts

Increasingly deep, confusion-limited radio surveys have pushed direct radio source-count measurements down to tens of $μ$Jy at 1.4 GHz. Confusion-noise P(D) analyses extend the statistical counts to below $1\,\mathrm{μJy}$. Radio source counts have allowed for constraints on the radio-derived star formation rate density (SFRD) history through models of the backwards evolution of the local radio luminosity function, using the radio-FIR correlation, $q \propto \log(L_{\mathrm{FIR}}/L_{1.4})$, to convert radio luminosities to FIR luminosities and hence star-formation rates. Recent deep radio source counts from MeerKAT suggest a potential tension in the SFRD history between radio and UV/IR measurements at $1\lesssim z\lesssim 2$. This corresponds to a $\gtrsim 3σ$ discrepancy between the predicted and measured source counts near the star forming galaxy source count $S^2n(S)$ peak of ${\sim}30\,\mathrm{μJy}$ under both a pure luminosity (PLE) and combined luminosity and density evolution (LADE). We consider what the requirement of agreement between radio source counts and the observed UV/IR SFRD indicates about the redshift evolution of the radio-FIR correlation and its intrinsic scatter out to $z=3$. We introduce a radio-luminosity based parameterization to $q_{\mathrm{FIR}}(z)$ based on changing thermal radio fractions alone that agrees with the observed stellar-mass dependent $q_{\mathrm{FIR}}(z)$ better than a non-evolving or decreasing $q_{\mathrm{FIR}}(z)$. Despite this, we find that a decreasing $q_{\mathrm{FIR}}(z)$ at fixed radio luminosity provides better agreement between source counts and the observed SFRD, while a $q_{\mathrm{FIR}}(z)$ that breaks down due to cosmic ray losses requires an intrinsic scatter up to $σ_q\approx 0.3\,\mathrm{dex}$.

astro-ph.GA↗

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR

Reinforcement learning with verifiable rewards (RLVR) improves the ability of large language model, yet headline accuracy gains often conceal a hidden cost: previously solved problems quietly become unsolvable as training proceeds. We frame this phenomenon as \emph{correct-set turnover}, representing the coupled dynamics of solution acquisition and regression over the mastered set. Under this view, retention becomes an explicit optimization target alongside acquisition. We analytically and empirically establish the \emph{repair-window principle}: the cost of restoring a regressed prompt grows sharply with review delay, defining a low-cost window that standard RLVR pipelines fail to exploit. To address this, we propose \textbf{\method{}}, a retention-aware review mechanism that tracks mastered prompts and periodically reintroduces them to \textbf{remind} the model of previous solutions. By utilizing pre-rollout batch replacement, \method{} incurs zero additional rollout overhead. Evaluated across 20 benchmarks spanning image-text, video, and text-only tasks with Qwen3-VL and Qwen2.5-Math, \method{} consistently improves performance over GRPO, DAPO, and replay baselines, demonstrating robust generalizability across modalities and algorithms.

cs.LG↗

RobotValues: Evaluating Household Robots When Human Values Conflict

While household robots are often evaluated based on task completion, everyday domestic environments involve value-conflicting situations where robots are expected to choose actions that prioritize diverse values such as human autonomy, efficiency, or social appropriateness. Yet, there are no benchmarks for evaluating robots' value preferences in such scenarios. We introduce RobotValues, a benchmark to evaluate household robot planners in 8K value-conflict scenarios. Each instance consists of a realistic, synthetically generated household image with multiple plausible robot actions that prioritize different human values. We construct ROBOTVALUES through LLM-assisted scenario generation, stakeholder-grounded value extraction, image generation and automatic quality control. We evaluate 10 VLMs used in robotics and find that models show default value preferences, including safety and accommodation, while underselecting privacy-prioritizing actions. When models are prompted to prioritize values that conflict with their preferences, models often fail to override the default actions, choosing incorrect actions 80% of the time on average across models. These findings highlight the need to go beyond task completion or safety evaluations and assess robots' decision-making capability when human values conflict.

cs.RO↗

On the finite field spherical restriction conjecture in four dimensions: the sharp endpoint and applications

Let $p$ be an odd prime. We prove the sharp extension estimate $R_{S_j}^*(2\to r)\lesssim_r 1$ for every sphere of nonzero radius $S_j\subseteq\mathbb{F}_p^4$ and every $r\geq3$, uniformly in $p$ and $j$. The proof combines arithmetic Hecke operator bounds with a refined fourth moment and an orthogonal decomposition. We also formulate a localized spherical restriction/extension conjecture that predicts the sharp dependence on the size of the physical support. This conjecture remains open and would yield distance estimates for almost every pin at the conjectured Erdős--Falconer exponent in four dimensions, with an arbitrarily small power loss in the size hypothesis. Unconditionally, our localized estimates imply that, for every $\varepsilon>0$, sets of size at least $p^{7/3+\varepsilon}$ determine $(1+o_\varepsilon(1))p$ distances from almost every pin, together with asymmetric two-set versions.

math.CA↗

A Common Antimatter Response in AMS-02 Positrons and Antiprotons

I develop a common history-overlap description that gives antiparticle time orientation a physical role beyond its use in field-theory amplitudes. In this framework, opposite orientation reduces the effective overlap with a finite, matter-defined environmental history, while standard local interactions, positive physical energies, and causal detector records are retained. The common geometry connects two contrasting cosmic-ray spectral patterns measured by AMS-02. For positrons, a geometric exposure estimate displaces the broad maximum of an empirical energy-weighted electron reference from 8--10 GeV to the 300--400 GeV region. A number-conserving calculation combines energy redistribution and spatial dilution to interpret the relative weighted height, of order $0.1$. For antiprotons, ordinary secondary production retains the parent-proton history. I extend the same three-direction overlap estimate to the accumulated logarithmic deformation after production. Within this approximation, an ordinary softening of a few tenths leaves a residual additional spectral index of order 0.01, yielding near preservation of the source index and a nearly flat antiproton-to-proton ratio in the production-scaling limit. The positron comparison uses no additional dominant source, and known astrophysical contributions remain allowed. The paired AMS spectra give this time-orientation framework concrete empirical direction: a common geometric description relates a lepton scale displacement to hadronic index preservation and defines explicit targets for independent theoretical and experimental tests.

hep-ph↗

Temporal Invariance Is an Illusion: Time-Dependent Influences of the Galactic Magnetic Field on UHECR Observations

Understanding the origin of the Ultra-High-Energy Cosmic Rays (UHECRs) requires explaining the features of their energy spectrum, mass composition, and arrival directions. Current modeling approaches neglect the time evolution of UHECR observables, a factor that is particularly important in the case of bursting UHECR sources. This study focuses on the influence of time delays caused by the galactic magnetic field (GMF) on the spectrum and arrival directions of UHECRs observed on Earth. Using CRPropa 3.2, we investigate the rigidity-dependence of the residence time of extragalactic cosmic rays entering our Galaxy. We find that UHECRs entering the Milky Way can experience delays of hundreds of kiloyears relative to light, and we demonstrate that these delays significantly alter the UHECR observables. Notably, a cutoff emerges in the transient scenario within the rigidity range of $10^{18}-10^{19}$ V, which coincides with the spectral break observed in data. We find a progressive shift in composition favoring heavier nuclei, as well as a delay distribution that is correlated with GMF strength. This causes the particles to be less correlated with their initial direction the larger their delays. A dipole-like anisotropy develops over timescales of about $\sim$100 kyr in certain bursts scenarios. Our results provide an alternative explanation for the UHECR spectral cutoff that does not invoke limits on source acceleration. This could potentially revise existing constraints.

astro-ph.HE↗

Voronoi-Hankel Transforms

Let $π$ be a generic irreducible representation of $\mathrm{GL}_n(\mathbf{F})$ for a local field $\mathbf{F}$. We introduce the Voronoi--Hankel transform $\mathcal{VH}_π$ for $π$, and prove it is equivalent to the $π$-Fourier transform introduced by Jiang--Luo. We give an extension $\widetilde{\mathcal{VH}}_π$ of $\mathcal{VH}_π$ to some larger class of functions on $\mathbf{F}^{\times}$, and prove the multiplicativity with respect to parabolic induction via the above-mentioned equivalence. For non-archimedean $\mathbf{F}$ we give a systematical study of the relevant kernel function, including the asymptotic behavior at $0$ and $\infty$, as well as some simple integral representation based on the local Langlands correspondences for the essentially tame supercuspidals. As an application in the non-archimedean case, we give an effective version of the stability theorem for the twisted local gamma factors.

math.NT↗

Beyond Missing Rates: Rethinking Incomplete Multi-View Clustering with Protocol Divergence

Incomplete multi-view clustering (IMVC) is typically evaluated by retraining separate models under different missing-view configurations. Evaluations indexed only by nominal missing rate can overlook differences in observation structure across missing-view protocols. We show that missing-data protocols with identical nominal missing rates can induce substantially different learning regimes, differing by approximately 50-fold in the proportion of fully observed samples. We formalize this phenomenon as protocol divergence, which quantifies structural disparities among missing-view protocols beyond marginal missing rates. Furthermore, we analyze support-gated reconstruction mechanisms and show that their optimization contribution is inherently limited by the frequency of eligible observations under explicit normalization and optimization conditions. Based on these observations, we propose CRAFT (Co-occurrence-free Robust Attention-masked Fusion Transformer), a train-once framework that combines representation learning with an architecture designed to process missing-view inputs. CRAFT combines (i) per-sample forward computation using each sample's observed views and shared parameters, and (ii) mask-aware fusion over nonempty observed-view subsets. The deployment evaluation starts from training data with all views available and reuses one final checkpoint per dataset and seed across missing-view protocols without retraining. Experiments on CUB and MultiFashion show that CRAFT achieves the strongest performance in 12 out of 13 information-matched settings. Additional deployment experiments across seven benchmarks and sixteen missing configurations demonstrate substantial computational savings through checkpoint reuse while maintaining competitive clustering performance. Code and evaluation tools: https://github.com/dk23lhl/CRAFT and https://github.com/dk23lhl/imvc-audit.

cs.LG↗

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe training outcomes. In real-world rubric-based RL, such hacking behaviors are often subtle and entangled with multiple judge biases, making them difficult to analyze, detect, and mitigate. In this paper, we introduce CHERRL, a Controllable Hacking Environment for Rubric-based RL. By injecting known biases into LaaJ, CHERRL enables stable reproduction of reward hacking, explicit observation of reward divergence, and identification of hacking onset. This provides a clean experimental testbed for studying the mechanisms and mitigations of reward hacking in rubric-based RL. To demonstrate its utility, we analyze different judge biases from the perspectives of discoverability and exploitability, and explore an agent for automatically detecting reward hacking onset from training logs. The code and environment are publicly available at https://github.com/THUAIS-Lab/CHERRL.

cs.LG↗

HVPE Growth of Si-Doped $β$-Ga$_2$O$_3$ on Sapphire: Influence of Substrate Offcut on Structural and Electrical Properties

Si-doped $β$-Ga$_2$O$_3$ films were heteroepitaxially grown on sapphire substrates using HVPE. The influence of sapphire offcut on growth kinetics, surface morphology, crystalline quality, and electrical transport properties was systematically investigated. Growth kinetics studies revealed a strong dependence of deposition rate on HCl flow, growth pressure, and source-to-substrate distance, with growth rates reaching up to 30 $μ$m/hr. Increasing sapphire offcut angle from 0$^\circ$ to 8$^\circ$ promoted a transition from multidirectional growth to highly aligned terrace-dominated surfaces, reducing the surface roughness from 14.69 to 2.74 nm. The improved surface morphology was accompanied by enhanced crystalline quality, with phase-pure (-201)-oriented $β$-Ga$_2$O$_3$ growth and a reduction in the rocking-curve full width at half maximum from 994 to 414 arcsec as the sapphire offcut increased. Electrical characterization of films grown on 6$^\circ$ offcut substrates yielded carrier concentrations ranging from $1.0\times10^{17}$ to $3.4\times10^{18}$ cm$^{-3}$. A maximum room-temperature electron mobility of 100cm$^2$/V$\cdot$s was achieved at a carrier concentration of $1.0\times10^{17}$cm$^{-3}$, representing the highest reported room-temperature mobility for HVPE-grown $β$-Ga$_2$O$_3$ on a foreign substrate. Analysis of the temperature-dependent transport characteristics yielded donor activation energies of 35 and 90 meV together with a low acceptor concentration of $3\times10^{15}$ cm$^{-3}$, consistent with the improved crystalline quality achieved on the offcut sapphire substrates. These results demonstrate that HVPE is capable of producing high-quality $β$-Ga$_2$O$_3$ heteroepitaxial layers with good crystalline quality and carrier transport characteristics, providing a promising pathway for scalable $β$-Ga$_2$O$_3$ epitaxy on low-cost foreign substrates.

physics.app-ph↗

Bentkus-type asymptotic e-values

Asymptotic e-values are emerging as a powerful alternative to asymptotic p-values, particularly in post-hoc inference and multiple testing, where significance levels may be data-dependent. Existing asymptotic e-values, however, suffer from the ``missing factor,'' a scaling inefficiency resulting in overly conservative inference. Drawing on the framework of near-optimal concentration inequalities developed by Bentkus in the 2000s, we introduce Bentkus-type asymptotic e-values and prove that they successfully eliminate the missing factor. We also demonstrate both theoretically and empirically that Bentkus-type e-values consistently deliver sharper inference than existing alternatives, leading to tighter post-hoc confidence intervals and higher rejection rates in multiple testing procedures.

math.ST↗

HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers

Humanoid loco-manipulation benefits from one controller to coordinate locomotion, arm movements, and fall recovery. These behaviors are difficult to learn together from scratch because they require different skills and objectives. We present HANDOFF, a training architecture that distills motion-tracking, locomotion, and recovery teachers into one 29-DoF policy, using context-conditioned KL losses over separate action slices and a soft mixture of experts. Commanded velocity continuously blends body supervision, while a binary recovery flag assigns the full action to the recovery teacher. The resulting controller takes a compact 10-D task-space command rather than a dense kinematic stream and does not switch policies at runtime. On the Unitree G1, HANDOFF matches state-of-the-art velocity tracking and obtains the largest robust manipulation workspace in a matched adapted-interface benchmark. The same controller executes natural-language-driven, multi-stage loco-manipulation tasks in simulation and hardware experiments without further data collection or fine-tuning.

cs.RO↗

Hierarchical Forecast Reconciliation for Urban Rail Transit Demand Prediction under Operational Disruptions

Accurate and coherent passenger demand forecasting is essential for Urban Rail Transit (URT) operations. Passenger demand is hierarchical: origin--destination (OD) flows aggregate to station-level inflows and outflows through conservation constraints. However, independently generated station- and OD-level forecasts may violate these constraints and limit information sharing across levels. This paper develops a hierarchical forecast reconciliation framework for joint station- and OD-level demand prediction. A neural Fully Connected Reconciler (FCR) maps incoherent base forecasts to coherent predictions with exact structural consistency by construction. We benchmark FCR against classical and machine-learning reconciliation methods using one year of Rejsekort smart-card and Banedanmark operational data from a 12-station Copenhagen S-train subnetwork, covering one-step, multi-step, and disruption forecasting. We also compare against a matched-parameter multi-task station--OD baseline with a soft consistency penalty. Reconciliation improves aggregate OD accuracy while enforcing exact coherence. Under standard conditions, FCR is competitive with statistical methods, while an oracle experiment using observed station demand reduces OD MSE by about 34\%. Under train delays and cancellations, reconciliation continues to improve OD accuracy, with FCR achieving larger reductions than MinT-Sample in disruption-affected subsets. Greater cross-level disagreement is associated with larger reconciliation gains, highlighting the value of hierarchical reconciliation under both regular and disrupted operations.

cs.LG↗

Projector Quantum Variational Ansatz

Quantum computing offers several algorithms for computing the ground state of a problem Hamiltonian. Algorithms in the Fault-Tolerant Quantum Computing (FTQC) regime, such as Quantum Phase Estimation (QPE) and Quantum Signal Processing (QSP), are the most desirable because they offer optimal asymptotic scaling. In the Noisy Intermediate Scale Quantum (NISQ) regime, the most realistic approaches involve Variational Quantum Eigensolver (VQE) algorithms and their variants. A central difference between the two regimes is that FTQC algorithms do not prepare the ground state via direct state transition. Instead, they build a project that identifies the ground state using ancillary qubits that flag the target subspace. The desired state is then recovered via amplitude amplification or post-selection. In this work, we propose a VQE ansatz whose structure is more similar to that of an FTQC algorithm. We introduce the Projector Variational Ansatz (PVA), which incorporates the projector structure into the adaptive variational setting. Depending on its parametrization, this ansatz can be equivalent to either an Adaptive Derivative-Assembled Pseudo-Trotter (ADAPT)-VQE when all ancilla mixing angles vanish, or to an Intermediate Scale Quantum (ISQ)-QSP quantum circuit structure. We benchmarked PVA against an ADAPT-VQE baseline using both qubit-excitation and fermionic-operator pools. We propose an implementation in which an L-layer PVA circuit is designed using 3L + 1 parameters and requires only one ancilla and two additional CNOT gates per Pauli generator used in ADAPT. Our experimental results show that this first proposal for PVA converges using fewer CNOT gates than the standard ADAPT-VQE on the tested system.

quant-ph↗

EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control

Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specified trajectories, and humanoid vision-language-action systems provide semantic interfaces, but neither offers a scalable and interactive prior for broad full-body behavior. We introduce EgoPriMo (Egocentric Motion Prior for Humanoid Robots), a unified framework that learns such priors from egocentric human demonstrations. Given egocentric observations and a text prompt, EgoPriMo reconstructs, generates, and forecasts SMPL-based full-body motion. Language is used as a high-level control signal rather than a complete motion specification. At the core of EgoPriMo is a Triple-stream DiT that jointly models body dynamics, egocentric visual context, and text; task-conditioning masks route different tasks and missing-modality data through the same checkpoint. Experiments on Nymeria and EgoExo4D show that one checkpoint improves egocentric motion generation over UniEgoMotion while supporting reconstruction and forecasting; the generated SMPL motions can also be executed by a Unitree humanoid controller. These results indicate a practical path from scalable egocentric observations to generalizable and interactive humanoid motion priors.

cs.RO↗