Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models

We present CombEval, a dynamic benchmark for evaluating combinatorial counting in large language models. CombEval represents each problem as a typed Cofola specification over entities, combinatorial objects, object dependencies, and constraints, enabling controlled generation of natural-language counting problems with exact solver-verified answers. Unlike static collections, CombEval supports systematic variation of object type, entity scale, constraint count, and reasoning depth. We evaluate 11 LLMs under direct and code-augmented settings and find that models remain brittle on ordered objects, indistinguishable elements, relatively positional constraints, and nested object dependencies. Error analysis further identifies failures in constraint interpretation and counting principles. CombEval provides a diagnostic testbed for studying when and why LLMs fail at combinatorial reasoning. The code and generated benchmark suites are publicly available at https://github.com/YuxuZhou-CN/combination-problem-generation.

cs.AI↗

Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory

Vision-Language-Action (VLA) models are increasingly deployed across diverse tasks, yet they remain black boxes whose physical interactions can cause irreversible harm, making generalizable and interpretable failure detection essential. We observe that successful and failed rollouts carry systematically different information-theoretic signatures. Building on this, we formalize VLA control as a closed-loop information pipeline and derive the Triple Information-theoretic (Tri-Info) signals that capture whether actions remain diverse, temporally consistent, and coupled to state transitions. Across six VLA models and three benchmark environments, Tri-Info matches the strongest baselines in-domain. Moreover, Tri-Info transfers across architectures, environments, and the sim-to-real gap without retraining with labeled data, reaching 70\% accuracy on real-world tasks. This establishes Tri-Info as a simple yet powerful method that not only detects failures with strong cross-domain generalization, but also delivers interpretable diagnostics of the underlying failure modes.

cs.RO↗

OctoNest: Adaptive Cross-Device Execution through Stateful Control

Computer use agents are expanding from single-device operation toward cross-device systems that coordinate tasks across heterogeneous environments. Execution conditions are often only partially known at planning time and revealed through interaction. Failures may require intra-device modality switching or inter-device reassignment; failing to distinguish these cases can lead to repeated failures or premature termination. However, existing systems primarily scale up single-device agents without sufficiently distinguishing device-level and modality-specific execution conditions. We propose OctoNest, which coordinates stateful cross-device orchestration and iterative device-local modality control. Device Agents refine subtasks and select modalities, while an Orchestrator uses execution feedback to revise plans and device assignments. We also introduce CAPEBench, comprising 158 instances from 23 cross-device seed tasks with controlled perturbations. OctoNest leads all three quality metrics, improving Perfect Pass over the strongest baseline by 18.35 percentage points and reducing token cost per perfect pass by 39.8\%. Further analyses support the complementary roles of local refinement and global revision and demonstrate CAPEBench's ability to distinguish control limitations under changing execution conditions.

cs.CL↗

Gas-induced gravitational-wave dephasing and accretion periodicities of live post-Newtonian massive black hole binaries: warm disk

We perform 3D hydrodynamical simulations of an equal-mass quasi-circular live $10^6~{\rm M}_\odot$ massive black hole binary (MBHB) embedded in a prograde, locally isothermal circumbinary disk (CBD) with $0.1$ aspect ratio. The binary evolves under the effect of gaseous torques and $2.5$ post-Newtonian dynamics. This approach allows us to track the influence of the CBD on a gravitational-wave (GW) driven MBHB inspiral all the way down to merger from $53$ Schwarzschild radii ($r_s$) over $1.5$ years. Comparing the GW inspiral rate with the viscous inflow, we find the binary to decouple from the CBD just $\approx2$ days before merger. We measure gas torques (gravitational and accretion) with and without concurrent GW emission, finding that their sums agree to within $\lesssim5\%$ down to $25~r_s$. We then measure a gas-induced phase-shift in the GW signal of $\approx9.0\times10^{-4}$ rad that accumulates over a year until $40~r_s$ and saturates afterward, which should be LISA-detectable at redshift $z{\lesssim0.2}$. We further characterize the mass accretion rate modulations. We recover the periodicity associated with the ``lump" at the inner cavity edge. The periodicities at the binary orbital period and at half of it are shifted to lower frequencies. We measure the former, due to relativistic apsidal precession, in an inspiraling binary for the first time, and newly identify the latter as due to the precession of the eccentric cavity. Our results have implications for multi-messenger astronomy, since observation of accretion rate modulation by LSST/Roman surveys and phase-shift by LISA will provide crucial information on the complex environment surrounding MBHBs.

astro-ph.GA↗

Wave-Energy Partition Governs Weak Collisional Damping in Cold Plasmas

Weak dissipation can control wave propagation, mode competition, and instability thresholds in plasmas, yet the physical origin of large branch-to-branch differences in collisional damping is often obscured by dielectric-tensor calculations. We show that weak collisional damping in cold plasmas is governed by wave-energy partition. In the one-rate cold-plasma model, the damping rate of a collisionless eigenmode is exactly the collision frequency multiplied by the fraction of the total wave energy stored in plasma motion. This result recasts the standard perturbative damping formula into a compact and physically transparent law, immediately explaining why field-dominated branches such as whistlers can be much less damped than the collision frequency, whereas quasi-electrostatic modes can exhibit damping of comparable magnitude. Analytic examples for Langmuir, transverse electromagnetic, whistler, and extraordinary waves show that the energy-partition form classifies weak collisional damping across distinct branches and provides a simple diagnostic for mode competition in multibranch plasma-wave systems.

physics.plasm-ph↗

Probing Light Dark Fermions in $B \to D^{(*)}\ell X_{\rm inv}$ via Rate Distributions

Experimental analyses of the semileptonic decays $ B \to D^{(*)} \ell \barν$ typically rely on the assumption that the missing energy originates from a massless neutrino, as predicted by the Standard Model. However, this assumption may not hold in scenarios where the invisible final-state particle is instead massive, such as a sterile neutrino or a dark-sector fermion. In this work, we explore how the presence of a massive dark sector fermion modifies the kinematic and angular distributions of these decays. Our analysis is carried out within the framework of a general weak effective theory, and we also discuss effective and simplified models in which these interactions may arise. In addition, we study the implications of these effects for the extraction of the CKM matrix element $ |V_{cb}|$. Overall, our results show that relaxing the standard assumption of a massless neutrino can lead to observable effects and provide a framework for systematically investigating their impact on semileptonic $B$- decay distributions.

hep-ph↗

FlowDPG: Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation

Real-world reinforcement learning for robotic manipulation remains challenging, and this difficulty is amplified for flow matching policies: policy gradients must be backpropagated through time (BPTT) along the multi-step ODE that maps noise to actions, which is computationally prohibitive and numerically fragile. We propose FlowDPG, a DDPG-style method for flow matching policies that bypasses BPTT entirely. FlowDPG evaluates the critic gradient at a one-step estimate of the clean action and distills the resulting correction into the velocity field at training time, leaving multi-step inference unchanged. Intuitively, it combines two complementary vectors: the demonstration-driven velocity that keeps the action feasible, and the critic-driven correction that steers it toward higher value. Our contributions are threefold: (1) a BPTT-free distillation framework for stable DDPG-style improvement of flow matching policies, (2) a formal connection between the FlowDPG update and the vanilla deterministic policy gradient via three explicit approximations, and (3) real-world validation on two long-horizon tasks on different robots. FlowDPG reaches 92% end-to-end success on dual-arm AirPods assembly with Franka arms and an 86% rubric score on scrambled-egg cooking with YAM arms, substantially outperforming recent RL methods. Videos and more results: https://flowdpg.github.io.

cs.RO↗

Holographic correlation functions of fermions in anisotropic plasma

By using the gauge-gravity duality, we study the holographic fermionic correlation functions in strongly coupled anisotropic plasmas. Starting from the isotropic AdS-Schwarzschild black brane background, we revisit the prescription for computing the retarded Green's function of a probe Dirac fermion and then generalize the formulas with respect to the anisotropic geometries. The method is applied to three distinct holographic models that capture different physical origins of anisotropy: axion-induced, magnetic-field-induced and unquenched-flavor-induced. Numerical results for the holographic correlation functions reveal direction-dependent corrections, negative dips in the imaginary part, Landau levels in the fermionic dispersion (magnetic field), and a momentum-independent pseudogap indicating an incoherent metallic phase (flavors). Our results complement and go beyond the hard thermal loop approximation, providing non-perturbative insights into fermionic excitations in strongly coupled anisotropic plasmas relevant for heavy-ion collisions and certain condensed matter systems.

hep-th↗

Slice Monte Carlo Integration

Numerical integration involving expensive target functions is a common bottleneck in Bayesian inference and simulation. When a cheap surrogate is available, standard approaches such as reweighting or importance sampling often suffer from high variance and inefficient use of function evaluations. We introduce Slice Monte Carlo integration (S$\ell$MC), a method that leverages a Nested Sampling-like procedure on the surrogate to partition the space into informative strata, or slices, while generating samples in the parameter space drawn from the prior within each slice. This enables stratified Monte Carlo integration of the expensive target function over the surrogate-induced partition, yielding an efficient estimate of the target integral. The surrogate level sets therefore make the induced ordering of the parameter space the relevant information for S$\ell$MC, rather than the pointwise target-to-surrogate ratios governing importance-sampling efficiency. Another key advantage of S$\ell$MC is the decoupling of slice volume estimation from the evaluation of the target, allowing the refinement of the slice-volume estimates without additional target evaluations. We investigate the properties of S$\ell$MC, demonstrate how to efficiently generate posterior samples, and assess its robustness in controlled Gaussian benchmark families with continuously tunable surrogate-target mismatch.

stat.ME↗

Fast mixing of all-to-all quantum systems at high temperatures

It is shown that arbitrary quantum $k$-local Hamiltonians with bounded strength interactions admit a quantum Gibbs sampler [CKG23] with a system-size independent spectral gap, at sufficiently high temperatures. As a consequence, such systems admit fully-polynomial time quantum approximation algorithms for partition functions and global expectation values.

quant-ph↗

Analogues of Grün's lemma and Baer's theorem for skew left braces

We prove in this paper some analogues of the well-known group-theoretical Grün's lemma, stating that in a perfect group the first and the second centre coincide, and Baer's theorem, stating that if the quotient by the $n$th centre of a group is finite, then so is the $(n+1)$th term of the lower central series, in the scope of infinite skew left braces. These results represent significant improvements over previous work. The trifactorised group associated with a skew left brace will be crucial for our proofs.

math.GR↗

What Survives When You Compress a Recursive Reasoner for the Edge?

Recursive reasoning models solve hard structured tasks with a few million parameters by iterating a latent state, but deploying them on edge hardware means compressing them -- and quantization noise compounds across recursive cycles rather than accumulating over output tokens, so single-pass intuitions fail. Here, we ask what survives. Across a full precision sweep and three tasks, aggressive compression preserves local prediction but destroys global reasoning: pruning, distillation, and linear attention drive puzzle-exact accuracy to zero while cell accuracy holds. Naive INT4 is different: its damage is architectural, striking MLP-mixing recursion but not attention on the same task, and per-channel calibration reverses it without retraining, where quantization-aware training does not. We also introduce carry-trajectory fidelity, the cosine similarity to the full-precision reasoning path, as a label-free signal that predicts both the damage and the recovery before any task evaluation, with a threshold that transfers across tasks. The result is a deployment recipe: flash-streamed embeddings remove a 99.4 MB bottleneck, one recursive cycle matches full-depth accuracy at 6x fewer FLOPs, and calibrated INT4 brings the backbone to 3.3 MB, inside a 4 MB budget.

cs.LG↗

Light selects non-equilibrium states of phase-separated droplets

Driving phase-separated droplets out of equilibrium can oppose coarsening and trigger division. Selecting and sustaining these distinct states remains an important experimental challenge. Here, we use reversible photoswitching to tune molecular driving in DNA-azobenzene coacervates confined in microfluidic droplets without chemical fuel depletion or waste accumulation. Single-wavelength illumination sustaining bidirectional molecular switching arrests droplet coarsening at a finite size. Dual-wavelength illumination with wavelength-dependent penetration depths generates spatially asymmetric molecular driving within uniformly illuminated droplets, producing persistent interfacial instabilities. Under low-salt conditions, droplets divide and undergo recurrent growth-division cycles. Kinetic and thermodynamic measurements reveal how flux balance, phase stability and droplet size govern these distinct non-equilibrium states. These findings establish optically controlled molecular driving as a route to selecting sustained non-equilibrium droplet dynamics.

cond-mat.soft↗

From Dataset Spectral Geometry to Network Weights: A Geometry-Aware Initialization for Sigmoidal MLPs in Image Classification

Classical universal approximation theorems (UAT) establish the expressive power of sigmoidal multilayer perceptrons, but they do not specify how the weights should be initialized. We study a supervised, data-dependent, geometry-aware initialization for one-hidden-layer sigmoidal MLPs that compiles labeled class geometry into network weights. The construction starts from the idea that sigmoid units can act as smooth half-space gates. For each class, we center the training samples at their mean, apply SVD to estimate principal directions and spectral scales, select retained directions by an energy threshold, and represent each retained direction by a pair of sigmoid slab gates. These class-specific gates are then concatenated into a shared hidden layer initialized directly from the training set. We also formulate a SVD-Mahalanobis subspace classifier as a non-neural geometric reference, which tests whether the estimated spectral class geometry is already discriminative before being embedded into the MLP. Experiments on MNIST, Fashion-MNIST, and CIFAR-10 show that the proposed initializer produces a substantially more informative zero-epoch state than a matched task-agnostic Xavier reference, while full training reaches comparable final accuracy. Frozen-hidden experiments and neutral-head ablations further show that the class-wise SVD gates remain useful fixed features, even after the initially aligned output layer is replaced by a Xavier-initialized one.

cs.LG↗

Characterization of the RF Board for microwave SQUID multiplexing readout electronics

Microwave SQUID multiplexing ($μ$MUX) is a widely used readout technique for large-scale transition-edge sensor (TES) arrays. It uses radio-frequency (RF) probe tones to interrogate cryogenic resonators, requiring frequency conversion between the baseband electronics and the cryogenic RF signal chain. This work describes the RF Board, a room-temperature frequency-conversion board deployed in the AliCPT $μ$MUX readout system. The board up-converts 0-4 GHz baseband I/Q signals to the 4-8 GHz RF band for injection into the cryogenic chain and down-converts returned RF signals to baseband I/Q for ADC digitization. For the current 1000-tone operation, with a DAC output tone power of -30 dBm/tone, the required power windows are -35 to -25 dBm/tone for the RF tones transmitted to the cryostat and -45 to -35 dBm/tone at the ADC input for the returned tones. The RF Board is characterized using swept single-tone measurements covering up-conversion, down-conversion, and RF loopback. Based on these measurements, the RF output power is calculated to be -31.06 to -25.53 dBm, satisfying the RF output window. Assuming a representative cryogenic-chain transmission of -40 dB, the loopback result gives an estimated returned power of -45 to -35 dBm, within the target range. These results show that the RF Board meets the wideband frequency-conversion and tone-power requirements for the $μ$MUX readout system.

astro-ph.IM↗

AI Training Manager: Bounded Closed-Loop Control of Adaptive Training Recipes

We present the AI Training Manager, a bounded LLM-based metacognitive monitoring-and-control layer for machine-learning training. The Manager asynchronously observes structured telemetry from an active training run, assesses the current training regime, and selects adaptive interventions through a constrained, deterministically verified action interface. We evaluate the feasibility of this architecture using deliberately induced, interpretable failure modes in supervised and reinforcement learning tasks. For supervised learning, we induce a multi-objective overfitting failure in a GPT-2-style model trained on TinyStories. The Manager prevents the resulting late-run validation collapse, reducing final validation loss by 54.3% relative to the same stressed recipe. For reinforcement learning, we study robotic reaching under opposing multifactor stress regimes. In a conservative regime, the Manager raises final deterministic safe success from 0.413 to 0.705. In an aggressive regime, it raises safe success from 0.121 to 0.764. The same Manager instructions and action surface are used in both regimes. These results suggest that bounded LLM reasoning can provide a practical metacognitive control layer over ongoing learning processes.

cs.AI↗

Quantum Lazy Sampling and Path Recording for Any Group

A central challenge in quantum algorithms and cryptography is reasoning about algorithms with oracle access to a random group element (e.g. a random function, permutation, or unitary). Can we efficiently simulate such algorithms? Can we determine what they know after t queries? A classical tool for this is lazy sampling: the oracle does not commit to the full group element upfront, but rather samples partial information about it on the fly. We study a quantum analog of lazy sampling: compressed oracles (or recording oracles). These are quantum data structures that allow on-the-fly simulation for quantum queries, originally introduced by Zhandry (CRYPTO '19) for random functions, and generalized to unitaries by Ma-Huang (STOC '25) and permutations by Carolan (STOC '26), and used to great effect in security proofs and lower bounds due to their interpretability. We define and analyze a general-purpose and interpretable path-recording oracle, derived from first principles, that perfectly simulates random elements of any closed subgroup of $U(N)$. Our oracle stores, in superposition, t input-output pairs, with updates described in terms of the commutant of the group's tensor power representation. This transparently records the information the algorithm has learned. Our oracle builds on recent work of Grinko-Yoshida (QIP '26), who gave a different general-purpose compressed oracle without clear interpretability. One interesting application of our path-recording is allowing direct comparisons between compressed oracles of different groups, giving a new technique for proving pseudorandomness results. For example, comparing $S_N$ and $U(N)$ yields what is arguably the simplest construction to date of pseudorandom unitaries: the product PC of a pseudorandom permutation and a random Clifford, improving on the prior PFC construction (Metger-Poremba-Sinha-Yuen, FOCS '24; Ma-Huang, STOC '25).

quant-ph↗

Provably Efficient Learning of Fermionic Correlations under Particle-Number Symmetry

Predicting local fermionic correlations is a central task in quantum many-body physics, as these correlations encode many physically relevant local observables. The ubiquitous particle-number symmetry imposes strong structural constraints on quantum states, suggesting that local correlations should be learned with fewer samples than by symmetry-agnostic approaches. However, it has remained unclear whether such a provable advantage exists in collective learning of local correlations. Here, we develop a framework of number-conserving fermionic-shadow tomography based on random orbital rotations. We prove that, for every given order $k$, we can simultaneously estimate all $k$-body fermionic correlations of an $N$-mode $η$-particle state with a given variance $\varepsilon^2$ using only $O_k(η^k/\varepsilon^2)$ samples, which are independent of the system size $N$. We further establish a matching information-theoretic lower bound $Ω_k(η^k/\varepsilon^2)$ for any adaptive protocol based on single-copy measurements, showing that the $(η^k,\varepsilon)$-dependence is optimal up to constants depending only on $k$. Furthermore, numerical studies show a 20-fold query reduction for one-body correlation estimation at $N=200$, $η=20$, and $\varepsilon=10^{-2}$, compared with the best alternative including Heisenberg-limited estimation methods. Relative to fermionic Gaussian-unitary shadows, sample costs for the square-lattice Fermi-Hubbard model at $N=288$ are reduced by factors of $148$ for energy density and $161$ for spin structure factor estimation. For molecular energy derivatives with respect to atomic coordinates, our method reduces sample requirements by factors of $53$ for $\mathrm{Cr}_2$ and $18.5$ for $[\mathrm{Fe}_2\mathrm{S}_2]^{2-}$ at $N=128$. These results demonstrate the practical advantage of particle-number symmetry for fermionic observable estimation.

quant-ph↗