Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 685 records · Page 38Linked to original sources

Population Scaling or Data Dilution? Dynamics of Local Topology Evolution in Decentralized Learning

Scaling decentralized learning changes not only the number of clients $N$, but also the dynamics of information propagation and consensus. We argue that the effect of increasing $N$ cannot be understood in isolation, because data allocation, topology-dependent mixing, and communication capacity may change simultaneously. We study these coupled effects on CIFAR-10 with $N\in\{10,50,100,200\}$, comparing a degree-two Ring, a Static Random graph, and Local-First Heuristic Evolution (LFHE), a locally adaptive topology process based on friend-of-friend discovery. The Ring provides an analytically transparent failure mode: its Metropolis spectral gap decays as $Θ(N^{-2})$, implying progressively slower contraction of model disagreement as the population grows. Experiments show that holding the nominal local dataset size fixed substantially reduces the apparent population penalty observed when a fixed total dataset is divided among more clients. The remaining degradation depends strongly on communication structure: Ring enters a high-disagreement regime, whereas Static Random and LFHE remain close to consensus. Increasing LFHE's degree threshold further improves accuracy and consensus, but at a substantially higher model-transmission cost. These results show that decentralized scaling is governed by coupled learning and communication dynamics, rather than by the number of clients alone.

cs.LG↗

Universality and Convergence of Generative Flows

Generative flows sample from an unnormalized target by training a flow to be balanced, and the training loss is the signal a practitioner watches. We ask what that signal is worth: whether a small loss certifies an accurate sampler, whether the loss can be driven to zero, and how fast gradient descent does so. The loss decides the first. Losses that compare the two sides of the balance by their difference bound, in total variation, the error of the sampler the flow implies, with explicit constants that do not involve the policy; flow-matching losses that compare them through a ratio admit no such bound, already on a single cycle, whenever their generator is continuous at balance. On graphs, the backward policy decides the other two. Once it is frozen, balance becomes invariance under the backward chain, so that existence is free on finite graphs, and one constant --- the norm of that chain's Green operator, which plays the role of an inverse spectral gap --- fixes the order of the curvature of the loss around the balanced flow, from above and below, and sets a floor under the rate at which training converges near it. The mechanism is that gradient descent diffuses the flow along the backward policy. For the squared-logarithm generator of detailed and trajectory balance, training the balance loss on states converges globally on every finite path-connected graph, from every positive initialization. The constant can be infinite while backward trajectories are short on average, and exact flow matching can then fail. The bounds and rates are tested by exact computation on enumerable state spaces, and every theorem carries a certification status computed from a Lean~4 development.

cs.LG↗

How Should a Prompt Optimizer Spend a Tight Budget? BudgetAPO with Noise-Adaptive Evaluation

Automatic prompt optimization (APO) has been widely employed to adapt large language models without updating their weights, yielding promising results. However, existing methods such as GEPA and OPRO assume hundreds to thousands of subject-model calls, far more than is practical behind paid, rate-limited APIs. Under tight budgets they fail in two ways: multi-stage pipelines can exhaust the budget and return the seed prompt unchanged, while single-stage methods compare candidates on fixed-size minibatches, regardless of each task's noise. As a remedy, we introduce BudgetAPO, a single-stage optimizer for the tight-budget regime. BudgetAPO incorporates (1) a noise-adaptive rule that sizes the evaluation slice to each task's noise, measured by a short probe; (2) a fixed slice that turns every accept/reject decision into a paired comparison; and (3) a reflective operator that rewrites reasoning strategy and output format jointly. Extensive results across seven benchmarks and five subject models demonstrate that BudgetAPO ranks first on every subject and beats every baseline under Holm-corrected paired tests, while returning the seed in 13% of runs at 250 calls against 86% for GEPA. On GPT-OSS-20B, GEPA needs 5 times as many calls to match \method's 100-call score.

cs.AI↗

Spatial delocalization of cold oscillators in hollow-core fibers

Trapped particles in hollow-core fibers enable long-range sensing, though advanced control of their in-fiber motion, such as multimodal cooling and squeezing, remains challenging. Here, we present optical interference-based adaptive techniques for controlling the motion of a fringe-trapped silica nanoparticle inside a fiber. After feedback-cooling its axial and radial motion, we induce axial delocalization (position anti-squeezing) via two complementary approaches. First, non-adiabatic fringe suppression expands the position variance by 11.83 ($\pm$0.7) dB to that of the initial cold-state while retaining Gaussian statistics. Second, multi-pass particle positioning at dark fringes increases the delocalization to 13.23 ($\pm$0.5) dB relative to the cold-state's variance via dark inverted optical potentials, in agreement with our Wiener stochastic model. Stronger delocalization produces non-Gaussian states of motion. Unlike prior inverted-trap implementations, our method requires neither auxiliary optical traps nor charged particles in Paul traps. Our results demonstrate fringe-trapped particles in hollow-core fibers as a versatile platform for long-range sensing and macroscopic quantum physics.

physics.optics↗

Computing Stable Matchings under Complementarities and Preference Misalignment

We study many-to-one, two-sided stable matching problems in which preferences are complementary and firms and workers may rank the same allocation differently. Here a coalition is a group of workers that can be jointly matched with a firm. With complementary preferences, a stable matching need not exist. We show that a stable matching exists for every instance when two conditions hold. The first requires each firm to have an anchor worker who is included in every feasible coalition that can be matched with that firm. The second is Transitive Alignment, which means there must be an overall ranking across different coalitions that is consistent with the workers' preferences. We propose CDAR (Combinatorial Deferred Acceptance with Reproposals), a Deferred Acceptance-style algorithm that permits reproposals, and prove its finite-time convergence and its output of a stable matching by constructing an $N$-digit potential function. Moreover, we show that, under strict preferences, the output of CDAR is Pareto optimal among stable matchings. The proposed framework naturally captures applications such as vehicle-route assignment for traffic safety and the misaligned interests of labor and management in wage negotiation.

cs.GT↗

BabelFake: A Multilingual Audio-Visual DeepFake Benchmark

Reliable and practical audio-visual DeepFake detection requires benchmarks that reflect diverse linguistic contexts and modern data synthesis pipelines for visual as well as audio manipulations. However, existing datasets predominantly contain footage of English-speakers, often include outdated manipulation types, or overlook the audio modality. Further, many datasets feature individuals who did not consent to be used in DeepFake creation. We introduce BabelFake, a multilingual audio-visual DeepFake benchmark recorded with consenting participants. BabelFake contains 399k clips (1,323 hours) from 496 individuals spanning five languages (English, German, Italian, French, Spanish). Our modular data generation pipeline pairs 11 modern video manipulation methods with 4 voice cloning engines, distinguishing visual-only (face swapping) and joint audio-visual manipulations (lip synchronization and portrait animation). By benchmarking state-of-the-art detectors, we show that detection difficulty depends on the audio-visual generation pairing, with substantial performance degradation when authentic audio is preserved. Cross-language/demographic evaluation reveals sensitivity varying across detector architectures and training data, while human evaluation reveals that perceived realism and machine-detection difficulty do not necessarily align.

cs.CV↗

DIALER: A Case for Improving Rare-Class Accuracy in Retraining-Free Edge Video Analytics

Edge video analytics with lightweight models is prone to accuracy degradation due to persistent distributional shifts in live video streams. While continuous learning (CL) addresses such data drift, it heavily strains the limited compute resources of edge servers originally provisioned for inference. Our empirical study reveals that emerging vision foundation models (VFMs) offer a practical, retraining-free alternative that delivers high average accuracy with remarkable compute savings. However, VFMs frequently misclassify specific rare classes, which often represent critical objects, as visually similar common classes. We design DIALER, a system that exploits the spare compute cycles freed by retraining-free VFM inference to mitigate rare-class misclassifications. Specifically, DIALER pre-builds multi-stage correction pipelines for dominant rare-to-common confusion pairs offline. At runtime, it routes correction candidates to the corresponding pipelines and executes as many stages as idle GPU headroom permits. Evaluation on four real-world driving datasets shows that DIALER improves rare-class accuracy by up to 14.0% without interfering with real-time VFM inference for multi-stream analytics.

cs.DC↗

SOL: Measuring Gaps between Text Distributions by Double Sliced Wasserstein Metrics

Evaluating text generation requires measuring how well the generated distribution matches the data distribution. For autoregressive models, this is done by the perplexity. Diffusion and flow-based language models can only provide a likelihood bound, whose tightness differs between model families. Sample-based substitutes such as generative perplexity with entropy do not consider the distribution fit. We propose SOL, a distance between text distributions. Each sequence is represented by the empirical measure of its hidden states under a fixed transformer and the distributions of these measures are compared by the double sliced Wasserstein distance. We prove that SOL is a metric if the transformer is injective. Experiments show that SOL detects distributional failures, recovers expected model trends, and provides stable sample-based estimates. We put forward SOL to fill the gap in the current evaluation protocol used for non auto-regressive models. As a first step we use SOL to re-evaluate a variety of models trained on OpenWebText.

cs.CL↗

The Review Lottery: Benchmarking an Observational Estimator of Peer-Review Noise (ICLR 2017-2025)

How much of a conference accept/reject decision would change if the same paper were reviewed by a different set of reviewers? Running a second independent program committee is the gold standard for answering this, but it is prohibitively expensive: done only twice (NeurIPS 2014 and 2021). We build an observational estimator of this quantity from public review data alone, calibrate it twice, and apply it to nine years of ICLR (2017-2025; 36,113 papers, 134,912 reviews). The estimator decomposes scores with a Bayesian ordered-probit model into paper quality and reviewer noise, maps scores to decisions with a logistic model, and simulates two independent committees (posterior draws B=1,000; committee sizes k=2,3,4). Estimated disagreement rates are 23-30% at k=2 and 18-24% at k=4; 30-50% of accepted papers would be rejected. External calibration: at the NeurIPS 2021 reviewer-count caliber (k=3), the simulated 2021 disagreement rate is 23.3% [21.7%, 25.0%] vs. reported 23.0% (bias +0.3pp); accept precision and committee correlation agree within 5pp and 0.04. Internal calibration: on 18,740 papers with 4+ reviews, random model-free 2+2 reviewer splits agree with the k=2 simulation within 1pp in 2018 and 2021-2025. Longitudinally, we find no robust time trend in reviewer noise over 2017-2025. The high accepted-paper flip rates of 2020 and 2021 have distinct mechanisms: the 2020 four-point scale compressed scores (23.7% of papers had zero within-paper variance), and a counterfactual shows coarsening the scale raises disagreement by about 7pp; 2021 instead combined the lowest signal-to-noise ratio in the sample with the most threshold-crowded acceptances. For the LLM era, a 2023 breakpoint test on within-paper score variance finds no break, but the design has almost no power, and no post-2022 review text or confidence data exist, so no LLM attribution is attempted.

cs.AI↗

Relaxation of a Vlasov gas to an inhomogeneous state due to phase space mixing in an axisymmetric potential: A Newtonian analogy of the Kerr orbital motion

We explore the dynamics of a Vlasov gas propagating in an external axisymmetric potential consisting of a central potential with an additional quadrupolar component which gives rise to two potential wells along the symmetry axis. Employing independently $N$-particle simulations and statistical methods, we show that an initially homogeneous and isotropic configuration evolves to an inhomogeneous final state with overdensities located at the wells. The quadrupolar component is chosen such that the equations of motion form an integrable Hamiltonian system, which allows one to compute the final state of the gas analytically. On the one hand, this allows one to compare the late-time behaviour of the simulations with an analytic prediction and on the other hand to understand the relaxation process through phase space mixing. Our model constitutes a Newtonian analogue of the particle motion in a Kerr spacetime, and we discuss possible applications to recently observed astrophysical phenomena.

astro-ph.HE↗

Learning to Decide, Not to Reason: Parameter-Efficient Decision Operators via Low-Rank Activation Steering

Injecting skills into a frozen language model currently costs a million parameters and a reinforcement-learning pipeline. We introduce DecSteer, a System-1 decision operator trained by behavior cloning that lowers this cost by roughly two orders of magnitude. The default operator uses 330K parameters to match a 1.33M-parameter operator trained with reinforcement learning, exceeds or achieve comparable performance, while collapsing 3,685-token deliberation into a 6-token decision with no loss in accuracy. A rank-4 variant with 23K parameters, 1/58 of the strongest published skill operator, suffices for SearchQA and near-suffices for LiveMath, where higher rank still helps; the same recipe transfers across five tasks and three backbones, with out-of-distribution gains persisting on LiveMath problems released months after training. The gap to prior work is trainability, and it is set jointly by initialization and architecture. The initialization of prior operators zeroes the gradient of both large factor matrices at the first optimization step, whereas our zero-initialized output projection inside a shared low-rank backbone receives a gradient immediately, which a gradient-flow probe confirms directly. The gain isn't chain-of-thought compression. 23 of 57 LiveMath points beat the base model's best-of-8 sampling, and a logit-lens probe shows the operator amplifies the answer along the model's existing late-layer pathway, not writing it earlier. Gains track the base model's headroom across 13 base-task pairs, and skills compose as approximately linear operators that can be added, interpolated, and hot-swapped at inference time.

cs.LG↗

ROMA: LLM System for Real-World Object-Centric Multi-Sensory Active Perception

Humans inherently understand the physical world through an active process. When sensory evidence is insufficient to infer physical properties, we naturally interact with the environment by deciding what information is missing, how to acquire it, and when sufficient evidence has been obtained. In stark contrast, existing multi-sensory robot systems mainly integrate sensory inputs rather than actively acquiring missing evidence through interactions. In this work, we introduce ROMA, an LLM-based system for Real-World Object-Centric Multi-Sensory Active Perception. ROMA integrates vision, audio, tactile, and force sensing into a reasoning-interaction-feedback loop. The model identifies missing evidence and determines the target objects, interactions, and modalities, while a physical interface executes the selected interactions and collects the multi-sensory feedback. To support this capability, we construct ROMI-2K, a large-scale real-world multi-sensory object interaction dataset covering nearly 2,000 objects and 6 atomic interactions with synchronized sensory feedback. Building on these data, we develop a two-stage training framework that aligns sensory modalities and equips the LLM to assess evidence sufficiency, select informative interactions, and reason over the multi-sensory feedback. We further characterize active perception as perception chains, where acquired evidence guides subsequent interactions and reasoning, and establish ROMA Bench to evaluate single-attribute, long-horizon multi-attribute, and intent-driven active perception. Experiments show that ROMA can actively acquire missing evidence and solve complex, long-chain multi-sensory perception tasks that existing methods struggle to handle, laying a strong perceptual foundation for active multi-sensory embodied agents.

cs.RO↗

Regulating a Monopolist through Capacity Control

We study monopoly regulation when capacity is contractible but subsequent output is not. A firm privately observes its production cost before installing costly capacity and then choosing output, while the regulator cannot use transfers. We provide sufficient conditions under which a single capacity floor implements an optimal mechanism with full participation. The floor raises output by changing investment while leaving utilization to the firm after installation costs are sunk. Under additional conditions, the allocation has three regions: laissez-faire, full utilization of the floor, and underutilization. High-cost types knowingly install capacity they will not fully use, yet produce more than under laissez-faire. This physical waste is a necessary evil of optimal regulation that expands service. We also show that underutilization can persist in optimal regulation with capacity-contingent subsidies or a common entry fee.

econ.TH↗

High-Angular-Resolution Survey of Stars within 18 pc of the Sun: Part I

We present the rationale, design, methodology, and observational facilities of a high-angular-resolution imaging survey of stars within 18~pc of the Sun. Our comprehensive sample consists of 2047 stars compiled from Gaia eDR3 and complementary catalogs. The main goal is to observe as many stars from this sample as possible using high-angular-resolution techniques, considering their inherent limitations, to perform astrometry of known binary systems and determine detection limits for secondary components for each observed star. This approach may also allows us to identify previously unresolved stellar companions. Observations were carried out using the 2.1-m telescope at OAN-SPM in the North and the 3.58-m NTT in the South, employing high-speed cameras for speckle imaging. We obtained observations of 703 stars with V < 14 mag at OAN-SPM, and 645 stars with I < 16 mag at the NTT, of which 140 stars were observed in both hemispheres, for a total of 1208 unique stars observed. In this paper: (1) we present the construction of the catalog of 2047 stars within 18 pc; (2) using a control sample of 70 known close binaries within 18 pc with separations 1", we show that selecting targets based on Gaia quality parameters commonly used as binarity indicators--RUWE, ipd_frac_odd_win, and ipd_frac_multi_peak--can miss a significant fraction of binaries, justifying our unbiased, volume-limited survey; and (3) we present the results of pilot observations conducted with the 2.1 m telescope at OAN-SPM in July 2022. A full analysis of the complete sample will be presented in of forthcoming papers.

astro-ph.SR↗

PhysLDM: Latent Diffusion for High-Fidelity Deformable Simulation

Neural simulation of high-fidelity deformable bodies is a foundational challenge in computer graphics and physical AI. Long-horizon prediction for high-resolution 3D volumetric meshes is difficult: autoregressive methods are susceptible to error accumulation, while direct multi-frame prediction at native resolution is computationally prohibitive. This motivates a compact spatiotemporal latent representation, which is largely unexplored for mesh-based volumetric physics. Meanwhile, it remains unclear whether deterministic regression or generative diffusion is the more appropriate predictive paradigm. To address these coupled challenges, we introduce PhysLDM, a unified latent-diffusion paradigm for one-shot volumetric deformable simulation. Its core is a holistic spatiotemporal VAE that avoids the "staircase" artifacts of standard temporal compression (as in common video VAEs), achieving ~2.48 mm reconstruction precision on meter-scale scenes at up to 78x token compression. Based on this reliable latent space, we systematically compare regression and diffusion methods. Our experiments uncover a key modeling insight: complex deformable dynamics are often chaotic, and in this regime deterministic regression tends to produce non-physical averages, whereas diffusion better models their distribution. Accordingly, we employ a latent diffusion model that effectively learns from the chaotic data to generate physically plausible trajectories. Trained purely kinematically on an Objaverse-scale dataset, a single PhysLDM generalizes zero-shot to unseen OOD datasets (GSO and Toys4K). Its differentiability further enables efficient solution of inverse problems and higher-order design optimization. To our knowledge, PhysLDM is the first high-fidelity spatiotemporal autoencoder and latent-diffusion paradigm for volumetric deformable dynamics, offering a scalable and robust approach to neural simulation.

cs.CV↗

Neuromotor Hierarchy Network: Physiological Inductive Biases for Robust Generalization in sEMG Decoding

Surface electromyography (sEMG) provides a wearable, noninvasive interface to neuromuscular activity for movement decoding and human-computer interaction. Population-scale decoding remains difficult because the relationship between sEMG and neuromuscular activity varies across users and sessions, while task-relevant dynamics span channels and multiple timescales. Learning waveform-to-output mappings from task labels leaves the distinction between recording variability and coordinated motor activity implicit. We introduce the Neuromotor Hierarchy Network (NHN), which learns a compact latent neuromotor state from task supervision to represent task-relevant neuromuscular coordination. NHN constructs this latent state through a hierarchy inspired by neuromotor organization. It adapts recording statistics while preserving relative intensity. Its spatiotemporal encoder uses parameter-efficient channel interactions and modulates features with multi-timescale history. The resulting features yield candidate activations of learned motor primitives, which are temporally integrated and continuously weighted to form the state. Theoretical analysis characterizes the efficiency, temporal behavior, and optimization of NHN's core mechanisms. We evaluate the architecture for both continuous hand-pose estimation on emg2pose and touch-typing recognition on emg2qwerty. On emg2pose, NHN reduces user-averaged angular error by 0.52% to 2.84% across all three generalization splits in both Regression and Tracking relative to Hadidi et al.'s best task-specific variants, using 48.42% to 48.51% fewer parameters. On emg2qwerty, NHN reduces beam-search character error rate by 19.40% zero-shot and 30.42% after fine-tuning relative to SplashNet-Upscale, using 65.86% fewer parameters. Physiology-guided inference of a latent neuromotor state supports parameter-efficient sEMG decoding.

cs.LG↗

A Dimension-Free Bound on the Poincaré Constant of Isotropic Log-Concave Measures

The Kannan--Lovász--Simonovits (KLS) conjecture asserts that isotropic log-concave probability measures have Poincaré constants bounded by a universal constant, independently of dimension. We give a deterministic variational proof with an explicit bound on the Poincaré constant $C_P(μ)\le25$, where $μ$ is any isotropic log-concave probability measure. Starting from elliptic moment estimates and a quadratic variance inequality, we establish geometric bounds on normalized Appell coefficient norms through \emph{joint} maximization over the measure and test function. Variation of the measure gives a maximum-principle inequality, while stationarity in the test function controls the highest-order cumulant terms. Two concave barriers constructed from quadratic and cubic polynomials close the induction. A weighted Helmholtz--Hodge decomposition then controls the curl correction of weighted divergence, yielding a curvature estimate for compatible symmetric tensor fields that is uniform in rank. Quadratic duality converts this estimate into an operator comparison for centered integration, linking the coefficient bounds to control of the inverse gradient. A spectral-radius estimate and a scalar growth inequality for adjoint iterates then yield the Poincaré bound. The final estimate is independent of the auxiliary positive curvature, allowing approximation to complete the proof for general isotropic log-concave measures.

math.PR↗

Charged soliton--black hole phase transitions in three-dimensional Einstein--Gauss--Bonnet gravity

We construct a magnetic soliton by double Wick rotating the static electrically charged black hole in three-dimensional Einstein--Gauss--Bonnet gravity. We match the boundary metrics and Maxwell sources under the chosen boundary condition and compute the difference of the renormalized Euclidean actions. Because the Maxwell potential grows logarithmically at infinity, the chosen finite Maxwell boundary term changes both the fixed Maxwell source and the renormalized action. The charge contributions for both the soliton and the black hole appear with plus signs in $ΔI^{(c_{\mathrm{fin}})}$. For the grand-canonical boundary condition $c_{\mathrm{fin}}=1$, the black-hole and soliton branches can cross with different slopes. For the example considered below, the black-hole branch satisfying the response conditions ends at a maximum temperature, and the coexistence curve has a $T\to0^+$ endpoint at two opposite values of the spatial Maxwell source. At a nonextremal crossing, the entropy difference is $S_b-S_s=S_b$. For these two static solutions, equality of both current components requires both charges to vanish.

gr-qc↗