Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Symmetries on vector parking functions via bounded lattice paths

Partly motivated by enumeration of parking functions and their variants, there is a long-standing interest in lattice paths refined by several statistics, including the notable $\mathsf{run}$ and $\mathsf{return}$. We consider bounded lattice paths, which are bounded by a given lattice path and are related to vector parking functions, for which we generalize $\mathsf{run}$ to composition runs, parameterized by a composition. By constructing involutions on these paths, we establish symmetries relating composition runs to some generalized return statistics. As an application, we settle an open problem of Dai, Fu, and Qiu on rational Dyck paths. The symmetries between generalized $\mathsf{run}$ and $\mathsf{return}$ are then transferred to vector parking functions, which suggest a new notion of prime decomposition. In the special case of $(a, b)$-parking functions, we compute the generating function refined by $\mathsf{run}$ and $\mathsf{pri}$, which is new even for classical parking functions.

math.CO↗

Scalar Communication via Random Direction Refreshing for Distributed Optimization

Distributed optimization over networks requires agents to repeatedly exchange decision variables with their neighbors. When the decision dimension $d$ is large, these exchanges dominate the communication cost, which is critical for bandwidth-constrained agents. Existing remedies quantize or sparsify the exchanged vectors, yet each message still scales with $d$ and the compression error must be compensated by additional states. To address this limitation, we propose a scalar-communication mechanism in which every neighbor message carries a single real number regardless of $d$. Agents regenerate a common random direction from a shared seed, transmit only the inner product of their state with that direction, and act on the resulting rank-one surrogate of their neighbors' states while retaining full local gradients. We develop and analyze the mechanism for an existing continuous-time distributed optimization algorithm. For strongly convex local costs with Lipschitz gradients, we show that the optimizer remains the unique consensus equilibrium, that a fixed direction admits spurious equilibria, and that refreshing the direction at a sufficiently high rate yields exponential mean-square and almost-sure convergence with constant gains and no residual error. The framework admits any isotropic fixed-norm direction distribution, including Rademacher, scaled-coordinate, and sphere-normalized Gaussian directions; all three attain lower fresh-encoding variance than unnormalized Gaussian directions. The effects of the direction distribution and the refresh interval are illustrated in~simulations.

cs.MA↗

Constraints on the Origin of Universal $1/f^2$ Photon-Count Spectra at Baseband

A series of 11.6-d-duration photon-counting experiments, employing a broad variety of light sources with different statistical properties and optical spectra, were carried out over a 1.5-yr period. All of the photon-count spectra at baseband followed a common $1/f^2$ form over the frequency range $1 \times 10^{-6} \leqslant f \leqslant 5 \times 10^{-4}$ Hz, corresponding to a timescale range $33$ min $ \leqslant T_f \leqslant 11.6$ d, where $T_f \equiv 1/f$. The lower and upper timescale limits were established by the photodetector noise floor and the duration of the individual experiments, respectively. Unlike ordinary Brownian motion, all of the measured photon-count sample paths exhibited irregular long-timescale fluctuations, with durations ranging from hours to days. It has been established that the photon-count spectra cannot plausibly be ascribed to fluctuations of the current, voltage, or temperature of the source or the detector, nor to technical sources of noise associated with the optical system or local environment. Although the physical origin of the photon-count fluctuations remains unresolved, several heterodox hypotheses are set forth. From a statistical point-of-view, the photon counts appear to follow a doubly stochastic Poisson process with a slowly varying random intensity. Analogous experiments that rely on ionizing radiation and direct-conversion solid-state radiation detectors are proposed.

physics.optics↗

An Algebraic Proof of the Non-Modularity of Solutions to the Canonical Modular Differential Equation

In this paper, we provide an algebraic proof of the condition for the existence of modular form solutions to the Kaneko-Zagier differential equation of weight $k$. Unlike previous representation-theoretic approaches relying on $SL_2(\mathbb{Z})$, our method employs the connection matrices of the principal congruence subgroup $Γ(N)$. By deriving a condition for simultaneous triangularizability from the commutators of the representation matrices and applying Galois theory, we prove that the associated two-dimensional representation of $Γ(N)$ is irreducible when the denominator $m$ of the fraction $(k+1)/6 = n/m$ satisfies $m=1$ or $m \ge 7$. Consequently, we identify the weights that admit no modular form solutions and obtain a complete classification of the dimensions of the spaces of such solutions.

math.NT↗

Universal Improvement of Channel Fidelity via Entanglement Assistance

Communication using quantum channels generally requires encoding and decoding on many identical uses of a quantum channel. In practical settings, decoherence may severely limit our ability to do so, potentially rendering each channel use nonidentical. Given $n$ nonidentical channels, we present an entanglement-assisted encoding and decoding strategy that yields a channel with higher worst-case fidelity than all of the $n$ given channels. The strategy is universal, in the sense that the improvement holds regardless of what channels are given, as long as they satisfy mild assumptions on their initial entanglement fidelities. Our protocol is optimal when the given channels are depolarizing channels, and is essentially the unique one that is universal. This idea can also be extended to classical channels, where shared randomness between the sender and the receiver allows universal enhancement of the probability of correct transmission. We then consider general quantum resource theories, providing a necessary condition and a sufficient condition for the possibility of such universal improvement of quantum resources.

quant-ph↗

Generate What You Can Trust: Content Credibility in Generative Recommenders

Generative recommendation (GR) represents items with semantic IDs (i.e., discrete token sequences) and generates target item tokens as recommendations. Despite its promising results, existing methods predominantly optimize for accuracy while neglecting the credibility of the recommendations they generate. This oversight inevitably exposes users to uncredible content (e.g., fake news) with serious societal consequences, including user distrust, reputation harm to platforms, and broader social instability. To address this critical yet underexplored challenge, we propose CreGR, the first credible GR model that jointly tackles content credibility across the two core stages of GR: tokenization and generation. In the tokenization stage, we design a new credibility-aware tokenizer that explicitly encourages the model to learn discriminative tokens respectively for credible and uncredible items, thereby disentangling credibility signals at the token level. Building on this, in the generation stage, we propose a novel accuracy-preserving and credibility-oriented generator grounded in discrete diffusion. Specifically, we introduce an asymmetric masking probability reduction strategy that selectively diminishes the contribution of tokens associated with uncredible content to the generation process, while leaving tokens encoding user preference signals unaffected so as to preserve recommendation accuracy. Experiments on three real-world datasets demonstrate the effectiveness of CreGR.

cs.IR↗

How Should a Prompt Optimizer Spend a Tight Budget? BudgetAPO with Noise-Adaptive Evaluation

Automatic prompt optimization (APO) has been widely employed to adapt large language models without updating their weights, yielding promising results. However, existing methods such as GEPA and OPRO assume hundreds to thousands of subject-model calls, far more than is practical behind paid, rate-limited APIs. Under tight budgets they fail in two ways: multi-stage pipelines can exhaust the budget and return the seed prompt unchanged, while single-stage methods compare candidates on fixed-size minibatches, regardless of each task's noise. As a remedy, we introduce BudgetAPO, a single-stage optimizer for the tight-budget regime. BudgetAPO incorporates (1) a noise-adaptive rule that sizes the evaluation slice to each task's noise, measured by a short probe; (2) a fixed slice that turns every accept/reject decision into a paired comparison; and (3) a reflective operator that rewrites reasoning strategy and output format jointly. Extensive results across seven benchmarks and five subject models demonstrate that BudgetAPO ranks first on every subject and beats every baseline under Holm-corrected paired tests, while returning the seed in 13% of runs at 250 calls against 86% for GEPA. On GPT-OSS-20B, GEPA needs 4.5 times as many calls to match BudgetAPO's 100-call score.

cs.AI↗

The Differential Hilbert Operator Between Weighted Bergman Spaces

In this paper, a complete characterization of the boundedness, compactness, norm and essential norm of the differential Hilbert operator $\mathcal{H}_2:A^2_α\to A^2_β$ is obtained. More precisely, $\mathcal H_2:A^2_α\to A^2_β$ is bounded if and only if $-1<α<0$ and $β\geqα+2$ and it is compact if and only if $-1<α<0$ and $β>α+2$. The norm and essential norm of $ \|\mathcal H_2\|_{A^2_α\to A^2_β}$ are also investigated. In particular, when $β=α+2$, \[ \|\mathcal H_2\|_{A^2_α\to A^2_{α+2}} =\|\mathcal H_2\|_{\mathrm e,A^2_α\to A^2_{α+2}} =\frac{π\sqrt{(α+2)(α+3)}} {\sin(π(α+2)/2)},\qquad -1<α<0. \] When $β>α+2$, $\|\mathcal H_2\|_{\mathrm e,A^2_α\to A^2_β}=0$. Furthermore, for every $1\leq p<\infty$, the operator $\mathcal H_2$ belongs to the Schatten class $\mathcal S_p$ if and only if it is compact.

math.CV↗

LocAttMamba: A Low-Complexity Mamba Framework with Attention-Based Multi-AP Fusion for Indoor Localization

Accurate and low-complexity indoor localization is important for location-based services in fifth generation (5G) and sixth generation (6G) networks, where positioning devices operate under limited computational budgets and non conditions. Indoor localization has been studied widely using traditional signal-level localization approaches. However, these techniques often show degraded performance in on-line-of-sight (NLoS) scenarios. Recently, artificial intelligence (AI)-based techniques, including transformers, have been applied to address these challenges. While transformer-based architectures can capture the dependencies within the measurements collected from distributed access points (APs), their high computational complexity results in a large number of multiply-accumulate operations and long inference time. In contrast, lightweight recurrent and convolutional models trade this cost for degraded accuracy. In this paper, we propose LocAttMamba, a low-complexity localization framework in which the channel impulse response (CIR) and time-based features of each AP are processed by a separate Mamba encoder with near-linear complexity, and the resulting per-AP embeddings are fused through a multi-head attention layer that weights each AP according to its importance at every time step. The framework jointly predicts the user location and its per-axis uncertainty, which is refined through a post-hoc calibration step. We evaluate the proposed framework using two real-world 5G and ultra-wideband (UWB) measurement datasets. Our numerical results reveal that LocAttMamba obtains a mean two-dimensional (2-D) positioning error of 1.004 m and 0.599 m, respectively, on 5G and UWB datasets, outperforming the second-best benchmark by 9.79% and 5.82%, while requiring the fewest multiply-accumulate operations among all evaluated models and being 4.4-16 times faster than the transformer-based benchmarks.

cs.NI↗

Bayesian Data Augmentation for DNN Retraining with Binomial Outcomes in Vision-Based UAV Landing

In GPS-denied or cluttered urban environments, vision-based landing is essential for reliable UAV missions. Real-world landing sites are often unstructured and highly variable, requiring strong generalization by the perception system. Deep Neural Networks (DNNs) trained with synthetic data augmentation offer a scalable solution for learning landing-site features across diverse vehicle and environmental states. However, computationally expensive DNN retraining, along with challenging performance validation via test flights, limits exhaustive model fine-tuning and necessitates an optimized retraining pipeline. In this work, we deploy a Bayesian data augmentation framework integrated with a photorealistic simulator featuring high-fidelity vehicle dynamics to iteratively retrain the helipad detector DNN, maximizing landing performance as the objective function. We validate our framework with experiments in a photorealistic simulator under different environmental conditions and vehicle states, demonstrating improved landing performance and tighter confidence intervals on predicted landing outcomes.

cs.RO↗

A Superprocess-based Approach to Rough CIR Processes and Feller Random Measures

Feller random measures generalize the Feller diffusion (the CIR process) by giving it memory. They arise as the scaling limits of nearly unstable Hawkes processes, and include the rough CIR process and its hyper-rough and discontinuous relatives. We show that every Feller random measure is the occupation measure of a Dawson--Watanabe superprocess whose spatial motion is a killed Lévy subordinator, integrated over the branching time. This is a continuum analogue of the Hawkes--Oakes cluster representation. It yields existence, stability in the parameters, and stochastic equations driven by an explicit noise. The structure of this noise depends on whether the kernel has an atom at the origin. Without an atom, the noise is a Brownian motion time-changed by the distribution function of the measure. This leads to a martingale characterization and to sharp results on densities and their Hölder regularity. With an atom, the noise is a compensated inverse-Gaussian process time-changed by the compensator of that distribution function, and the measure is purely atomic.

math.PR↗

An evolutionary origin of collective decision making in humans and machines

Groups of individuals can solve collective problems more accurately than any single member, by aggregating their opinions. Recent theoretical work has identified individual-level reward schemes that allow uninformed individuals to evolve collective intelligence from the bottom up, through social learning. Yet these results are restricted to linear prediction problems and simple averaging, while the decision tasks that real groups confront are often non-linear, and the institutions that aggregate opinions are seldom single-layer averages: districts elect representatives who in turn vote on policy, referees advise editors who decide on publication. Here we develop a framework for the evolution of collective intelligence in multi-layer voting populations, where individuals observe limited information and groups recursively aggregate their opinions by majority rule. We prove that single-layer voting cannot solve non-linear classification problems under any individual reward scheme. We then identify a "marginal feedback" payoff structure, which rewards individuals only when their opinion is pivotal in their group, and at every layer above them. This reward scheme induces a layered population to evolve accurate collective solutions to complex, non-linear decision tasks through individual-level peer imitation alone. The collective behavior that emerges is equivalent to a multi-layer perceptron in machine learning. Our results provide a naturalistic account of hierarchical institutions, in which the outsize importance of swing voters is the incentive that sustains collective accuracy; and they identify the credit-assignment rule in machine learning as not just an engineered solution but a natural evolutionary outcome.

physics.soc-ph↗

Learning Coordinated Visuomotor Box-Pushing from Solo Demonstrations

Multi-robot imitation learning, particularly in settings where visuomotor policies are deployed in a communication-free, onboard decentralised style, represents an attractive paradigm. However, its realisation remains insufficiently understood, largely due to the difficulty of collecting collective demonstrations, since a single operator cannot control many robots simultaneously. Meanwhile, unlike coupled collaborative manipulation, many coordinated tasks achieve system-wide efficiency primarily through minimising inter-robot interference. This structure motivates us to study whether data collected by a teleoperated single-robot can be leveraged for large-scale coordinated box-pushing as a testbed. We systematically investigate dataset creation strategies and lightweight policy architectures. In particular, experiments with up to 40 robots highlight the difficulty of acquiring effective coordination solely through passive observation of other operating robots, revealing a concrete bottleneck for multi-robot research.

cs.RO↗

Dataset-Free Compliant Humanoid Loco-Manipulation with Dynamic Online Posture

Most humanoid loco-manipulation controllers require human motion data to learn whole-body coordination and posture, leaving policies reliant on external sources to provide this data. We present OCLO (Online-posture Compliant LOco-manipulation), a humanoid loco-manipulation system trained without human motion data and commanded only through two end-effector targets. Because these targets do not uniquely determine whole-body posture, OCLO generates pelvis height and torso orientation online using an analytic reachability prior, further refined through policy-in-the-loop sampling with a task-agnostic cost. OCLO also learns whole-body compliance by displacing end-effector references according to measured forces through a spring-damper model, encouraging the legs, waist, and pelvis to yield to external loads. In simulation, using the reachability prior leads to a 77.8% success rate in acquiring the commanded reference, a vast improvement over the 37.8% success rate accomplished without the prior. Further, refinement reduces end-effector orientation error across all evaluated tasks. The same posture module improves a pretrained SONIC controller on four of five tasks. Without compliance training, policies tend to lose balance under disturbances rather than sacrifice tracking. On a Unitree G1, OCLO maintains balance under end-effector disturbances that cause its ablations to fail and performs seven loco-manipulation tasks, including crouched walking and picking up a box from a low surface. Project website: https://oclo-humanoid.github.io/

cs.RO↗

Sharp Integrality Gaps in Calibration Distance

We study the offline gap between deterministic calibration distance C and its fractional relaxation L for binary unit-weight sequences under total absolute-change cost. We sharpen the offline comparison C <= L + O(sqrt(T)) (Qiao and Zheng, 2024, Theorem 2) to the sharp worst-case order Theta(T^(1/3)). If Delta_T is the supremum of C - L over length-T inputs, then T^(1/3)/1000 <= Delta_T <= 41T^(1/3) for T >= 216. The upper bound holds for every input, while each T >= 216 has a rational lower-bound input. For every input with m distinct forecasts, C <= L + m, and the unrestricted-sample worst-case sparse order is Theta(m). For rational forecasts and accuracy, with binary-encoded multiplicities of separately assignable unit identities, a grid-free polynomial-bit-time procedure returns B <= L <= U, U - B < eta, and an exactly calibrated compact repair of cost at most U + m <= L + m + eta.

cs.LG↗

Instrumental Variable Analysis in Underrepresented Subpopulations Powered by Knowledge Transfer

Instrumental variable (IV) analysis in underrepresented target populations suffers from low statistical efficiency and weak IV bias due to limited sample size. Leveraging external knowledge from a source population with larger samples offers an appealing solution to this problem. We propose in this paper the KNowledge-powered IV Estimator for underrepresented Subpopulations (KNIVES), a framework that enhances the IV analysis for underrepresented subpopulations by transferring knowledge from external source population data. KNIVES constructs a knowledge-transferred IV score by combining multiple candidate models through an effective signal-to-noise ratio criterion tailored to the downstream IV estimator, rather than prediction accuracy alone. To address weak associations between the IV score and the exposure, as well as a high correlation among candidate models in the target population, KNIVES employs a bias correction technique to mitigate the resulting bias and ensure robust performance. Theoretical investigation demonstrates that our method is robust to weak IV bias and achieves improved efficiency compared to the causal effect estimator based on any single candidate model. In addition, KNIVES avoids negative transfer by guaranteeing that its asymptotic variance is not larger than that of the estimator relying solely on target population data. Extensive numerical studies demonstrate the superior finite-sample performance of our method over existing methods. Two Mendelian randomization studies on ethnic minority subgroups in UK Biobank data further illustrate the advantages of our method.

stat.ME↗

Automatic Speech Recognition for Low-Resource Sinhala: A Critical Review of Methods, Challenges, and Future Directions

Automatic speech recognition (ASR) for low-resource languages remains a major challenge. Sinhala, the primary language of Sri Lanka with about 16 million speakers, illustrates the difficulty: agglutinative morphology, a 54-phoneme inventory, subject-object-verb (SOV) syntax and scarce annotated speech data limit both conventional and modern ASR systems. This paper presents the first critical review of Sinhala ASR research, tracing its development from Hidden Markov Models (HMMs) through deep neural networks to self-supervised pre-trained models such as wav2vec 2.0, XLS-R, Whisper and Massively Multilingual Speech (MMS). We compare existing Sinhala systems with related low-resource ASR work on Tamil, Malayalam and Hindi in terms of architecture, training data, word error rate (WER) and robustness to real-world acoustic conditions, and we assess self-supervised and transfer learning as responses to scarce labeled data. We show that most reported WERs are not directly comparable because they differ in corpus, data split and scoring, and that the only controlled comparison in the literature attributes an 18.1% relative WER reduction to corpus correction alone. We also discuss context-aware ASR that draws on phonological, syntactic and semantic knowledge. We identify six research gaps: (1) the lack of large annotated corpora covering multiple dialects and acoustic conditions; (2) weak contextual modeling of Sinhala morphosyntax; (3) high WER in real-world conditions; (4) the absence of standardized benchmarks; (5) the lack of parameter-efficient fine-tuning studies; and (6) the absence of annotated code-switched Sinhala-English speech resources. We outline a research agenda to address these gaps, intended as a roadmap for researchers working on Sinhala and other morphologically rich languages.

cs.CL↗

Knowing the Rules, Applying the Rules: Evaluating Language Models on Traditional Chinese Bazi

Knowing domain rules does not guarantee applying them to a case. We study this distinction in traditional Chinese Bazi through 3,000 Chinese multiple-choice questions spanning 14 Theory and 11 Case categories. Six endpoint systems are evaluated, with primary results reported on a 2,492-item model-informed refinement. Theory accuracy exceeds Case accuracy for every system, and gaps of 16.60-29.56 percentage points remain when invalid responses are excluded. The contrast is more specific than a general case-reasoning deficit. Across six systems, Twelve Stages and Nayin reach mean accuracies of 89.10% and 88.62%, while Shensha Basics reaches 75.96%. Within Case, Luck Pillars averages 84.62%, but Career and Family Relations average only 36.98% and 38.19%. Overall rankings also conceal different category strengths. On the original 3,000 items, paired DeepSeek native/disabled comparisons associate native configurations with Theory gains of 6.53 and 12.20 points for Flash and Pro, respectively; Case changes are -3.67 and +1.27 points. These are provider-configuration associations, not isolated causal effects of reasoning. The results motivate task-specific evaluation of cultural-domain applications rather than reliance on aggregate knowledge scores. The benchmark measures agreement with a model-generated, model-verified answer key, not real-world predictive validity. Final-set results are post-selection descriptions, and incomplete provenance and expert validation constrain their interpretation.

cs.CL↗