Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 127 records · Page 7Linked to original sources

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation

We introduce Robust Filter Attention (RFA), a formulation of self-attention as a robust state estimator. Each token is treated as a noisy observation of a latent trajectory governed by a linear stochastic differential equation (SDE), and attention weights are determined by consistency under this model rather than static feature similarity. Under isotropic noise and decay assumptions, RFA matches the computational complexity of standard attention. On language modeling benchmarks, RFA achieves lower perplexity than RoPE within the training window while remaining stable under zero-shot extrapolation to longer contexts. The framework also provides a dynamical interpretation of standard positional mechanisms, connecting rotational embeddings and recency biases to transport and uncertainty propagation induced by dynamics.

cs.LG↗

Local points on twists of $X(p)$ with applications

Let $E/\mathbb Q$ be an elliptic curve and $p \geq 3$ a prime. The modular curve $X_E^-(p)$ parametrizes elliptic curves with $p$-torsion modules anti-symplectically isomorphic to~$E[p]$. We give a complete classification of when $X_E^-(p)(\mathbb{Q}_\ell)$ is non-empty, for all primes $\ell$. We give two different applications. First, we classify CM curves $E/\mathbb{Q}$ for which the modular curve $X_E^-(p)$ is a counterexample to the Hasse principle for infinitely many~$p$. Assuming the Frey--Mazur conjecture, we prove that for at least $60\%$ of rational elliptic curves $E$, the modular curve $X_E^-(p)$ is a counterexample to the Hasse principle for at least $50\%$ of primes~$p$. Secondly, we introduce a new technique to the elimination stage of the modular method and apply it to show that $x^3+y^3=5^αz^p$ has no non-trivial primitive solutions for various primes $p$ satisfying $(α/p)=-1$. Moreover, as a by-product of our work, we simplify the assumptions of several local symplectic criteria due to the first author and Alain Kraus.

math.NT↗

Stochastic Analysis of Overlapping Generations Models Under Incomplete Markets

We provide a stochastic analysis of an overlapping generations model under incomplete markets. By casting individual optimization with idiosyncratic income risk into a forward backward stochastic differential equation (FBSDE) system, we (i) establish existence and uniqueness of the dynamic general equilibrium interest rate and (ii) derive analytical and semi explicit formulas for both the equilibrium interest rate path and the natural borrowing limit; defined as the discounted expected shortfall of future income. Our FBSDE based approach yields tractable policy functions and equilibrium mappings without relying on high dimensional PDE methods, offering clear insights into how income dynamics and demographic structure drive interest rate fluctuations and credit constraints.

math.PR↗

An Efficient Subspace Algorithm for Federated Learning on Heterogeneous Data

This work addresses the key challenges of applying federated learning to large-scale deep neural networks, particularly the issue of client drift due to data heterogeneity across clients and the high costs of communication, computation, and memory. We propose FedSub, an efficient subspace algorithm for federated learning on heterogeneous data. Specifically, FedSub utilizes subspace projection to guarantee local updates of each client within low-dimensional subspaces, thereby reducing communication, computation, and memory costs. Additionally, it incorporates low-dimensional dual variables to mitigate client drift. We provide convergence analysis that reveals the impact of key factors such as step size and subspace projection matrices on convergence. Experimental results demonstrate its efficiency.

cs.LG↗

The Game is the Game: Dynamic network analysis and shifting roles in criminal networks

Objectives: This paper incorporates time as a crucial variable to identify key players in criminal networks and explores how actors' positions change over time. It then assesses the accuracy of the results against the uncertainty around network data collected from criminal justice records. Methods: Network data are from a judicial document for a two-year investigation targeting a drug trafficking and distribution network. We use Katz centrality in its dynamic version to explore changes in relationships and relative importance of network actors. We then use a novel method of introducing new edges to the network using Bernoulli random trials to simulate missing data and assess the extent to which node rankings based on Katz centrality change or remain the same when introducing some level of uncertainty to our observed network. Results: We identify actors who consistently held a central role over the course of the two-year investigation and differentiate them from actors who provided key contributions to the group's activities, but only for a limited period. We show that compared to centrality measures commonly used in criminal network analysis, dynamic Katz centrality is helpful to differentiate individual contributions even among central nodes and explore individual trajectories over time, even when data are incomplete. Conclusions: This paper demonstrates the value of key player identification using temporal network data and offers an additional analytical tool to both organised crime scholars trying to capture the complex nature of criminal collaboration and law enforcement agencies aiming at identifying appropriate targets and disrupting criminal groups.

cs.SI↗

Enabling Regulatory Multi-Agent Collaboration: Architecture, Challenges, and Solutions

Large language models (LLMs)-empowered autonomous agents are transforming both digital and physical environments by enabling adaptive, multi-agent collaboration. While these agents offer significant opportunities across domains such as finance, healthcare, and smart manufacturing, their unpredictable behaviors and heterogeneous capabilities pose substantial governance and accountability challenges. In this paper, we propose a blockchain-enabled layered architecture for regulatory agent collaboration, comprising an agent layer, an off-chain computation layer, and an on-chain anchoring layer. Within this framework, we design three key modules: (i) an agent behavior tracing and arbitration module for automated accountability, (ii) a dynamic reputation evaluation module for trust assessment in collaborative scenarios, and (iii) a malicious behavior forecasting module for early detection of adversarial activities. Our approach establishes a systematic foundation for trustworthy, resilient, and scalable regulatory mechanisms in large-scale agent ecosystems. Finally, we discuss the future research directions for blockchain-enabled regulatory frameworks in multi-agent systems.

cs.AI↗

Optimal Transport Based Testing in Factorial Designs

We introduce a general framework for testing statistical hypotheses in factorial designs for probability measures supported on discrete spaces. The suggested methodology is based on the pairwise comparison of measures using optimal transport (OT). The formulation of hypotheses is intuitive: It is a direct extension of those underlying the analysis of variance (ANOVA) and its nonparametric counterparts to test for linear relationships between (discrete) probability measures in factorial designs. To this end, means or cumulative distribution functions simply will be replaced by measures. We derive under the null hypotheses and under (local) alternatives the asymptotic distribution of the corresponding empirical OT test statistic, which is the optimal value of a linear program with random objective function. It turns out that this requires to extend existing techniques from probability measures to signed measures, and we show directional Hadamard differentiability and the validity of the functional delta method. We discuss computational issues, permutation and bootstrap tests, and back up our findings with simulations. We illustrate our methodology on datasets from cellular biophysics and from biometric fingerprint identification.

math.ST↗

TransforMARS: Fault-Tolerant Self-Reconfiguration for Arbitrarily Shaped Modular Aerial Robot Systems

Modular Aerial Robot Systems (MARS) consist of multiple drone modules that are physically bound together to form a single structure for flight. Exploiting structural redundancy, MARS can be reconfigured into different formations to mitigate unit or rotor failures and maintain stable flight. Prior work on MARS self-reconfiguration has solely focused on maximizing controllability margins to tolerate a single rotor or unit fault for rectangular-shaped MARS. We propose TransforMARS, a general fault-tolerant reconfiguration framework that transforms arbitrarily shaped MARS under multiple rotor and unit faults while ensuring continuous in-air stability. Specifically, we develop algorithms to first identify and construct minimum controllable assemblies containing faulty units. We then plan feasible disassembly-assembly sequences to transport MARS units or subassemblies to form target configuration. Our approach enables more flexible and practical feasible reconfiguration. We validate TransforMARS in challenging arbitrarily shaped MARS configurations, demonstrating substantial improvements over prior works in both the capacity of handling diverse configurations and the number of faults tolerated. The videos and source code of this work are available at https://github.com/RuiHuangNUS/TransforMARS

cs.RO↗

BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots

Scale-consistent ego-motion estimation is fundamental for autonomous ground robots. Bird's-Eye-View (BEV) representation naturally addresses the scale drift problem of monocular visual odometry (MVO) by providing a metric-scaled planar workspace, enabling the simplification of 6-DoF ego-motion to a more robust 3-DoF model. However, existing BEV-based methods suffer from two key limitations: sparse supervision signals from pose-only training, and information loss during perspective-to-BEV projection. We present BEV-ODOM2, an enhanced framework in the BEV-ODOM line that addresses both limitations without supervision modalities beyond the pose ground truth. Our approach introduces (1) pose-derived dense rigid BEV flow supervision, which reparameterizes the 3-DoF pose ground truth into a pixel-level training signal, and (2) Perspective View (PV)-BEV fusion, which computes correlation volumes before projection to retain additional motion cues and help alleviate projection ambiguity. An enhanced rotation sampling strategy further balances diverse motion patterns during training. We evaluate on four datasets with varied spatial scales: KITTI, Oxford, NCLT, and our newly collected ZJH-VO benchmark. BEV-ODOM2 reduces the RTE of BEV-ODOM by 39% on average across the four datasets. It further enables closed-loop navigation on a physical robot with centimeter-level cross-track error. Real-time inference on an NVIDIA Jetson AGX Orin confirms edge deployment feasibility. The code and the ZJH-VO dataset are publicly released to facilitate future research.

cs.RO↗

Formation of Cavity-Polaritons via High-Order Van Hove Singularities

We consider polaritons formed by hybridizing a continuum of interband particle-hole excitations of an insulating phase with a cavity photon at subgap frequencies, where absorption is suppressed. The strength of the hybridization is driven by the Van Hove singularity in the joint density of states (JDOS) at the band gap: the stronger the singularity, the more a photon is hybridized with the interband transitions. In order to increase the singularity and thus the polariton hybridization without absorption, we propose to engineer a nonparabolic momentum dispersion of the bands around the gap in order to implement a high-order Van Hove singularity (HOVHS) in the JDOS. Ultracold atoms in tunable optical lattices are an ideal platform to engineer two-dimensional gapped phases with nontrivial band dispersions at the gap. Moreover, the intrinsic noninteracting nature of polarized fermionic atoms prevents the emergence of subgap excitations, which are common in solid-state systems and could otherwise spoil the absence of absorption below the gap. Our findings identify band-engineering at the gap edge as a promising route for polariton control with applications in quantum nonlinear optics.

cond-mat.quant-gas↗

eVGGT: An Efficient Geometry-Aware Vision Encoder for Visuomotor Policies

Geometry-grounded vision models, such as VGGT, have emerged as robust visual encoders, providing essential geometric priors for robotic manipulation. However, the high computational cost of these models often leads to slow inference, limiting their practical applications in real-world robotics. This paper introduces eVGGT, a lightweight geometry-aware vision encoder distilled from the high-performing VGGT. Our findings demonstrate two primary advantages: i) integrating eVGGT into imitation learning frameworks (including ACT and Diffusion Policy) yields up to a 6.3% improvement in success rate over standard 2D encoders across bimanual and single-arm tasks in both simulation and real-world settings with variable viewpoints; ii) eVGGT achieves a nearly 5 times speedup and a 63% reduction in memory usage compared to state-of-the-art geometry-aware encoders while maintaining comparable task performance. These results suggest that eVGGT substantially alleviates the performance-latency bottleneck that has limited geometry-aware visuomotor policies in real-world deployment.

cs.RO↗

Matricial ranges, dilations, and unital contractive maps

Let $J_n$ be the Jordan block of size $n$ with all eigen values zero. Arveson introduced the notion of the matricial range of an operator in his remarkable article called Subalgebras of $C^*$-algebras II (Acta Math, 128, 1972) and established that every unital positive map on the operator system generated by $J_2$ is completely positive. This describes the matricial range of $J_2$ as the set of all matrices with numerical radius at most $\frac{1}{2}$. Later, Choi and Li generalize this result of Arveson and prove that every unital positive map on the operator system generated by any $2\times 2$ matrix or any $3\times3$ matrix with a reducing subspace is completely positive. After fifty years of the above result of Arveson, the matricial range of $J_n$ for $n\geq 3$ has not been characterized. This article aims to investigate this long-standing open problem for $n=3$. We begin by establishing a structure theorem for a dilation of an operator $B$ satisfying $BB^*+B^*B=I$ and then investigate whether every $B\in\mathbb{M}_n$ satisfying $BB^*+B^*B\leq I_n$ admits a dilation $\widetilde{B}$ for which $\widetilde{B}\widetilde{B}^*+\widetilde{B}^*\widetilde{B}=I$. This study plays the central role to the development of this paper. We use this to prove that every unital contractive map on the operator system generated by $J_3$ is $2$-positive and obtain some partial results towards characterizing the matricial range of $J_3$. Next, we study unital contractive maps on operator systems generated by $4\times 4$ normal matrices, and show that this is equivalent to studying a unital contractive map on the operator system generated by $T=\text{diag}(λ,-1,i,-i)$, where $\Re{(λ)}\geq 0$. We prove that every unital contractive map on the operator system generated by $T=\text{diag}(1,-1,i,-i)$ is completely positive.

math.FA↗

The Platonic Universe: Do Foundation Models See the Same Sky?

We investigate when foundation models converge towards shared representations, and how this convergence depends on model capacity, training regime, and model architecture. We take a `science-for-AI' approach, using astronomy as an experimental instrument to test the Platonic Representation Hypothesis and its Aristotelian refinement against an external physical reference. The historical success of astrophysics is evidence that a compact, modality-invariant description of galaxy observables exists, and so representation convergence toward reality should be measurable against the physical parameters astronomers already use. Given this framework, we evaluate eleven foundation model families (spanning classification, self-distillation, joint-embedding prediction, autoencoding, vision-language pre-training, and astro-specific architectures from $\mathcal{O}$(10M)${\to}\mathcal{O}$(10B) parameters) on crossmatched JWST, HSC, and Legacy imagery, and DESI spectroscopy. All models are evaluated frozen, with no astronomy-specific fine-tuning. We probe redshift, stellar mass, and sSFR via linear probes, and local (MKNN) and global (CKA) embedding geometry within families, between modalities, and across architectures. We find that physics performance scales predictably with capacity; probe directions align consistently with expected astrophysical correlations and selection effects; and local (not global) embedding alignment tracks physics performance, including between DESI spectra and HSC imagery---modalities that share essentially no low-level statistics. Our results support the ARH over the strict PRH, demonstrate astronomy's value as an experimental framework for neural representation learning, and suggest that astro-foundation models can build on general-purpose pre-trained architectures, capitalizing on the broader open machine learning community's already-spent computational investment.

astro-ph.IM↗

Maximum principle for robust utility optimization via Tsallis relative entropy

This paper investigates an optimal consumption-investment problem featuring recursive utility via Tsallis relative entropy. We establish a fundamental connection between this optimization problem and a quadratic backward stochastic differential equation (BSDE), demonstrating that the value function is the value process of the solution to this BSDE. Utilizing advanced BSDE techniques, we derive a novel stochastic maximum principle that provides necessary conditions for both the optimal consumption process and terminal wealth. Furthermore, we prove the existence of optimal strategy and analyze the coupled forward-backward system arising from the optimization problem.

q-fin.MF↗

One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning

Large Language Models (LLMs) are increasingly deployed across multilingual and multicultural settings, yet it remains unclear whether changing language leads models to adopt community-specific moral reasoning or merely changes how shared learned abstractions are expressed. We conduct a controlled multilingual evaluation across six geographically, culturally, and linguistically diverse languages (Arabic, Chinese, English, Hindi, Russian, and Spanish), using parallel moral reasoning benchmarks with English-origin, Chinese-origin, and natively elicited ground-truth judgments. Across 13 open-weight LLMs spanning 2B-70B parameters, we find substantial cross-lingual divergence in moral judgments, with English generally achieving the highest performance even when ground-truth judgments originate in Chinese or are collected natively in each language. Yet the reasoning underlying these divergent judgments is considerably more convergent: Utilitarianism dominates in five of six languages, reasoning follows broadly shared stages, and language-specific moral-value associations correspond only sparsely and inconsistently to values measured in the corresponding human communities. Finally, a large-scale OLMoTrace analysis of pretraining data sources reveals little direct reproduction of training text across languages, while the corpus composition, training stage, and cultural provenance of retrieved training evidence vary substantially by response language. Thus, similar moral reasoning structures emerge even from heterogeneous and often linguistically localized training evidence. Our findings, collectively, reveal a central disconnect in multilingual moral reasoning: language changes models' moral judgments and the training evidence associated with their reasoning, but does not correspondingly localize the moral abstractions they apply.

cs.CL↗

WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning

Technical-domain benchmarks constructed from arXiv papers can overlap the same public text used in LLM pretraining. Auditing this risk at training-corpus scale requires searching billions of corpus n-grams while retaining per-item evidence that users can inspect and recompute. We contribute a reverse-probe audit at a fixed 13-gram threshold: it indexes benchmark prompts, streams public pretraining corpora, and emits per-problem prompt-surface lexical-overlap metadata with memory that scales with the benchmark. We instantiate the protocol in WirelessMathBench-XL, a 4,027-problem wireless mathematical-reasoning benchmark built from 836 retained arXiv papers across 20 subfields. Against 12.9B streamed 13-grams from RedPajama-arXiv, the audit identifies a strict zero-hit view S0 covering 3,853 problems (95.7%). Filtering to S0 changes accuracy by less than 1 pp for every evaluated model; frontier calibration rows form one high-accuracy cluster between 86.5% and 91.3%, not a resolved rank order. Only 30/800 test items carry detected overlap. Under an all-flagged-correct counterfactual, their largest possible positive score inflation is 0.31-0.51 pp for the frontier rows, so full-versus-S0 is a bounded, structurally underpowered stability summary rather than a contamination-effect test or cleanliness claim. The audit channel does not cover paraphrase, target-answer, post-training, or closed-corpus exposure. The release includes source-paper identifiers, verifier-facing ground truths, audit and threshold metadata, filtered views, a paper-disjoint sensitivity view, evaluation traces, paired-bootstrap scripts, training recipes, Croissant metadata, and a Datasheet for Datasets.

cs.LG↗

Graph Your Own Prompt

We propose Graph Consistency Regularization (GCR), a novel framework that injects relational graph structures, derived from model predictions, into the learning process to promote class-aware, semantically meaningful feature representations. Functioning as a form of self-prompting, GCR enables the model to refine its internal structure using its own outputs. While deep networks learn rich representations, these often capture noisy inter-class similarities that contradict the model's predicted semantics. GCR addresses this issue by introducing parameter-free Graph Consistency Layers (GCLs) at arbitrary depths. Each GCL builds a batch-level feature similarity graph and aligns it with a global, class-aware masked prediction graph, derived by modulating softmax prediction similarities with intra-class indicators. This alignment enforces that feature-level relationships reflect class-consistent prediction behavior, acting as a semantic regularizer throughout the network. Unlike prior work, GCR introduces a multi-layer, cross-space graph alignment mechanism with adaptive weighting, where layer importance is learned from graph discrepancy magnitudes. This allows the model to prioritize semantically reliable layers and suppress noisy ones, enhancing feature quality without modifying the architecture or training procedure. GCR is model-agnostic, lightweight, and improves semantic structure across various networks and datasets. Experiments show that GCR promotes cleaner feature structure, stronger intra-class cohesion, and improved generalization, offering a new perspective on learning from prediction structure. [Project website](https://darcyddx.github.io/gcr/) [Code](https://github.com/Darcyddx/graph-prompt)

cs.LG↗