Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 163 records · Page 9Linked to original sources

Geodesic switches and exceptional times in dynamical Brownian last passage percolation

We consider Brownian last passage percolation evolving dynamically via a discrete resampling procedure. Using $Γ_{(0,0)}^{(n,n),r}$ to denote a geodesic from $(0,0)$ to $(n,n)$ at time $r$, we prove that the expected total number of coarse-grained changes (or "switches") accumulated by $Γ_{(0,0)}^{(n,n),r}$ away from its endpoints during a time interval $[s,t]$ is at most $n^{5/3+o(1)}(t-s)$; we expect the exponent $5/3$ to be tight. Using the above estimate, we establish that the set $\mathscr{T}$ of exceptional times at which a non-trivial bi-infinite geodesic exists a.s. has Hausdorff dimension at most $1/2$. Further, for any fixed direction $θ$, we show that the set $\mathscr{T}^θ\subseteq \mathscr{T}$ of times at which a non-trivial bi-infinite geodesic directed along $θ$ exists a.s. has Hausdorff dimension equal to $0$.

math.PR↗

Soft Gravitons, Hard Truths: Infrared Safety of Particle Processes in a Gravitational-Wave Background

Gravitational waves are thought to propagate unattenuated through matter due to a cancellation between graviton absorption and stimulated emission inferred from leading-order soft-graviton arguments. We revisit this reasoning and show that it fails for the converse problem: the effect of a gravitational-wave background on matter. At leading order, real graviton emission \emph{and} absorption appear to enhance decay rates of unstable particles. By extending the soft-graviton framework describing real and virtual processes in a gravitational wave background, and resumming them to all orders, we show that inclusive decay rates remain essentially unchanged. Bose-enhanced emission and absorption in graviton spectra persist but compensate each other to all orders; the mutual transparency between matter and gravitational radiation hence follows from infrared safety.

hep-ph↗

Disciplined Biconvex Programming

We introduce disciplined biconvex programming (DBCP), a modeling framework for specifying and solving biconvex optimization problems. Biconvex optimization problems arise in various applications, including machine learning, signal processing, computational science, and control. Solving a biconvex optimization problem in practice usually involves heuristic methods based on alternate convex search (ACS), which iteratively optimizes over one block of variables while keeping the other fixed, so that the resulting subproblems are convex and can be efficiently solved. However, designing and implementing an ACS solver for a specific biconvex optimization problem usually requires significant effort from the user, which can be tedious and error-prone. DBCP extends the principles of disciplined convex programming to biconvex problems, allowing users to specify biconvex optimization problems in a natural way based on a small number of syntax rules. The resulting problem can then be automatically split and transformed into convex subproblems, for which a customized ACS solver is then generated and applied. DBCP allows users to quickly experiment with different biconvex problem formulations, without expertise in convex optimization. We implement DBCP in the open-source Python package dbcp, as an extension to the well known domain specific language CVXPY for convex optimization.

math.OC↗

HAGI++: Head-Assisted Gaze Imputation and Generation

Mobile eye-tracking is crucial for capturing human visual attention in real-world and XR settings, supporting research and human-computer interaction. Yet blinks, pupil-detection errors and lighting changes create missing values that hinder gaze analysis. We present HAGI++, a multi-modal diffusion-based imputation method that, for the first time, leverages integrated head-orientation sensors to exploit the natural correlation between head and eye movements. Using a transformer-based diffusion model, it learns cross-modal dependencies between eye and head data and can additionally incorporate wrist/hand motion when such wearable signals are available. Evaluations on the large-scale Nymeria, Ego-Exo4D and HOT3D datasets show that HAGI++ consistently outperforms traditional interpolation and deep-learning time-series imputation baselines. Statistical analysis confirms that its gaze-velocity distributions closely match real human behaviour, yielding realistic imputations. Even when 100% of gaze data are missing (pure gaze generation), HAGI++ exceeds methods that rely on the visual inputs and the methods rely on full-body motion capture by incorporating wrist motion from commercial wearables. Our approach enables more complete, accurate eye-gaze recordings in real-world contexts, enhancing gaze-based analysis and interaction across many applications. Our code is available at https://git.cai.simtech.uni-stuttgart.de/public-projects/HAGI

cs.HC↗

Type II embeddings for $d=6$ Einstein-Maxwell gauged supergravity

Bi-spinor and G-structure methods are used to classify the possible consistent truncations of type II supergravity to $d=6$ Einstein-Maxwell (gauged) supergravity, and its consistent sub-sectors. In the absence of R-symmetry gauging and a tensor multiplet we establish that every supersymmetric Mink$_6$ solution defines an embedding of the $d=6$ theory. Adding a tensor multiplet places restrictions on these embeddings, but embeddings still exist. In the presence of R-symmetry gauging the internal spaces of the embeddings are neither related to Mink$_6$ or AdS$_6$. Under the assumption that the internal space contains a single U(1) isometry housing the $d=6$ gauge field we classify the possible embedding manifolds. We find two classes of embedding for the entire theory, one of which is governed by a Toda-like equation and contains at least one prototype bounded embedding, albeit containing likely un-physical singularities. In the absence of a tensor multiple the classes of embeddings become more permissive, though the PDEs governing them become more complicated in general.

hep-th↗

Constrained convex clustering for interpretable spatial domain detection in spot-based spatial transcriptomics

Popular technologies for generating spatially resolved transcriptomic data measure gene expression at the resolution of a "spot", i.e., a small tissue region 55 microns in diameter. Each spot can contain many cells of different types. In typical analyses, researchers are interested in using these data to identify and profile discrete spatial domains in tissue. In this paper, we propose a new method, DUET, which simultaneously identifies discrete spatial domains and estimates each spot's expected cell-type proportion. This allows the identified spatial domains to be characterized in terms of the underlying expected cell-type proportions, which affords interpretability and biological insight. DUET utilizes a constrained version of model-based convex clustering, and as such, can accommodate Poisson, negative binomial, normal, and other types of expression data. Moreover, our convex clustering-type criterion allows for both the number of clusters and degree of spatial smoothness to be controlled by a single tuning parameter, which can be chosen in a data-driven fashion. Through simulation studies and a real data application, we show that DUET can achieve better clustering and deconvolution performance than some existing methods.

stat.AP↗

NeuroCLIP: Brain-Inspired Prompt Tuning for EEG-to-Image Multimodal Contrastive Learning

Recent advances in brain-inspired artificial intelligence have sought to align neural signals with visual semantics using multimodal models such as CLIP. However, existing methods often treat CLIP as a static feature extractor, overlooking its adaptability to neural representations and the inherent physiological-symbolic gap in EEG-image alignment. To address these challenges, we present NeuroCLIP, a prompt tuning framework tailored for EEG-to-image contrastive learning. Our approach introduces three core innovations: (1) We design a dual-stream visual embedding pipeline that combines dynamic filtering and token-level fusion to generate instance-level adaptive prompts, which guide the adjustment of patch embedding tokens based on image content, thereby enabling fine-grained modulation of visual representations under neural constraints; (2) We are the first to introduce visual prompt tokens into EEG-image alignment, acting as global, modality-level prompts that work in conjunction with instance-level adjustments. These visual prompt tokens are inserted into the Transformer architecture to facilitate neural-aware adaptation and parameter optimization at a global level; (3) Inspired by neuroscientific principles of human visual encoding, we propose a refined contrastive loss that better model the semantic ambiguity and cross-modal noise present in EEG signals. On the THINGS-EEG2 dataset, NeuroCLIP achieves a Top-1 accuracy of 63.2% in zero-shot image retrieval, surpassing the previous best method by +12.3%, and demonstrates strong generalization under inter-subject conditions (+4.6% Top-1), highlighting the potential of physiology-aware prompt tuning for bridging brain signals and visual semantics.

cs.IR↗

Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy

Modern AI accelerators rely on matrix multiply-accumulate units (MMAUs), such as NVIDIA Tensor Cores and AMD Matrix Cores, to accelerate deep neural network workloads. MMAUs expose only instruction-level or API-level interfaces of matrix multiply-accumulate (MMA) operations, while leaving internal floating-point arithmetic behavior undocumented. Consequently, MMAUs across vendors and architectural generations often produce numerical discrepancies for identical inputs, and sometimes exhibit reduced numerical accuracy that can cause training instability. Diagnosing and understanding the root causes of these effects is challenging without white-box models of their arithmetic behavior. This paper proposes closed-loop feature probing (CLFP), a generic and systematic framework for constructing bit-accurate arithmetic behavior models of MMA operations. Based on this framework, we analyze all MMA instructions on ten GPU architectures spanning NVIDIA Volta through RTX Blackwell and AMD CDNA1 through CDNA3, and derive the first bit-accurate arithmetic models for these MMAUs. Our models explain previously observed cross-platform numerical discrepancies and accuracy issues, enable white-box numerical error analysis, reveal four types of precision bottlenecks and one type of numerical asymmetry, and inform software workarounds as well as design suggestions for future MMAUs. This work is open-source at https://github.com/microsoft/MMA-Sim

cs.AR↗

Beyond Accuracy: Behavioral Dynamics of Agentic Multi-Hunk Repair

Automated program repair has traditionally focused on single-hunk defects, overlooking multi-hunk bugs that are prevalent in real-world systems. Repairing these bugs requires coordinated edits across multiple, disjoint code regions, posing substantially greater challenges. We present the first systematic study of LLM-driven coding agents (Claude Code, Codex, Gemini-cli, and Qwen Code) on this task. We evaluate these four state-of-the-art agents on 404 multi-hunk bugs from the PolyHunk dataset, yielding 1,616 repair trajectories for large-scale behavioral analysis. We employ fine-grained metrics to assess localization, repair accuracy, regression behavior, and operational dynamics across agents. We find that localization capability varies substantially, with Codex achieving the highest success rate (75.3%) and Qwen Code the lowest (40.4%). Repair accuracy also differs widely, ranging from 26.98% (Qwen Code) to 92.82% (Claude Code), and consistently declines with increasing bug dispersion and complexity (hunk divergence and spatial proximity). High-performing agents (Claude Code and Codex) demonstrate superior semantic consistency, achieving positive average regression reduction, whereas lower-performing agents often introduce new test failures. Notably, agents do not fail fast; failed repairs consume substantially more resources (33%-440% more input tokens) and require longer execution time (35%-330%). Additionally, we developed Maple to provide agents with repository-level context. Empirical results show that Maple improves repair accuracy of Gemini-cli by ~21% through enhanced localization. By analyzing fine-grained metrics and trajectory-level analysis, this study moves beyond accuracy to explain how coding agents localize, reason, and act during multi-hunk repair. Our findings underscore the impact of bug divergence and spatial proximity on multi-hunk repair success for coding agents.

cs.SE↗

D-GAP: Improving Out-of-Domain Robustness via Dataset-Agnostic and Gradient-Guided Augmentation in Frequency and Pixel Spaces

Out-of-domain (OOD) robustness is challenging to achieve in real-world computer vision, especially in unsupervised domain adaptation scenarios, where shifts in image background, style, and acquisition instruments often degrade model performance. Generic augmentations show inconsistent gains under such shifts, whereas dataset-specific augmentations require expert knowledge and prior analysis. Moreover, prior studies show that neural networks adapt poorly to domain shifts because they exhibit a learning bias to domain-specific frequency components. Perturbing frequency values can mitigate such bias but overlooks pixel-level details, leading to suboptimal performance. To address these limitations, we propose D-GAP, a Dataset-agnostic and Gradient-guided augmentation method for the Amplitude spectrum (in frequency space) and the Pixel values. Unlike conventional handcrafted augmentations, D-GAP computes sensitivity maps in the frequency space from task gradients, which reflect how strongly the deep models respond to different frequency components, and uses the maps to adaptively interpolate amplitudes between source and target samples. We further propose a dual-space augmentation that jointly controls spectral bias and spatial fidelity by introducing a complementary pixel-space blending branch. This way, D-GAP turns augmentation from fixed, random, or manually designed perturbation into a model-response-adaptive intervention. Extensive experimental results show that the proposed method consistently outperforms both generic and dataset-specific domain adaptation methods, improving average OOD performance by +5.3% on four real-world datasets and +1.9% on three benchmark datasets. Code is available at https://github.com/RapidsAtHKUST/D-GAP.

cs.CV↗

Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations

Explanations are often promoted as tools for transparency, but they can also foster confirmation bias; users may assume reasoning is correct whenever outputs appear acceptable. We study this double-edged role of Chain-of-Thought (CoT) explanations in multimodal moral scenarios by systematically perturbing reasoning chains and manipulating delivery tones. Specifically, we analyze reasoning errors in vision language models (VLMs) and how they impact user trust and the ability to detect errors. Our findings reveal two key effects: (1) users often equate trust with outcome agreement, sustaining reliance even when reasoning is flawed, and (2) the confident tone suppresses error detection while maintaining reliance, showing that delivery styles can override correctness. These results highlight how CoT explanations can simultaneously clarify and mislead, underscoring the need for NLP systems to provide explanations that encourage scrutiny and critical thinking rather than blind trust. All code will be released publicly.

cs.CL↗

Generalized Borel Sets

Generalizing classical descriptive set theory opens foundational questions about the Borel hierarchy. In this paper we systematically study those questions, working in the general framework of Polish-like spaces relative to an uncountable cardinal $κ$, possibly singular, satisfying $2^{<κ}=κ$. We provide fundamental properties of the $κ^+$-Borel hierarchy of any regular Hausdorff space of weight at most $κ$, and establish sufficient conditions for its non-collapse. We highlight a unique phenomenon that arises in the case of singular cardinals, namely, the existence of a second, distinct Borel hierarchy, the $κ$-Borel hierarchy: we prove that it is strictly finer than the $κ^+$-Borel hierarchy, and then characterize the precise relationship between the two. Finally, for regular cardinals, we resolve three questions about the behavior of the $κ^+$-Borel hierarchy on subspaces of the generalized Baire space ${}^κκ$, constructing various models via forcing where several nontrivial constellations for the length of the $κ^+$-Borel hierarchy on the space are realized.

math.LO↗

Anyon Quasilocalization in a Quasicrystalline Toric Code

An exactly solvable model of a quantum spin liquid on a quasicrystal, akin to Kitaev's honeycomb model, was introduced in Kim \textit{et al.}, \href{https://doi.org/10.1103/PhysRevB.110.214438}{\text{Phys. Rev. B} \textbf{110}, 214438 (2024)}. It was shown that in contrast to the translationally invariant models, such a spin liquid stabilizes a gapped ground state with a finite irrational flux density. In this work, we analyze the strong bond-anisotropic limit of the model and demonstrate that the aperiodic lattice geometry naturally generates a hierarchy of exponentially separated coupling constants in the resulting toric code Hamiltonian. Furthermore, a perturbative magnetic field leads to anomalous localization properties where an anyonic excitation sequentially delocalizes over subsets of sites forming equipotential contours in the quasicrystal. In addition, certain background flux configurations, together with the underlying geometry, give rise to strictly localized eigenstates that remain decoupled from the rest of the spectrum. Using numerical studies, we uncover the key mechanisms responsible for this unconventional localization behavior. Our study highlights that topologically ordered phases, in the presence of geometrical constraints can lead to highly anomalous localization properties of fractionalized charges.

cond-mat.str-el↗

SkillWrapper: Generative Predicate Invention for Task-level Robot Planning

Generalizing from individual skill executions to long-horizon tasks is a core challenge in building autonomous robots. A promising direction is learning high-level, symbolic representations of low-level robot skills, enabling abstract reasoning independent of the low-level state space. Recent advances in foundation models have made it possible to generate symbolic predicates that operate on raw sensory inputs-a process we call generative predicate invention-to facilitate downstream representation learning. However, prior work learns these abstractions using heuristic or ad-hoc procedures, leaving unclear which formal properties they ought to satisfy, and how these properties can guide representation learning. We address these questions by characterizing conditions under which learned representations support sound and complete task-level planning, and using them to guide the design of SkillWrapper, a system that autonomously learns symbolic representations of black-box skills without predefined tasks, predicates, or operators. Our approach leverages foundation models to actively collect robot data and learn human-interpretable, plannable representations directly from RGB observations. Our extensive empirical evaluation in simulation and on real robots shows that SkillWrapper learns abstract representations that enable robots to compose black-box skills to solve unseen, long-horizon tasks in the real world.

cs.RO↗

Diagram-to-Circuit QNLP for Financial Sentiment Analysis

We study a \emph{QDisCoCirc}-inspired, chunked diagram-to-circuit quantum natural language processing (QNLP) model for three-class sentiment classification of financial texts. In our classical simulations, we keep the Hilbert-space dimension manageable by decomposing each sentence into short contiguous chunks. Each chunk is mapped to a shallow quantum circuit, and the resulting Bloch vectors are used as a sequence of quantum tokens. Simple averaging of chunk vectors ignores word order and syntactic roles. We therefore add a small Transformer encoder over the raw Bloch-vector sequence and attach a CCG-based type embedding to each chunk. This hybrid design preserves physically interpretable semantic axes of quantum tokens while allowing the classical side to model word order and long-range dependencies. The sequence model improves test macro-F1 over the averaging baseline and chunk-level attribution further shows that evidential mass concentrates on a small number of chunks, that type embeddings are used more reliably for correctly predicted sentences. For real-world quantum language processing applications in finance, future key challenges include circuit designs that avoid chunking and the design of inter-chunk fusion layers.

q-fin.GN↗

Reinforcement Learning to Initialize Newton-Raphson for AC Power Flow with Quantum Annealing-Based Environment Updates

The Newton-Raphson (NR) method is widely used for solving power flow (PF) equations due to its quadratic convergence. However, its performance deteriorates under poor initialization or extreme operating scenarios, e.g., high levels of renewable energy penetration. We propose the use of reinforcement learning (RL) to optimize the initialization of NR, and introduce a quantum-enhanced RL environment update mechanism that addresses the combinatorially large action space at each RL timestep by formulating the voltage adjustment task as a Quadratic Unconstrained Binary Optimization (QUBO) problem, solved with an Ising machine. RL initialization is benchmarked against flat start and start from the DC (linearized) PF solution on a standard 4-bus system, Iwamoto's ill-conditioned 11-bus system, and the IEEE 118-bus system under normal and stressed loading and reactive power limits, with verified operational solutions. On all systems, a supervised initializer refined by RL requires fewer NR iterations than flat and DC starts and than the same initializer without RL, for all seeds. For example, on the 118-bus system under normal and stressed loading, it reached 2.04 and 2.86 NR iterations, compared with 3.02 and 5.13 from DC start and 2.61 and 3.09 without RL. In wall-clock time, this pays off only for an initializer integrated into the solver and reused for many solves on a fixed topology. On the 4-bus system, a quantum-enhanced RL agent with a quantum-inspired annealer moved challenging initial states that required 29 and 44 NR iterations to initializations that required three NR iterations within one RL timestep.

eess.SY↗

Anchor to Expand: Semantic Anchoring for Personalized Text-to-Image Diffusion Models

Personalizing text-to-image diffusion models extends pretrained models to represent novel user-specific concepts from only a few reference images. However, learning a new concept while building on the prior knowledge of the pretrained model remains a key challenge. When personalization focuses on learning the target concept, the model tends to overfit the reference examples and degrade its general capability. In contrast, emphasizing prior preservation can hinder capturing distinctive personalized attributes. In this paper, we address this challenge by viewing a personalized concept as an underrepresented concept whose semantic counterpart is well represented in the pretrained model. Rather than treating the learning of a new concept and prior preservation as separate objectives, we reformulate them as a single anchored learning problem. We therefore introduce Semantic Anchoring Personalization (SAP), which keeps concept learning grounded in the pretrained semantic structure while capturing subject-specific attributes. The proposed objective offers a simple yet effective formulation that can be applied across different model backbones without architectural modifications or auxiliary networks. Extensive experiments across various settings demonstrate that SAP achieves a better balance between subject fidelity and text-image alignment than baseline methods. Further ablation studies validate the contribution of semantic anchoring to personalization.

cs.CV↗