Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Communication-Aware Synthesis of Safety Controller for Networked Control Systems

This paper studies communication-aware safe control for discrete-time linear multi-agent systems under limited information exchange. The main challenge lies in the coupling between remote-state estimation and safe controller design, since estimation errors affect the state evolution through the controller gains, while the controller design must account for the resulting observer-induced state perturbations to guarantee safety. To address this challenge, a distributed input predictor is incorporated into a $k$-hop state observer to reconstruct unavailable remote states, and the resulting estimation errors are characterized to quantify the observer-induced state perturbations. An overall robust safety invariant (RSI) set is then constructed by jointly accounting for individual agent state constraints and prescribed safety constraints on the relative states of interacting agents under these perturbations. A linear matrix inequality (LMI)-based optimization method is developed to jointly synthesize the distributed observers, local controllers, and the RSI set. A case study illustrates the effectiveness of the proposed method.

eess.SY↗

The Hitchhikers Guide to Rubric Quality Understanding and Enrichment

Rubrics distill notions of expert quality and measure agent performance. However, the quality of rubrics themselves have not been systematically measured and are often left to downstream performance. We import apparatuses from measurement theory built for exactly this: quantitative signals based on the rubric's content, and introduce the RubrIc-Failure Taxonomy (RIFT), of nine possible ways a rubric fails, organized under reliability and content validity. Every mode leaves a distinct signature. To show the signals track failure causally, we seed 720 corruptions, injecting each RIFT mode into clean rubrics at known severity levels. A linear probe over the signals identifies which mode was injected at $75.0\%$ accuracy, beating $56.7\%$ for a frontier model asked to name the failure directly. Surprisingly across GDPval and Terminal-Bench, 10 of 48 expert-authored rubrics weight their criteria backwards, putting more of the score on requirements an expert panel judged less essential. This means a response can fail what matters most and still be graded well. This paper serves as a comprehensive guide on how to understand failure modes in rubrics and create better versions using quality signals, causal experiments, and provides a taxonomy with its rules and examples.

cs.AI↗

Distributed Variational Quantum Linear Solver

This paper develops a distributed variational quantum algorithm for solving large-scale linear equations. For a linear system of the form $Ax=b$, the large square matrix $A$ is partitioned into smaller square block submatrices, each of which is known only to a single noisy intermediate-scale quantum (NISQ) computer. Each NISQ computer communicates with certain other quantum computers in the same row and column of the block partition, where the communication patterns are described by the row- and column-neighbor graphs, both of which are connected. The proposed algorithm integrates a variant of the variational quantum linear solver at each computer with distributed classical optimization techniques. The derivation of the quantum cost function provides insight into the design of the distributed algorithm. Numerical quantum simulations demonstrate that the proposed distributed quantum algorithm can solve linear systems whose size scales with the number of computers and is therefore not limited by the capacity of a single quantum computer.

quant-ph↗

Learning an Interpretable Risk Scoring System for Maximizing Decision Net Benefit

Risk scoring systems are widely used in high-stakes domains to assist decision-making. However, existing approaches often focus on optimizing predictive accuracy or likelihood-based criteria, which may not align with the main goal of maximizing utility. In this paper, we propose a novel risk scoring system that directly optimizes net benefit over a range of decision thresholds. The model is formulated as a sparse integer linear programming problem which enables the construction of a transparent scoring system with integer coefficients, and hence, facilitates interpretation and practical application. We also establish fundamental relationships among net benefit, discrimination, and calibration. Specifically, we derive bounds relating the area under the net benefit curve to a ROC functional, both evaluated on a fixed threshold grid, and show that post-processing can achieve moderate calibration on the training data without decreasing the area under the net benefit curve on that grid. We evaluated our method on multiple public datasets as well as on a large-scale credit risk dataset. This computational study demonstrated that our interpretable method can effectively achieve high net benefit while maintaining competitive discrimination and calibration performance.

cs.LG↗

Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity

Automated fact-checking is a crucial task that supports a responsible information ecosystem. While recent research has progressed from text-only to multimodal fact-checking, a prevailing assumption is that incorporating visual evidence universally improves verification accuracy. In this work, we challenge this assumption and show that the indiscriminate use of visual evidence can reduce accuracy. Building on this finding, we propose AMuFC, a modular fact-checking framework that employs two collaborative vision-language models with distinct roles to enable the adaptive use of visual evidence. Experimental results on three datasets, including WebFC, introduced in this study, demonstrate the effectiveness of adaptive visual evidence use in fact-checking.

cs.CL↗

Measuring Depth of Matroids

Motivated by recently discovered connections between matroid depth measures and block-structured integer programming [ICALP 2020, 2022], we undertake a systematic study of recursive depth parameters for matrices and matroids, aiming to unify recently introduced and scattered concepts. We propose a general framework that naturally yields eight different depth measures for matroids, prove their fundamental properties and relationships, and relate them to two established notions in the field: matroid branch-depth and matroid tree-depth (a natural depth counterpart of matroid tree-width). In particular, we show that six of our eight measures are mutually functionally inequivalent, and among these, one is functionally equivalent to matroid branch-depth and another to matroid tree-depth. We also prove that all these depth measures, possibly except one, coincide on matroids and on matrices over any field, which, perhaps surprisingly, is not a trivial finding. Finally, we provide a comparison between the matroid parameters and classical depth measures of graphs.

math.CO↗

Quantum Algorithms for Heterogeneous PDEs: The Neutron Diffusion Eigenvalue Problem

We develop a quantum algorithm to solve a type of linear reaction-diffusion equation, the neutron diffusion (generalized) k-eigenvalue problem that establishes nuclear criticality. The algorithm handles an equation with piecewise constant coefficients, describing a problem in a heterogeneous medium. We apply uniform finite elements and show that the quantum algorithm provides significant polynomial end-to-end speedup over classical uniform finite element methods in the cases tested, although nonuniform classical methods can perform much better than uniform ones. Our work leverages recent advances in quantum linear systems--fast inversion and quantum preconditioning--and uses Hamiltonian simulation as a subroutine. Our results suggest that quantum algorithms may provide speedups for heterogeneous PDEs, though the extent of this advantage over the fastest classical algorithm depends on the effectiveness of other classical approaches such as nonuniform or adaptive meshing for a given problem instance.

quant-ph↗

$π^2$: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models

We study a QA curation pipeline for improving long-context complex reasoning in large language models (LLMs). Our approach, $π^2$, constructs high-quality reasoning data through rigorous QA curation: 1) extracting and expanding tables from Wikipedia, 2) from the collected tables together with relevant metadata, generating complex reasoning questions whose answers are automatically determined and validated through dual-path code execution, 3) finally, back-translating chain-of-thoughts solutions grounded in realistic context. Supervised fine-tuning with gpt-oss-20b and Qwen3-4B-Instruct-2507 on $π^2$ yields consistent improvements across four long-context reasoning benchmarks and our alike $π^2$-Bench, with average absolute accuracy gains of +6.25% and +3.37% respectively. Through deeper analyses, we observe that reasoning style contributes little, while faithful reasoning patterns discovered by back translation and grounded realistic long context, as $π^2$ is designed for, are crucial for the improvement. Our code, data, and models are fully open-source at https://github.com/vtpss/pi-squared.

cs.CL↗

$ξRϕ^2$ non-minimal coupling, and the long range gravitational potential for different spin fields from 2-2 scattering amplitudes

In this paper, we investigate the long range gravitational effect of curvature-scalar field non-minimal coupling in the form of $ξR ϕ^2$, in the perturbative quantum gravity framework. Such coupling is most naturally motivated from the renormalisation of a scalar field theory with a quartic self interaction in a curved spacetime background. This coupling results in two scalar-$n$ graviton vertices which contain no explicit momenta of the scalar, qualitatively different from the usual, e.g., $κh^{μν}T_{μν}$-type minimal matter-graviton vertices. Assuming the dimensionless coupling parameter $ξ$ to be small, we compute the 2-2 scattering Feynman amplitudes between such scalars up to ${\cal O}(G^2 ξ)$. From the non-relativistic limit of these amplitudes, we compute the corresponding long range gravitational potential. There exists no tree level contribution $({\cal O}(ξG))$ here, and hence the one loop ${\cal O}(G^2 ξ)$ result is leading. Recently, the effect of a cosmological constant in such non-minimal interaction and the subsequent gravitational potential was computed. In this work, we take the cosmological constant to be vanishing. The resulting potential is found to have $r^{-4}$ leading behavior. We further extend these results for scalar-massive spin-1 and massive spin-1/2 scattering. Spin and polarisation dependence of the two body potential have been explicitly demonstrated. We discuss some possible physical implications of these results.

hep-th↗

Pruning the Augmented Graphs of Convex Sets for Scalable Joint Task and Motion Planning

We present top-down and bottom-up approaches for solving large-scale instances of joint task and motion planning problems in Graphs of Convex Sets (GCS). Planning what tasks to perform and how to move between them can be encoded exactly as a Shortest Path Problem (SPP) using an Augmented GCS (AGCS), but this graph grows exponentially with the number of tasks, limiting prior work to only 11 tasks. Both of our approaches significantly improve scalability by carefully pruning the AGCS before it is constructed. The top-down approach replaces the complete graph with a sparse Delaunay graph that maintains a natural nearest-neighbor connectivity, reducing the number of vertices, edges, and subgraphs relative to the full AGCS. The bottom-up approach uses classical Traveling Salesman Problem (TSP) heuristics to create an initial ordering, then expands it within a fixed search window, pruning the AGCS to a bounded maximum width independent of the problem size. Both approaches can obtain optimal or near-optimal solutions in a fraction of the time required by the original AGCS. The bottom-up approach can solve instances with 1000 tasks in about two minutes. We additionally discuss lower bounds based on minimum 1-trees to quantify suboptimality of the proposed approaches.

eess.SY↗

Direct-detection constraints on inelastic dark matter with a scalar mediator

We phenomenologically calculate direct-detection constraints on inelastic dark matter (DM) for a scalar portal scenario with leptophilic couplings. The p-wave velocity suppression of the annihilation cross section of scalar-mediated inelastic Dirac DM implies the opening of viable regions of DM parameter space in the MeV-GeV mass range. Xenon-based experiments can provide a constraints on scalar-mediated inelastic fermion dark matter for sub-MeV mass splitting, via endothermic and exothermic spin-independent DM-electron scattering. To estimate the relevant constraints, we use public data from the XENON1T, PandaX-4T, and LZ liquid-xenon experiments that measure ionization electron signals.

hep-ph↗

What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning

Existing Graphical User Interface (GUI) reasoning tasks remain challenging, particularly in UI understanding. Current methods typically rely on direct screen-based decision-making, which lacks interpretability and overlooks a comprehensive understanding of UI elements, ultimately leading to task failure. To enhance the understanding and interaction with UIs, we propose an innovative GUI reasoning paradigm called UI-in-the-Loop (UILoop). Our approach treats the GUI reasoning task as a cyclic Screen-UI elements-Action process. By enabling Multimodal Large Language Models (MLLMs) to explicitly learn the localization, semantic functions, and practical usage of key UI elements, UILoop achieves precise element discovery and performs interpretable reasoning. Furthermore, we introduce a more challenging UI Comprehension task centered on UI elements with three evaluation metrics. Correspondingly, we contribute a benchmark of 26K samples (UI Comprehension-Bench) to comprehensively evaluate existing methods' mastery of UI elements. Extensive experiments demonstrate that UILoop achieves state-of-the-art UI understanding performance while yielding superior results in GUI reasoning tasks.

cs.AI↗

Space-time correlations of passive scalars in colored-noise flows

The space-time correlation of a passive scalar advected by a Gaussian colored-noise velocity with wavenumber-dependent correlation times and power-law spatial spectra is investigated in the present paper. Within the inertial-convective subrange, we derive an analytical solution for the space-time correlation. This solution validates the elliptic approximation (EA) model [He and Zhang, Phys. Rev. E 73, 055303(R) (2006)], demonstrating that the iso-correlation contours are self-similar in the co-moving space-time frame $(r-Uτ, Vτ)$, with a universal spatial-to-temporal intercept ratio of 1.55. Unlike the classic Kraichnan white-noise model, our formulation simultaneously recovers the Obukhov--Corrsin scaling for spatial correlations (when the velocity obeys Kolmogorov scaling) and reproduces the random-sweeping mechanism, yielding Gaussian (rather than exponential) temporal decorrelation of scalar Fourier modes. Our results clarify the underlying decorrelation mechanism of passive scalars: mean-flow advection and large-scale sweeping dominate temporal decorrelation, and small-scale distortion dominates spatial decorrelation.

physics.flu-dyn↗

Inverse Laplace and Mellin integral transforms modified for use in quantum communications

Integral transformations are useful mathematical tool to work out signals and wave-packets in electronic devices. They may be used in software protocols. Necessary knowledge may come from quantum field theory, in particular from quantum chromodynamics, in which the optic theorem and the renormalization group equation can be solved by a unique contour integral written in two different "dual" ways related between themselves by a complex map in the complex plane of Mellin variable. The inverse integral transformation should be modified to be applied for these contour integral solutions. These modified inverse transformations may be used in security protocols for quantum computers. Here we do a brief review of the basic integral transforms and propose their modification for the extended domains.

quant-ph↗

What do your logits know?

Recent work has shown that probing model internals can reveal a wealth of information not apparent from the model generations. This poses a risk of unintentional or malicious information leakage, where model users are able to learn information that the model owner assumed was inaccessible. Using vision-language models as a testbed, we present the first systematic comparison of information retained at different representational levels as it is compressed from the rich information encoded in the residual stream through two natural bottlenecks: low-dimensional projections of the residual stream obtained using tuned lens, and the final top-k logits most likely to impact model's answer. We show that even easily accessible bottlenecks defined by the model's top logit values can leak task-irrelevant information present in an image-based query, in some cases revealing as much information as direct projections of the full residual stream.

cs.AI↗

Finite-temperature quantum Krylov method from real-time overlaps

Accurately evaluating finite-temperature properties of quantum many-body systems remains a central challenge. Many existing quantum approaches require thermal-state preparation at each target temperature, making low-temperature calculations especially demanding in terms of circuit depth and accuracy. In this work, we present the finite-temperature quantum Krylov (FTQK) method, which combines stochastic trace estimation with a Hermitian generalized eigenvalue problem based on $\cos(\tilde H)$. The only quantum-side input required is the temperature-independent real-time overlap sequence $g_n=\langleϕ|e^{-inτH}|ϕ\rangle$. All benchmark results presented in this work were obtained from classical numerical simulations of the quantum-circuit protocol. For periodic spin-$\frac{1}{2}$ Heisenberg chains with $N=14$ and $18$, FTQK accurately reproduces the specific heat, magnetic susceptibility, and entropy in the noiseless case. It also accurately reproduces the specific heat of the $N=14$ honeycomb Kitaev model, demonstrating that FTQK can be applied beyond simple one-dimensional chains to a model on a two-dimensional lattice with bond-dependent interactions. Furthermore, in a proof-of-principle finite-shot study for the $N=14$ Heisenberg chain, the main thermodynamic features remain well preserved even at $σ=10^{-3}$. Here, $σ=1/\sqrt{N_{\mathrm{shot}}}$, where $N_{\mathrm{shot}}$ is the number of shots used to estimate an overlap. These results demonstrate the potential of FTQK for finite-temperature calculations using reusable, temperature-independent real-time overlap data.

quant-ph↗

Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis

Learned classifiers deployed in agentic pipelines face a fundamental reliability problem: predictions are probabilistic inferences, not verified conclusions, and acting on them without grounding in observable evidence leads to compounding failures across downstream stages. Software vulnerability analysis makes this cost concrete and measurable. We address this through a unified cross-language vulnerability lifecycle framework built around three LLM-driven reasoning stages-hybrid structural-semantic detection, execution-grounded agentic validation, and validation-aware iterative repair-governed by a strict invariant: no repair action is taken without execution-based confirmation of exploitability. Cross-language generalization is achieved via a Universal Abstract Syntax Tree (uAST) normalizing Java, Python, and C++ into a shared structural schema, combined with a hybrid fusion of GraphSAGE and Qwen2.5-Coder-1.5B embeddings through learned two-way gating, whose per-sample weights provide intrinsic explainability at no additional cost. The framework achieves 89.84-92.02% intra-language detection accuracy and 74.43-80.12% zero-shot cross-language F1, resolving 69.74% of vulnerabilities end-to-end at a 12.27% total failure rate. Ablations establish necessity: removing uAST degrades cross-language F1 by 23.42%, while disabling validation increases unnecessary repairs by 131.7%. These results demonstrate that execution-grounded closed-loop reasoning is a principled and practically deployable mechanism for trustworthy LLM-driven agentic AI.

cs.SE↗

Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions

Using behavioural science, health interventions focus on behaviour change by providing a framework to help patients acquire and maintain healthy habits that improve medical outcomes. In-person interventions are costly and difficult to scale, especially in resource-limited regions. Digital health interventions offer a cost-effective approach, potentially supporting independent living and self-management. Automating such interventions, especially through machine learning, has recently gained considerable attention. Ambivalence and hesitancy (A/H) play a primary role for individuals to delay, avoid, or abandon health interventions. A/H are subtle and conflicting emotions that place a person in a state between positive and negative evaluations of a behaviour, or between acceptance and refusal to engage in it. They manifest as affective inconsistency across modalities or within a modality, such as language, facial, vocal expressions, and body language. While experts can be trained to recognize A/H, integrating them into digital health interventions is costly and less effective. Automatic A/H recognition is therefore critical for the personalization and cost-effectiveness of digital health interventions. Here, we explore the application of deep learning models for A/H recognition in videos, a multi-modal task by nature. In particular, this paper covers three learning setups: supervised learning, unsupervised domain adaptation for personalization, and zero-shot inference via large language models (LLMs). Our experiments are conducted on the unique and recently published BAH video dataset for A/H recognition. Our results show limited performance, suggesting that more adapted multi-modal models are required for accurate A/H recognition. Better methods for modeling spatio-temporal and multimodal fusion are necessary to leverage conflicts within/across modalities.

cs.CV↗