Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

ADF-EA: A Unified Execution Assurance System for Agent Device Foundation

Agents based on large language models (LLMs) can access heterogeneous devices through tools and APIs, but reliable execution must account for unmet effects, uncertain outcomes, and changing prerequisites. A command may be acknowledged without producing its intended effect, while missing feedback may obscure an action that has already succeeded. We present Agent Device Foundation--Execution Assurance (ADF-EA), an architecture that connects agent planning and device execution through shared capability contracts. Device Capability Contracts (DCCs) unify invocation conditions, intended effects, evidence requirements, and recovery rules across heterogeneous interfaces. Agents use these contracts to plan, while the runtime applies the same semantics to authorize actions, verify effects, and govern continuation and completion. Persistent execution state retains verified progress, unresolved outcomes, and remaining budgets across plan revisions, enabling observation-based recovery, authorized retries, and necessary state repair. We formalize the execution lifecycle and establish conditional soundness properties for completion and recovery authorization. Evaluations span multiple LLMs, five agent frameworks, and simulated process-control, household, and robotic manipulation domains. Compared with direct invocation and existing execution-checking approaches, ADF-EA reduces false completion and unnecessary repetition, supports necessary state repair, prevents calls to unavailable capabilities, and preserves permitted task completion and recovery. These results demonstrate DCCs as a reusable semantic foundation for agent autonomy across heterogeneous devices, unifying capability-based planning, evidence-grounded execution, and authorized recovery within one architecture.

cs.MA↗

LUMO (Lightweight Unified Multilingual Orchestrator): A Privacy Preserving Offline Voice Assistant

Reliable voice interaction is essential in environments with limited internet connectivity and strong privacy. However, most existing voice assistants depend on cloud-based services, which leads to latency issues, dependency on internet access, and privacy vulnerabilities. This research presents LUMO (Lightweight Unified Multilingual Orchestrator), a privacy preserving offline voice assistant designed for edge computing environments. This system integrates local Automatic Speech Recognition (ASR), locally deployed quantized Large Language Model (LLM), and Text-to-Speech (TTS) synthesis into a fully offline pipeline running on a Raspberry Pi 5 with 8 GB RAM. To enable efficient operation on resource constrained hardware, the language model is compressed using 4-bit GGUF quantization, which reduces memory usage while preserving practical conversational capability. Existing edge based voice assistants Mycroft provides partial offline functionality without a generative LLM, with an approximate latency of ~5 s and power consumption of ~12 W, while Rhasspy supports full offline operation but lacks generative capabilities, with ~3 s latency and ~11 W power usage. In contrast, LUMO achieves a Word Error Rate (WER) of 6.8% for short English utterances in low noise conditions, an end-to-end response latency of 2.0-4.0 s, and a lower peak power consumption of approximately 9.0 W. The system also achieves effective offline recognition for Bangla speech, supporting multilingual accessibility in low resource settings. By operating entirely offline, LUMO provides strong data privacy, reduced need for cloud connectivity, and suitability for privacy sensitive edge execution such as rural healthcare, education, and disaster response scenarios.

cs.LG↗

Convergence rates in the periodic homogenization of vanishing-viscous Hamilton-Jacobi equations

We study convergence rates when periodic homogenization and vanishing viscosity occur at the same scale in Hamilton--Jacobi equations whose momentum Hessians are positive definite at every point. For bounded Lipschitz initial data and a class of smooth Hamiltonians, we establish an optimal rate $O(\varepsilon|\log\varepsilon|)$ on fixed time intervals. The second-order term plays a key role: the elliptic cell problem provides higher derivative bounds for the effective Hamiltonian $\bar{H}$ and its Legendre dual $\bar{L}$. The proof is purely PDE, based on corrector expansions and viscosity comparison, while its guiding idea comes from the control interpretation of the Hopf--Lax formula. For a fixed target $(x,t)$, a minimizing origin selects a characteristic velocity and its dual momentum, these guide the construction of smooth comparison profiles, whose gradients supply the momentum argument of the first-order corrector. Thus we never need to differentiate the effective solution, even at points where it is nonsmooth.

math.AP↗

Training-Free Contextual ASR via SpeechLLM-Based Error-Aware Selective Retrieval

Recognition of domain-specific and low-frequency terms remains challenging for automatic speech recognition (ASR). Although contextual biasing can improve their recognition, directly providing a large terminology dictionary introduces many irrelevant biasing terms. Retrieval-based contextual biasing addresses this issue by selecting candidate terms from an external dictionary, but querying many recognized words requires numerous dictionary lookups and may yield poorly targeted candidates. We propose a training-free contextual ASR framework in which a pretrained speech large language model (SpeechLLM) jointly generates an ASR hypothesis and localizes error spans likely to involve domain-specific terms. Only the localized spans are used to retrieve phonologically similar terms from an external terminology dictionary. The same SpeechLLM then re-recognizes the audio conditioned on the first-pass hypothesis and the retrieved terms, without task-specific model training. To assess applicability across domains, we evaluate the framework on medical, air traffic control, and financial speech. The proposed method substantially reduces dictionary queries while improving the recall and ranking of relevant terminology candidates and second-pass ASR performance across all three domains.

cs.SD↗

Multi-Objective Human-in-the-Loop Bayesian Optimization of a Lower-Limb Exoskeleton

Human-in-the-loop optimization (HILO) is a common approach for optimizing the control of assistive devices to account for the wearer's unique biomechanics and subjective preferences. However, despite research suggesting that a person may have a different prioritization of objectives depending on time-varying factors such as the environment, their mood, or energy levels, existing HILO approaches only consider a single objective or enforce a fixed weighting on a set of objectives. Neither approach is capable of representing an individual's preferences over objectives. In this work, we propose Multi-Objective Human-in-the-loop Bayesian Optimization (MO-HILBO), which builds on explicit multi-objective Bayesian optimization to efficiently infer a personalized set of Pareto-optimal controllers. We compare our approach with an existing multi-objective HILO method and experimentally demonstrate MO-HILBO on a lower-limb exoskeleton across two objectives: metabolic cost (efficiency) and ordinal human feedback (comfort). We find that MO-HILBO (1) discovers Pareto-optimal controllers, and (2) that the pairwise ordering of points on the Pareto front itself is consistent with validation trials. Lastly, we open-source mohilo, a Python package for running both HILO and MO-HILBO on wearable devices: https://dynamicmobility.github.io/mohilo/.

cs.RO↗

Learning Vision-Based Agile Gap Traversal: Differentiable Simulation with a Warm-Started Critic

Traversing narrow gaps is challenging for autonomous quadrotors, especially when control commands come directly from high-dimensional visual observations. Existing end-to-end methods often rely on behavior cloning or full-rollout backpropagation through time (BPTT) via differentiable simulation, which can limit policy performance or incur high training costs. We propose a two-stage reinforcement learning framework for more efficient ego-centric visuomotor gap-traversal policy training, leveraging quasi-analytical policy gradients (QPG) via differentiable simulation and critic warm-starting. The framework utilizes QPG to avoid backpropagation through visual rendering, reducing computation and memory costs while improving sample efficiency. In the first stage, an expert actor and critic are trained using privileged observations, including gap geometry. Unlike prior gap-traversal approaches, our training utilizing QPG does not require resetting the agent along optimized reference trajectories. In the second stage, a visual policy is trained using binary gap masks from two ego-centric cameras and low-dimensional observations, while its privileged critic is warm-started from the first stage. This substantially improves training efficiency and traversal success compared with cold-starting the critic or using full-rollout BPTT. Our framework does not require retraining the expert actor when system parameters change, enabling more efficient generalization across drone platforms than state-of-the-art visual gap-traversal methods based on action supervision. The learned visual policy also generalizes to gaps with unseen shapes. Extensive real-world experiments further demonstrate robust gap traversal using binary masks rendered online. Beyond gap traversal, the proposed framework is generic and can be extended to other visuomotor robot learning tasks.

cs.RO↗

The second gap for self-shrinkers with constant norm of the second fundamental form

Let $X: M^{n}\to \mathbb{R}^{n+1}$ be a complete self-shrinker with constant squared norm of the second fundamental form $S$. In this paper, we prove that if $S\leq \frac{10}{7}$, then $S=1$ or $S=0$, and the self-shrinker is isometric to either a round sphere $\mathbb{S}^n(\sqrt{n})$ with the center at the origin, or a cylinder $\mathbb{S}^k(\sqrt{k})\times \mathbb{R}^{n-k},~~1\leq k\leq n-1$, or a plane through the origin.

math.DG↗

MM-VeriAgent: Learning to Use Extensive Tools to Verify Multimodal Misinformation with Reinforcement Learning

Real-world multimodal misinformation often involves mixed forgery sources, requiring sample-specific detection strategies. Existing tool-augmented methods rely on predefined workflows or inference-time planning, limiting adaptability or increasing inference cost. To address this issue, we introduce \textbf{MM-VeriAgent}, which learns to verify mixed-source multimodal misinformation with tools. We first build \textbf{MM-VeriTools}, a specialized toolkit for misinformation detection agents. By benchmarking various candidate models and methods on the sub-tasks required by mixed-source detection, we select the strongest for textual, visual, and cross-modal forgery analysis and encapsulate them as callable tools with a unified interface. On top of this toolkit, we train the LVLM agent with reinforcement learning to teach it how to use these tools to better solve mixed-source detection. Since many of the tools are specialized models whose online execution at every rollout severely limits RL efficiency, we further introduce \textbf{Tool-Execution Cache}, which pre-executes candidate tool calls and reuses their cached outputs during training. This preserves multi-step rollouts while reducing online tool execution, largely improving the training efficiency.Experiments on MMFakeBench demonstrate substantial accuracy gains over the base model without explicit tool search at inference time. Ablation and efficiency analyses further validate the learned tool-use policy and show that Tool-Execution Cache reduces online tool executions during training.

cs.CV↗

"If You're Not Doing It, Somebody Else Is": Active Negotiation and the Invisible Labor of Sustained LLM Use

Large language models (LLMs) have become fixtures of academic work even as their users describe them as degrading their writing, thinking, and skills. Dominant adoption frameworks read continued use as evidence of satisfaction, and cannot explain continued use of a distrusted tool. We interviewed 36 graduate student workers, balanced between English-as-a-foreign-language (EFL) and non-EFL speakers, and introduce the Active Negotiation framework: a model of sustained LLM use as a recurring cycle of risk, mitigation, and justification. A failure surfaces a risk, mitigation labor addresses it, and a justification renders the residual risk tolerable until the next failure reopens the cycle. The cycle runs across three dimensions: practical, auditing output; internal, auditing one's own cognition and identity; and social, managing how peers and institutions perceive use. EFL participants invoke linguistic parity as a further justification. We reframe continued adoption as compliance sustained by invisible labor.

cs.HC↗

Coarse geometry of metric measure spaces

Using ideas from optimal transport theory, we introduce a notion of measured coarse equivalence for metric measure spaces and define a corresponding variant of the uniformly finite homology of Block and Weinberger, called weighted $\ell^\infty$ homology, for large-scale doubling metric measure spaces. We prove that this homology is invariant under measured coarse equivalence and that the vanishing of its zeroth homology is equivalent to weighted non-amenability. The proofs combine techniques from optimal transport theory and the disintegration of measures.

math.MG↗

Propagation delays and regional intensity changes in lensed hotspot images

Propagation delays cause an image recorded at a single observer time to combine radiation emitted at different stages of the source evolution. Comparable changes in total intensity can accompany distinct and even opposite changes in apparent image size. Using regional intensities, centroids, and covariances, we apply the law of total covariance to separate changes in regional intensity weights from changes in internal widths and centroid separation. Ray tracing simulations of a finite Gaussian hotspot moving along a prescribed strong field trajectory show two events with comparable attenuation in screen integrated intensity but opposite changes in second moment size. In the contracting event, the regional weights move away from balance; in the expanding event, they move toward balance even as the centroids approach each other. Differential propagation delays therefore drive these opposite size responses by redistributing intensity between spatially separated image regions.

gr-qc↗

From Bilateral Trade to Matching Markets: Sharp Gains from Trade

We study gains from trade in matching markets with independent private values and costs, Bayesian incentive compatibility, interim individual rationality, and no expected budget deficit. A second-best guarantee for finite bilateral trade extends without loss to matching markets with independent Borel priors, arbitrary downward-closed feasibility, and finite expected first-best gains. For bounded buyers with monotone hazard rates and arbitrary bounded sellers, we determine the exact worst-case ratio of second-best to first-best gains, approximately $0.72490721$. For binary buyers and sellers with at most $m$ types, we determine the exact ratio for every $m$, including $8/9$ when $m=2$ and a limit of $4/5$ as $m$ grows. Both families of bounds are tight already in bilateral trade.

cs.GT↗

SAGE: Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion

Spatial misregistration and cross-modal discrepancies often cause ghosting, structural blurring, and content imbalance in RGB-T fusion. Existing methods typically decouple appearance adaptation, geometric alignment, and information fusion, limiting dependency propagation across stages. We propose Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion (SAGE), a unified framework integrating frequency equalization, hierarchical alignment, and subband fusion. SAGE employs invertible joint encoding and source-specific low-frequency modulation to derive structural and gain guidance while preserving source information. Hierarchical frequency collaborative alignment estimates global affine geometry from low-frequency approximations and transfers geometric and contextual cues to high-frequency correlation reasoning for reliability-aware residual refinement. Guided subband fusion jointly aggregates the aligned frequency coefficients under propagated source and alignment guidance, coordinates complementary low- and high-frequency information, and reconstructs the fused image through the inverse wavelet transform. Extensive experiments on RGB-T datasets with real-world and synthetic misalignments demonstrate consistently competitive performance in alignment and fusion, validating the effectiveness of source-anchored guidance for weakly registered RGB-T images.

cs.CV↗

From Visual Search to Movement Control: A Priority Field for Artificial Agents

Human spatial attention is widely conceptualized as being guided by a priority map that integrates perceptual salience, current goals, and past experiences. Here, we extend priority-based computation to movement control in artificial agents. We first introduce a lightweight model of visual search based on an integrated priority map. Trained on human saccades, it reproduced key behavioral patterns, including oculomotor suppression and history-driven selection. Extending the search model, we equipped an artificial agent with a priority field and evaluated its performance in a reach-avoid task that required reaching a goal destination while avoiding moving obstacles. Compared with alternative architectures, priority-field agents trained more efficiently and performed better in unseen, complex scenarios, even from simple demonstrations. Adding a simple memory mechanism also produced human-like, history-driven effects in anticipating the likely location of the upcoming goal. These findings suggest that priority-based computation may provide a promising foundation for movement control in artificial agents.

cs.RO↗

The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?

While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each formation date, prompts, output formats, and portfolio construction fixed. Our evaluation covers a full year of U.S. equities under three input conditions: numerical, identifiable news, and masked news. It includes more than 800,000 asset predictions and repeated model generations. Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns. For DeepSeek, where we examine the full progression from no reasoning to maximum reasoning, performance is nonmonotonic. Repeated generations also produce unstable treatment effects and portfolio selections, even when overall scores remain similar. These findings show that additional reasoning can change financial decisions without reliably improving their economic value, motivating validation for each task before deployment.

cs.AI↗

LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information

"System One" decision models such as TypeSafe's Jev and its open counterpart Laya answer typed questions about a text in a single forward pass with calibrated probabilities, but they cannot ask for missing information: when a first message does not say what separates two departments, they guess. We present LAVOIR (Laya with Value-Of-Information Routing), which places the candidate pieces of missing information (slots) in the input next to the answer options, so that one forward pass returns both the decision distribution and, for every slot, the expected gain in the probability of the correct decision if the user were asked about it. VOI targets need no human labels: gold decisions come from schema rules, an LLM only verbalizes messages and answers, a model from another family checks every text, and pairing each message with several profiles makes regression on realized gains estimate the expected gain. A Gini-impurity cap bounds the predicted value by what a calibrated model can still gain. In a controlled study, decisions on seen schemas are statistically indistinguishable from the Bayes ceiling. The final model's question policy matches a greedy oracle VOI policy on seen schemas (AUC 0.799 vs. 0.797), and with at most 0.5 questions per conversation it is 14.1 points more accurate than never asking. On real ABCD conversations, one real exchange raises accuracy by 8.3 points where LAVOIR asks and leaves it unchanged where it does not; on SGD the cap lowers the asking rate from 93% to 8.6%. On Laya's twelve benchmarks LAVOIR is above Laya's reported scores on seven, and it answers a question in 31 ms (median, GH200).

cs.AI↗

A Monolithic Discontinuous Galerkin Framework for Darcy Optimal Control with Radon-Measure Tracking and Pointwise Control Constraints

We study an elliptic optimal control problem governed by Darcy's equation in heterogeneous porous media, with pointwise box constraints on the control. The objective functional is formulated via a Radon measure, which allows the desired pressure state to be tracked on observation sets of varying dimension, including points, curves, and subdomains, within a single formulation. The state and adjoint equations are discretized by a symmetric interior penalty discontinuous Galerkin method, yielding a locally mass-conservative approximation that is robust across strong permeability discontinuities, while the control is approximated by piecewise constants. The state, adjoint, and control are retained as primary unknowns in a single coupled optimality system and solved monolithically by a primal-dual active set strategy. We establish stability and well-posedness of the discretization and derive a priori $L^2$ error estimates for both variables. The principal difficulty is the reduced regularity of the adjoint state induced by the measure-valued tracking data. The analysis controls the resulting adjoint-control coupling through an intermediate adjoint driven by the continuous optimal state. Numerical experiments confirm the predicted convergence rates for point, curve, and subdomain observations, exhibit mesh-independent primal-dual active set iteration counts, and demonstrate local mass conservation.

math.NA↗

Combining General and Domain-Specific Pretext Tasks for Brain MR Image Segmentation

A key challenge in medical image analysis is the scarcity of large annotated datasets for specific populations and diseases. As deep learning models rely heavily on labeled data, effective transfer learning strategies are needed to reduce the dependence on manual annotations. Self-supervised learning has emerged as a promising approach for developing foundation models by enabling the learning of transferable feature representations from large-scale unlabeled medical imaging datasets. In this study, we investigate voxel-level brain age prediction as a domain-specific self-supervised pretext task and compare it with image inpainting, a widely used non-domain-specific alternative. We further propose a multitask self-supervised pretraining framework that jointly optimizes both objectives to learn complementary neuroimaging representations. The pretrained models are evaluated on three downstream magnetic resonance image segmentation tasks: multiple sclerosis lesion segmentation, ischemic stroke lesion segmentation, and cortical brain structure segmentation. Overall, the proposed multitask pretraining framework consistently outperformed the single-task pretrained models and training from scratch across most experimental settings, demonstrating the benefit of combining domain-specific and general self-supervised learning pretext tasks for the development of generalizable neuroimaging foundation models.\ Code Availability: The source code used in this study is publicly available at https://github.com/TasneemN/Combining-General-and-Domain-Specific-Pretext-Tasks-for-Brain-MR-Image-Segmentation/

cs.CV↗