Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Lattice-ordered algebras admitting a polynomial growth continuous function calculus

We characterize the Archimedean lattice-ordered algebras with identity that admit a polynomial growth continuous function calculus. More precisely, for an $n$-tuple $\mathbf{x}=(x_1,\dots,x_n)$ in an Archimedean lattice-ordered algebra $X$ with identity $1_X$, we prove that the existence of a lattice-algebra homomorphism from the algebra $PG_n$ of continuous functions on $\mathbb{R}^n$ of polynomial growth, sending the coordinate projections to $x_1,\dots,x_n$ and the constant function to $1_X$, is equivalent to the existence of $f\ge 1_X\vee |x_1|\vee \cdots \vee |x_n|$ and an $f\!$-subalgebra $Y$ of $X$ such that $1_X,x_1,\ldots ,x_n \in Y$ and, for every $m \in \mathbb{N}$, the norm $\|{\cdot }\|_{f^{m}}$ is complete on $Y\cap I_{f^{m}}$. This result may be viewed as an analogue, for lattice-ordered algebras, of the characterization of positively homogeneous continuous function calculus for Archimedean vector lattices due to Laustsen and Troitsky. As a by-product, we describe the finitely generated free objects in the category of uniformly complete Archimedean $f\!$-algebras and also show that the existence of a nontrivial polynomial growth continuous function calculus on a vector space forces it to be a commutative $f\!$-algebra.

math.FA↗

Exact baryon-meson relation for $b \to s ν\barν$

We derive an exact relation among the normalized branching fractions of $Λ_b \to Λν\barν$ and $B \to K^{(\ast)} ν\barν$, within the lepton-number-conserving effective Hamiltonian in which only active neutrinos contribute to the invisible final state. Although the effective Hamiltonian contains 18 independent Wilson coefficients, the three normalized rates depend on only two independent combinations of them, leading to an exact algebraic relation. Remarkably, the central values of the coefficients of this baryon-meson relation are close to $1/4$ and $3/4$, which appear in the heavy-to-heavy $b\to c$ semileptonic sum rule among the branching fractions of $Λ_b \to Λ_c τ\barν$ and $B \to D^{(\ast)}τ\barν$. Once the decay rate of $B \to K^{\ast} ν\barν$ is measured, the decay rate of $Λ_b \to Λν\barν$ can be determined in a model-independent manner for new-physics scenarios involving only active neutrino interactions, thereby providing a clean target prediction for future experiments. This clearly demonstrates that observables in baryonic and mesonic $b \to s ν\barν$ transitions will serve as a compact consistency test and a powerful probe for discriminating among new-physics scenarios.

hep-ph↗

Model Predictive Communication for Timely Status Updates in Low-Altitude Networks

Timely information delivery in low-altitude networks is critical for many time-sensitive applications, such as unmanned aerial vehicle (UAV) navigation, inspection, and surveillance. The key challenge lies in balancing three competing factors: stringent data freshness requirements, UAV onboard energy consumption, and interference with terrestrial services. Addressing this challenge requires not only efficient power and channel allocation strategies but also effective communication timing over the entire operation horizon. In this work, we propose a model predictive communication (MPComm) framework, enabled by advanced channel sensing techniques, in which the channel conditions that the UAV will experience are largely predictable. Within this framework, we formulate a constrained bi-objective optimization problem to achieve a desired trade-off between energy consumption and terrestrial channel occupation, subject to a strict timeliness constraint. We solve this problem using Pareto analysis and show that the original non-convex, mixed-integer problem can be decomposed into a two-layer structure: the outer layer determines the optimal communication timing, while the inner layer determines the optimal power and channel allocation for each communication interval. An efficient algorithm for the inner problem is developed using non-convex analysis, with asymptotic optimality guarantees, while the outer problem is solved optimally via a simple graph search, with edges characterized by inner solutions. The proposed approach applies to a broad class of problem variants, including objective transformations and single-objective specializations. Numerical results demonstrate the efficiency of the proposed solution, achieving up to a six-fold reduction in terrestrial channel occupation and a 6dB energy saving compared to benchmark schemes.

eess.SY↗

Control of deterministic breakdown to turbulence of hypersonic boundary layer with spanwise non-uniform surface temperature

Direct Numerical Simulation (DNS) of a Mach 6 boundary layer over a flat plate is performed to assess the effect of spanwise non-uniform surface temperature on breakdown to turbulence under deterministic forcing. The streamwise location of laminar to turbulent transition in hypersonic boundary layers has a significant influence on viscous drag and aerodynamic heating of external surfaces of hypersonic vehicles. Previous work investigated the stabilization of hypersonic boundary layers by optimally growing streaks. More recently, DNS for a hypersonic boundary layer showed that it is possible to generate streaks through a spanwise non-uniform surface temperature distribution. The laminar computations showed the control method can stabilize the second Mack mode and it is robust across a range of Mach numbers and wall temperature ratios. In this work, two scenarios are investigated where two-dimensional (second Mack mode) and oblique (first Mack mode) disturbances dominate the initial linear stage of transition. It is found that weak control streaks with amplitude below 5% of the freestream velocity can reduce high-frequency shear-stress due to the second Mack mode by approximately 30% relative to the uncontrolled configuration, and delay transition. For first Mack mode dominated breakdown, the control streaks have no effect on transition location, but the peak amplitude of the spanwise-integrated wall heat flux is reduced. For the first and second Mack mode-dominated scenarios, the mean and high-frequency peak heat transfer are reduced approximately by 15% and 34%, respectively. The dominant mechanisms are identified and attributed to the pressure work contribution to turbulent kinetic energy and the second Mack mode dilatation work.

physics.flu-dyn↗

SGP-SAM: Self-Gated Prompting for Transferring 3D Segment Anything Models to Lesion Segmentation

Large segmentation foundation models such as the Segment Anything Model (SAM) have reshaped promptable segmentation in natural images, and recent efforts have extended these models to medical images and volumetric settings. However, directly transferring a 3D SAM-style model to lesion segmentation remains challenging due to (i) weak spatial representational capacity for small, irregular targets in intermediate features, and (ii) extreme foreground-background imbalance in 3D volumes.We propose SGP-SAM, a self-gated prompting framework for efficient and effective transfer to 3D lesion segmentation. Our key component, the Self-Gated Prompting Module (SGPM), performs conditional multi-scale spatial enhancement: a lightweight multi-channel gating unit predicts whether the current features require additional multi-scale fusion, and only then activates a Multi-Scale Feature Fusion Block to enrich spatial context. To further address small-lesion learning, we design a Zoom Loss that up-weights lesion-focused supervision by combining Dice and a voxel-balanced focal term.Experiments on MSD Liver Tumor and MSD Brain Tumor (enhancing tumor) show consistent gains over strong transfer baselines based on SAM-Med3D. On MSD Liver Tumor, SGP-SAM improves mDice by 7.3% over fine-tuning.

cs.CV↗

Computational Method for Desensitized Optimal Guidance Using Direct Collocation

A computational method is developed for desensitized optimal guidance using adaptive Gaussian quadrature collocation. The method computes a reference trajectory that reduces the sensitivity to uncertainties in the dynamic model by augmenting the objective functional to explicitly penalize the sensitivity of the state with respect to uncertain parameters. Using this desensitized reference trajectory as a starting point, the desensitized optimal guidance method developed in this paper computes a new optimal control on the remaining horizon at specified guidance update times. This shrinking horizon optimal control problem is solved using a Legendre-Gauss-Radau collocation method where, at each guidance update, a reduced-horizon mesh is determined by remapping the mesh to the remaining horizon and deleting the portion of the mesh associated with the expired portion of the horizon. The resulting guidance solution is found to improve robustness to external disturbances and modeling errors. The method is demonstrated on two numerical examples. The first example is Zermelo's navigation problem which illustrates the behavior of the method on a simple example. The second example is an atmospheric reentry problem that demonstrates the performance of the method on a more complex problem. For both examples, the dynamics are simulated in the presence of parameter uncertainties in the dynamic model, and Monte Carlo analysis is performed. The results show that the method developed in this paper produces tighter trajectory envelopes and smaller terminal state errors without significantly increasing the computational burden when compared with a method that does not penalize sensitivities.

math.OC↗

Fine Oxide Dispersoids Modulate Phonon Drag and Dislocation Relaxation in Dynamically Deformed Superalloys

Plastic deformation in metallic alloys is primarily governed by the motion of dislocations, which are atomic scale line defects that move through the crystal lattice under an applied stress. Superalloys containing fine oxide particles withstand extreme temperatures for prolonged durations by resisting dislocation motion. This thermally controlled mechanism conventionally involves dislocations first climbing over the particle and slowly detaching from it, a process that typically occurs on the order of seconds. However, these mechanisms drastically change when dislocation velocities increase, potentially exceeding half the shear wave speed of the material. At such extreme speeds, dislocations can interact with lattice vibrations, leading to pronounced phonon interactions. We leverage a pulsed laser to drive rigid microspheres at controlled velocities towards superalloy substrates containing a dense oxide dispersion. Synchronized high speed imaging allows precise mapping of deformation events, allowing high throughput decoupling and modeling of plasticity contributions. We find that the oxide network produces a dual and counterintuitive effect. Our modeling framework indicates that rapidly moving dislocations bypass oxide particles by bowing rather than climbing, thereby suppressing departure side dislocation relaxation. At the same time, the dense oxide network confines fast moving dislocations within the critical interparticle distance, thereby reducing their interaction with phonons. These findings shed light on new plasticity mechanisms in oxide particle containing superalloys when line defects accelerate and dissipate energy on picosecond timescales.

cond-mat.mtrl-sci↗

Extended State-dependent Hawkes Process for Limit Order Books: Mathematical Foundation and the Reproduction of Volatility Signature Plots

This paper proposes an Extended State-Dependent Hawkes Process (ExsdHawkes) to model the intricate dynamics of Limit Order Books (LOBs). Our theoretical contribution lies in relaxing traditional constraints by allowing for state disappearances---a phenomenon frequently observed in high-frequency trading. We mathematically prove, using Karush--Kuhn--Tucker (KKT) conditions, that the maximum likelihood estimation remains separable, justifying an efficient two-step procedure. In the empirical section, we apply our model to three months of high-frequency tick data of Mitsubishi UFJ Financial Group (8306). We demonstrate that ExsdHawkes successfully replicates the characteristic upward slope of the volatility signature plot by capturing the ``local super-criticality'' triggered during disequilibrium states. Crucially, we clarify that the transition out of equilibrium is deterministically triggered by Aggressive Market Orders (AMS/AMB), while Marketable Limit Orders (MLO) function as a critical liquidity-depletion catalyst within the expanded spread. Comparative analysis reveals that models lacking physical constraints (e.g., standard SD-Hawkes) suffer from explosive spectral radii and fail to maintain simulation stability. Our findings suggest that physical consistency is not merely a mathematical nicety, but a prerequisite for accurately modeling macro-level volatility. By enforcing the physical geometry to `pause' the residual accumulation during inadmissible periods, ExsdHawkes maintains statistical integrity where unconstrained models succumb to structural bias and simulation instability.

stat.AP↗

SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs

Mixture-of-Experts (MoE) has become a prevalent backbone for large vision-language models (VLMs), yet how modality-specific signals should guide expert routing remains under-explored. Existing routing strategies are either hand-crafted or modality-agnostic, relying on idealized priors that ignore the layer-dependent modality fusion patterns in MoE-VLMs and provide little guidance for expert specialization. We propose Soft Modality-guided Expert Specialization (SMoES), which consists of dynamic soft modality scores that capture layer-dependent fusion patterns, an expert binning mechanism aligned with expert-parallel deployment, and an inter-bin mutual information regularization that encourages coherent modality specialization. Our method leverages attention-based or Gaussian-statistics modality scores to optimize mutual information regularization. Experiments across four MoE-based VLMs and 16 benchmarks demonstrate improvement on both effectiveness and efficiency: 0.9% and 4.2% average gain on multimodal and language-only tasks, 56.1% reduction in EP communication overhead, and 12.3% throughput improvement under realistic deployment. These results validate that aligning routing with modality-aware expert specialization unlocks MoE-VLM capacity and efficiency.

cs.CV↗

On the nonexistence of good involutions on symplectic quandles

We investigate a necessary and sufficient condition for the existence of good involutions on symplectic quandles, which are defined on free $R$-modules equipped with an antisymmetric bilinear form. In particular, we discuss the nonexistence of such good involutions under various algebraic conditions.

math.GT↗

Clustering in co-evolving opinion dynamics: reduced SPDE models

Clustering is a fundamental collective phenomenon in agent-based models (ABMs) of opinion dynamics. To study clustering in systems with co-evolving social and opinion variables, we derive stochastic partial differential equation (SPDE) models that describe the evolution of clusters. Specifically, these SPDE models are reduced in the sense that their domain accounts only for social variables, while opinion variables are instead incorporated in a multiplicative (Eulerian velocity-type) fashion. The model reduction is primarily guided by computational efficiency, as well as model interpretability. We consider two settings: one in which opinions do not affect social interactions, and another one in which a feedback mechanism couples the two. Our approach extends reduced PDE modelling to a stochastic framework, which is essential for capturing long-term cluster behaviour. Numerical experiments indeed demonstrate that the proposed reduced SPDEs substantially decrease computational cost compared to full-state SPDE models, such as the Dean--Kawasaki equation, while still accurately reproducing the clustering behaviour of the underlying ABM. As a result, these reduced models provide an efficient tool for studying systems with large populations, including those arising in the analysis of real-world data: in particular, we provide an application related to the large-scale General Social Survey (GSS), which comprises opinion and social data of the US population since 1972.

physics.soc-ph↗

Query-Dependent Use of Generated Descriptions for Reliable Visual Question Answering

Vision-Language Models (VLMs) hallucinate objects that are not present, and a growing line of work tries to curb this by feeding the model its own generated caption as auxiliary evidence -- assuming that a caption, once available, is something to consume. We show this fails: naively appending a caption can lower accuracy rather than raise it, dropping Qwen2.5-VL-3B† on HallusionBench by nearly ten points. To understand why, we build GD-Probe, a diagnostic set that pairs a global and a detail question on the same image, so that any difference in caption effect is attributable to the question alone. Caption utility proves to be a per-query property: the same caption helps global questions and harms detail ones, through a single mechanism -- an embedded caption competes with the image for attention and pulls the model's evidence onto its own text -- whose sign is set by whether the caption covers the queried content. Crucially, this regime is readable from quantities the decoder already emits, with no attention access or grounding. We turn this into GEASS (Gated Evidence-Adaptive Selective Caption Trust), a training-free, logit-level module that decides per query how much of the caption to trust, gating it by the clean path's confidence, weighting it by the entropy reduction it induces, and raising the evidence bar when the two pathways disagree. Across four VLMs and two benchmarks (POPE and HallusionBench), GEASS improves over both vanilla inference and contrastive decoding under a single fixed setting, adding only two forward passes and no parameters.

cs.CV↗

Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time

Multimodal large language models (MLLMs) achieve strong performance on vision- and audio-language tasks, yet can generate responses that conflict with the given visual or auditory inputs, a problem known as multimodal hallucinations. Prior work suggests that this occurs when models rely more on textual cues and learned language patterns than on evidence from the perceptual input. To obtain a more direct account of this imbalance, we apply Layer-wise Relevance Propagation (LRP), which attributes predictions to individual input tokens, and use the resulting relevance scores to analyze and mitigate hallucinations. First, we examine whether this imbalance leads to multimodal hallucinations. We find that hallucinations often arise when the model relies less on perceptual inputs, and that changing this reliance affects its predictions. We further leverage LRP and propose a training-free framework that shifts relevance toward perceptual tokens by optimizing key-value representations during decoding, without modifying model parameters or requiring training data. We call this method Learning Inference-time Modality Enhancement (LIME). Despite using no spatial or temporal supervision, LIME concentrates relevance on query-relevant regions. We evaluate LIME across multiple multimodal benchmarks in both vision and audio domains, demonstrating consistent reductions in hallucinations and enhanced grounding while preserving generation quality.

cs.LG↗

Temporal State Tomography via Quantum Snapshotting the Temporal Quasiprobabilities

Quantum tomography is a cornerstone of quantum information science, enabling the reconstruction of states and channels from experimental data. Here we introduce a new paradigm, temporal state tomography (TST), for reconstructing quantum processes across multiple times. Our approach is based on temporal quasiprobability distributions (TQDs), which, in the informationally complete setting, provide a complete description of multi-time quantum processes and uniquely determine temporal states. We formulate TST as a unified framework for reconstructing both density operators and quantum channels within a single scheme. We show that any TQD can be obtained via classical post-processing of measurement outcomes generated by a fixed set of quantum instruments, thereby establishing a direct operational route to accessing TQDs experimentally. For informationally complete TQDs, the associated temporal state can be reconstructed via a temporal Bloch-type representation. Leveraging this correspondence, we derive the sample complexity of TST, thereby quantifying its statistical efficiency.

quant-ph↗

Transferability of Error Bounds and the Kurdyka-Łojasiewicz Property under $C^1$ Partial Smoothness

This work studies how a generalized error bound (GEB) property in the ambient Euclidean space transfers to the active manifold under $C^1$ partial smoothness. The GEB generalizes the standard error bound, which upper bounds the distance to a subset of critical points using first-order stationarity residuals, by allowing the distance to be composed with a gauge function. In the ambient space, these residuals are measured by the subdifferential or the slope of the original function, while on the manifold they are measured by the Riemannian gradient or the slope of the restricted function. We develop complementary geometric and metric analyses to show that the ambient and manifold GEBs are equivalent with the same gauge function. The geometric argument shows that the projection of the subdifferential onto the tangent space coincides with the Riemannian gradient, while the metric approach combines slope estimates, identifiability, and a new uniform linear-growth result for $C^1$ manifolds. We further extend our analysis to establish the analogous equivalence for the Kurdyka--Łojasiewicz (KL) property, with the same desingularizer, and separately characterize their connections with the GEB. Finally, for regularized optimization, we relate subdifferential and proximal error bounds, thereby placing the previously established equivalence between proximal and manifold error bounds for $\ell_1$-regularization within a broader framework.

math.OC↗

FoR-Net: Focus-on-Regions Network for Semantic Segmentation

This paper presents Focus-on-Regions Network (FoR-Net), an efficient semantic segmentation framework that explicitly focuses on hard regions through a selector-driven Top-K mechanism. Instead of relying on heavy global modeling, FoR-Net selectively enhances structurally informative regions. Multi-scale reasoning branches are introduced to aggregate spatial context efficiently. Experiments on the Cityscapes dataset demonstrate that FoR-Net achieves competitive performance while maintaining a lightweight architecture.

cs.CV↗

Diffusion Masked Pretraining for Dynamic Point Cloud

Dynamic point cloud pretraining is still dominated by masked reconstruction objectives. However, these objectives inherit two key limitations. Existing methods inject ground-truth tube centers as decoder positional embeddings, causing spatio-temporal positional leakage. Moreover, they supervise inter-frame motion with deterministic proxy targets that systematically discard distributional structure by collapsing multimodal trajectory uncertainty into conditional means. To address these limitations, we propose Diffusion Masked Pretraining (DiMP), a unified self-supervised framework for dynamic point clouds. DiMP introduces diffusion modeling into both positional inference and motion learning. It first applies forward diffusion noise only to masked tube centers, then predicts clean centers from visible spatio-temporal context. This removes positional leakage while preserving visible coordinates as clean temporal anchors. DiMP also reformulates point-wise inter-frame displacement supervision as a DDPM noise-prediction objective conditioned on decoded representations. This design drives the encoder to target the full conditional distribution of plausible motions under a variational surrogate, rather than collapsing to a single deterministic estimate. Extensive experiments demonstrate that DiMP consistently improves downstream accuracy over the backbone alone, with absolute gains of 11.21% on offline action segmentation and 13.65% under causally constrained online inference.Codes are available at https://github.com/InitalZ/DiMP.git.

cs.CV↗

Deco: Extending Cherished Physical Objects into AI Companion Agents through Dual Embodiment

Physical objects (e.g., plush toys) can transcend materiality to become emotional anchors and provide companionship. However, these bonds remain one-sided because most physical objects cannot reciprocate. AI companions offer responsiveness and personalization, but typically entail building bonds from scratch. We investigated how AI companions might inherit and extend users' existing bonds with physical objects. A formative study (N=9) informed four design principles (Faithful Identity, Calibrated Agency, Ambient Presence, Reciprocal Memory), shaping our Dual-Embodiment Companion Framework. We instantiated it as Deco to create digital embodiments of physical companions. In a within-subjects lab study (N=25), Deco was rated higher than a personalized digital-only companion on six companion-related measures (all p<.01). A subsequent seven-day field deployment (N=17) showed sustained engagement, higher post-deployment well-being (p=.040), and three key relational patterns: digital activities retroactively vitalized physical objects, bond deepening centered on emotional engagement depth, and participants sustained bonds while navigating companions' AI nature. Dual embodiment offers a promising framework for revitalizing physical objects with AI agents.

cs.HC↗