Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation

Following retrieved evidence does not guarantee factual correctness: misleading evidence can induce a model to replace an answer it previously gave correctly. Standard accuracy measures obscure this behavior by combining answer replacement with preexisting errors. We introduce RAG-Stress, a controlled diagnostic protocol for examining the limits of evidence reliance in retrieval-augmented generation. The protocol holds the question and reference answer fixed, edits one assertion to support a designated incorrect answer, and crosses two source priority policies with three positions of the answer span within the evidence text. We measure misleading rate (MR) on each model's subset of questions answered correctly without retrieval, alongside clean accuracy on the full evaluation set. We evaluate fifteen systems spanning API models, open models, and search agents trained with reinforcement learning on TriviaQA-RC, HotpotQA, and SearchQA, with additional English and Chinese MedQA evaluations. Instructions that prioritize documents consistently produce higher MR than those permitting reliance on prior knowledge. Averaged over models and positions, the gap ranges from 10.9 to 13.5 percentage points across the three QA datasets. Mean MR follows End $>$ Beginning $>$ Middle under both policies, although individual models do not uniformly follow this ordering. A separate paired audit of 500 questions and two checkpoints supports increased harmful override without establishing a corresponding improvement in beneficial correction. These findings distinguish evidence adherence from factual reliability and motivate evaluating whether retrieved evidence preserves, replaces, or corrects a model's answers.

cs.CL↗

WorldFact-Bench: Beyond Image-Internal Plausibility to Image-World Consistency

Advances in image generation have made visual authenticity increasingly difficult to assess. Although image forensics now examines both generation artifacts and higher-level visual inconsistencies, a plausible image can still contradict real-world facts or rules. We introduce WorldFact-Bench to evaluate image-world consistency from a single image, without a predefined claim or verification target. The benchmark contains 1,274 source-aligned real-fake pairs across four verification regimes and ten semantic domains. Each pair introduces a specific, evidence-supported factual conflict while seeking to preserve non-target content and visual plausibility. Images are evaluated independently, and pair accuracy requires both members of a pair to be classified correctly. We further propose PERSIST-Agent, which organizes iterative verification around a persistent state linking candidate facts, visual observations, evidence, and verification statuses. This state guides subsequent inspection and retrieval while retaining unresolved candidates. With backbone weights fixed, harness self-optimization refines the agent's prompts and execution rules through validation feedback. Experiments reveal strong label biases in several detectors and uneven gains from retrieval. On the evaluated 8B backbones, PERSIST-Agent improves pair accuracy over both direct judgment and retrieval-augmented baselines, while ablations support the role of persistent verification state. These findings highlight the value of state-guided verification and the remaining gap between visual plausibility and factual correctness.

cs.CV↗

Predictive Multiplicity in Cell-Fate Assignment: Label-Free Rashomon Sets and the Limits of Per-Cell Certification

Single-cell trajectory inference maps transcriptomic measurements onto developmental continua, yet configurations that fit the data equally well can assign conflicting cell fates. FateMultiplicity is a label-free framework that constructs a statistically admissible model set, or Rashomon set, without lineage labels, by evaluating model discrepancy on cross-fitted held-out genes under non-inferiority testing calibrated against random-seed variation. Multiplicity is large and depends more on the diversity of the model space than its size: twelve configurations of a second algorithm expose 20.0% of cells where twenty-four of the first expose 3.8%. Whether the per-cell certified fate margin FM yields more reliable assignments than the fitted model already provides is then tested, and it does not. On simulation ground truth, on the same cells, FM discriminates misassignment at AUC 0.682, against 0.965 for the baseline configuration's own decision margin (p = 0.003) and 0.854 for a seed-dispersion baseline. Informativeness is governed by the breadth of the admitted set, not its cardinality: at cardinality four, seed refits give 0.933 and hyperparameter-perturbed sets 0.701. Relaxing the infimum to a q-quantile recovers discrimination but converges toward the single model's own confidence; the supremum reaches 0.973 because theta*'s membership bounds it from below, while the infimum is unanchored. Multiplicity in trajectory inference is worth measuring and reporting, but per-cell certification over a label-free Rashomon set is not a route to more reliable fate calls. Two constructions survive: a margin-erosion ratio separates real from spurious branch points in simulation (AUC 0.890, untested on real data), and against clonally observed fate, uncertified cells disagree with their clone's outcome 16.4 percentage points more often than certified cells (p < 0.001).

cs.LG↗

Improved E-Value Thresholds with Applications to Multiple and Sequential Testing

E-values provide a flexible framework for statistical inference, but the universal threshold $1/α$ can be conservative when the null distribution has additional structure. We study how structural restrictions limit the worst-case concentration underlying Markov's inequality, using a nondecreasing density model to constrain tail allocation and an $L$-Lipschitz density model to control local concentration. The nondecreasing model gives the minimax rejection threshold, and the Lipschitz condition yields a sharper closed-form threshold with an $O(L^{-1/2})$ relative improvement over $1/α$. We further show that this threshold is asymptotically sharp as $α\to0$, in the sense that no uniformly valid threshold can improve on it by a fixed positive $L$-dependent amount. We then incorporate the Lipschitz calibration into multiple testing procedures with FDR control and use the same structural information to construct calibrated conditional e-values for anytime-valid sequential inference. Numerical experiments illustrate the resulting gains in discoveries and the trade-offs of the sequential procedures.

stat.ME↗

RGBD-to-3D Object Mesh Refinement via Depth Matching and Symmetry Propagation

Single-view 3D reconstructors often produce plausible meshes that disagree with the input view, especially near depth discontinuities and self-occlusions. We present a lightweight, plug-and-play RGBD-to-3D refinement that improves any RGB-to-3D reconstructor without retraining. Given a depth map, we correct the visible surface by bipartite matching to back-projected depth points, mirror these corrections onto the occluded side across a detected symmetry plane, and propagate them with a smoothness solver. Every stage is closed-form, making the method orders of magnitude faster than optimization-heavy test-time refinement. On GSO and OmniObject3D with five backbones, it yields consistent gains, also with monocular pseudo-depth, benefits more from symmetry on symmetric objects, and compares favorably with prior refinement in accuracy and runtime. It further improves an RGB-D-to-mesh reconstructor and transfers to real captures with noisy sensor depth.

cs.CV↗

What to Admit and How to Present: Governing Persistent Memory in LLM Agents

Persistent memory can improve personalization in LLM agents but can also induce sycophancy and cross-domain leakage. We distinguish two governance decisions: admission, which determines what recalled information enters the working context, and presentation, which determines how admitted information is expressed. We implement two inference-time designs without retraining: factor-compiled admission (FC), which assesses whole memory entries, and permission-semantic admission (PS), which decomposes entries into typed units; both translate adjudicated attributes into eligibility decisions via deterministic policies. We evaluate on a four-backbone development suite and an external benchmark with four tasks of 300 samples each. Relative to verbatim injection, FC and PS reduce pooled judge-assessed failure rates on the external benchmark by 6.7 and 8.8 percentage points (p = 2.7e-7 and 4.1e-12), and development-set cross-domain leakage falls by up to 29.5 percentage points. A query-conditioned gating baseline shows no significant change in objective-fact failure or pooled failure. Under matched admission budgets, PS outperforms random and relevance-based selection on external objective-fact judgment after Holm correction. Holding presentation fixed, tightening admission cuts cross-domain failure by a further 17.5 percentage points (p = 1.6e-4); in contrast, no comparison between two renderings of identical adjudicated outputs survives multiple-comparison correction. Both designs increase personalization failures, and PS misses the preregistered improvement and personalization-preservation criteria. These results support evaluating admission and presentation separately: selection quality provides task-specific safety gains, while preserving beneficial memory use remains unresolved.

cs.AI↗

Classical cylindrical contact homology is not invariant over $\bZ$

We show that classical cylindrical contact homology over $\mathbb{Z}$, equipped with its decomposition by free homotopy classes, is not invariant under changes of hypertight contact form. Specifically, we construct two nondegenerate hypertight perturbations of the standard Seifert contact form on $Σ(2,4,5)$ whose cylindrical contact homologies in the regular-fiber class are respectively torsion-free and contain $\mathbb{Z}/5$-torsion.

math.SG↗

Integrity and Credibility in Navigation: From Error Characterization to Operational Assuran

Navigation integrity and estimator credibility use error and uncertainty information to address different questions. This perspective distinguishes actual positioning error from reported statistical characterizations, formulates integrity as control of hazardous unwarned-use risk under the requirement of a specified operation, and describes credibility as the justification of an uncertainty or risk claim by its evidence usage and modeling process. Three hypothetical scenarios show how deteriorated accuracy, a contradicted uncertainty report, and a hazardous event can lead to different credibility and integrity judgments. The resulting assessment asks what happened, why the reported claim should be believed, and whether the information is sufficient for the intended operation.

eess.SP↗

Cooperative STT and SOT switching in perpendicular magnetic tunnel junctions: Role of the pulse-end magnetization state

Spin--orbit torque MRAM (SOT-MRAM) is a leading candidate for next-generation nonvolatile memory, offering high speed, endurance, and architectural compatibility. However, in conventional SOT switching, the magnetization remains near the in-plane region at pulse termination, making the final state highly sensitive to post-pulse relaxation dynamics and prone to back-switching. To overcome this, we propose a field-free scheme in which the transverse SOT drives large-angle precessional excitation while the perpendicular spin-transfer torque (STT) biases the trajectory toward the reversed $-z$ state. Micromagnetic simulations reveal a nonlinear switching boundary in the $J_{\mathrm{STT}}$--$J_{\mathrm{SOT}}$ parameter space, originating from the distinct dynamical roles of the two torques: SOT primarily governs the excitation and crossing of the dynamical separatrix, whereas STT controls the terminal trajectory and final-state selection. An analytical macrospin model, based on the stability analysis of the current-induced equilibrium, reproduces the critical-boundary trends as functions of current density, Gilbert damping $α$, and uniaxial anisotropy $K_\mathrm{u}$, and distinguishes dynamic anti-damping and static instability branches. Systematic analyses of pulse duration, damping, anisotropy, and the STT--SOT balance further demonstrate that reliable ultrafast switching requires not only sufficient excitation to cross the separatrix before pulse termination, but also precise control of the pulse-end magnetization state to minimize post-pulse relaxation. These results establish that the pulse-end state, rather than the instantaneous torque amplitude, is the decisive factor governing switching speed and reliability in coupled STT--SOT systems.

cond-mat.mes-hall↗

Implications of tachyonic phase transition in classically scale invariant general $U(1)_{X}$ models

We investigate the complementarity between collider searches and gravitational-wave (GW) observations in the classically scale-invariant general $U(1)_X$ extension of the Standard Model. The $U(1)_X$ scale is generated radiatively through the Coleman-Weinberg mechanism, while the electroweak scale is induced through the Higgs portal, without explicit mass terms in the scalar potential. Three right-handed neutrinos are introduced to ensure anomaly cancellation and acquire Majorana masses upon $U(1)_X$ symmetry breaking, which also generates the mass of the neutral gauge boson $Z'$. In the strongly supercooled regime, a tachyonic $U(1)_X$ phase transition can produce a stochastic GW background. Assuming three heavy Majorana neutrinos, we study their effects on the phase-transition dynamics and reheating, and estimate the resulting GW signals. We identify the regions of the $U(1)_X$ gauge coupling and $Z'$ mass accessible to future GW observatories, including LISA and DECIGO, and compare them with existing LEP and LHC constraints. We find that GWs from tachyonic phase transitions can probe regions with small gauge couplings and heavy $Z'$ bosons that are difficult to access through collider searches, demonstrating the complementarity of these probes.

hep-ph↗

Local Existence and Uniqueness for the 3D Compressible Navier-Stokes/Cahn-Hilliard System with Vacuum

We establish local existence and uniqueness of strong solutions to the compressible Navier-Stokes/Cahn-Hilliard system in a bounded smooth domain in $\mathbb R^3$, allowing the initial density to vanish. The initial data satisfy suitable regularity and compatibility conditions for the momentum and phase equations. The main difficulty arises from the simultaneous degeneracy of the momentum equation and the density-weighted phase equations, which complicates the construction of compatible positive-density approximations and uniform initial estimates. We overcome this difficulty by introducing a time-independent residual into the chemical relation. The residual preserves the phase compatibility exactly, keeps the initial phase field unchanged, and vanishes in $H^1$ as the approximation parameter tends to zero. For each positive-density approximation, we construct a local strong solution using a semi-discrete Galerkin scheme that discretizes only the phase variables and retains the full transport and momentum equations. We then derive a priori estimates independent of the density lower bound, which yield a common lifespan and allow us to pass to the vacuum limit. Uniqueness is proved using a modified difference energy that accounts for the coupling between the density and the chemical potential. To our knowledge, this is the first local well-posedness result for the three-dimensional compressible Navier-Stokes/Cahn-Hilliard system with initial vacuum.

math.AP↗

OmniDex: Scaling Dexterous Hand Grasping to Diverse Cluttered Scenes

Dexterous grasping is the foundational primitive in embodied AI, demanding massive data to train robust models. As real-world data collection is expensive, simulation has become the mainstream paradigm. Yet, while cluttered scenes best reflect real-world applications, learning to grasp within them is bottlenecked by a critical scarcity of large-scale data. To resolve this, we curate high-quality 3D objects and supporting bases, proposing a scalable seed-and-filter strategy that bypasses sluggish scene-level optimization. This yields an unprecedented benchmark comprising over 2.6 million scenes and 0.4B scene-specific grasp ground truths, featuring diverse realistic layouts paired with rich semantic and geometric observations. Furthermore, we introduce the OmniDex model to overcome the grasp multimodality and last-millimeter precision errors plaguing current generative models. By coupling Soft Winner-Takes-All learning with human-inspired physical constraints during training, and utilizing physics-driven ranking, our approach achieves robust dexterous grasping without the latency of post-optimization. Experimental results show that OmniDex model achieves state-of-the-art performance and strong generalization across diverse scenes, views, and unseen objects.

cs.RO↗

Retromorphic Testing of Quantum Compiler Passes

Quantum compilers play a critical role in transforming high-level quantum programs into optimized, hardware-compatible circuits. However, verifying the correctness of compiler passes remains challenging, as determining the expected output of large, deeply entangled quantum circuits is computationally intractable. This challenge is further amplified when compiler passes modify already complex circuit structures, making manual validation of transformed circuits impractical. In this work, we perform a systematic analysis of unit tests for quantum compiler passes in four quantum programming frameworks (PennyLane, Qiskit, Cirq, and pytket). Our findings indicate validation is dominated by program-content and program-metric assertions, and test circuits are generally small and shallow. Motivated by these observations, we introduce a testing methodology for automated validation of quantum compiler passes based on retromorphic testing and principles from the Hadamard test. This methodology analyzes a compiler pass, test circuit, and expected pass behavior to verify semantic preservation and intended structural modifications. We implement our methods in a framework, RetroQ, and apply it to compiler passes in PennyLane and Qiskit. Experimental evaluation reproduced several existing bugs as well as uncovered previously undetected defects, such as flawed symbolic parameter handling, incorrect commutation logic, failure to recognize self-adjointness of gates, and runtime crashes. These findings highlight the need for compiler-pass-specific testing methodologies to improve the reliability of the evolving quantum software stack.

quant-ph↗

Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models

Large audio-language models (LALMs) exploit multimodal evidence, yet task-irrelevant audio can alter text-reasoning decisions when listening is unnecessary. Aggregate Accuracy can hide this paired drift because audio-induced repairs and damages may cancel. Paired drift analysis and targeted interventions identify architecture-specific, intervention-sensitive late audio pathways as actionable control points. We introduce ICAP-Gate, which applies mechanism-guided, task-conditioned control to each model's pathway. Across four LALMs, two reasoning benchmarks, and environmental-sound and natural-speech interference, ICAP-Gate has lower point estimates for Influence Rate and Answer Flip than ungated inference in all 16 full-split model--condition evaluations. Fixed suppression degrades automatic speech recognition (ASR) across all four models, whereas ICAP-Gate matches ungated ASR performance by preserving the pathway for explicit audio-demand instructions. ICAP-Gate has lower paired-drift point estimates than mitigation prompting in all four evaluated settings and provides competitive stabilization relative to eight-sample Self-Consistency while using one generation per query; in controlled ARC measurements, Self-Consistency incurs $7.0$--$9.2\times$ ungated latency. These results establish selective modality influence control as a design principle for robust multimodal reasoning.

cs.SD↗

Randomized Scores and Diverse Timbres: Augmenting Automatic Music Transcription with Online-Generated Data

Automatic music transcription (AMT) is limited by the scarcity of audio recordings paired with precise symbolic annotations. Synthetic data can provide supervision at scale, but it remains unclear whether effective transfer depends on realistic score structure or broad timbral coverage. We study these factors separately through an online sampler--renderer pipeline. A unified corruption sampler ranges from unmodified MIDI clips through partial corruption to deeply randomized note-event distributions. The renderer converts these events to audio while independently controlling instrument and timbral coverage. A fixed transcription model is trained jointly on offline recordings and newly rendered examples. Controlled ablations reveal an asymmetry between the two factors: moderate corruption of the note-event distribution does not impair transfer and can improve it, whereas broader renderer-side timbral support consistently improves out-of-domain generalization under a fixed note-event distribution. Finally, online-rendered examples complement real and existing synthetic data in a strong combined-data regime. These results suggest that synthetic AMT data should prioritize coverage of note-level attributes and their timbral realizations over realistic joint score structure.

eess.AS↗

3DTexMOR: 3D Gaussian Multi-Object Removal via Texture-Space Inpainting

3D object removal aims to remove target objects from reconstructed scenes and complete the geometry and appearance of occluded regions. Existing NeRF- and 3DGS-based methods typically inpaint 2D images to guide 3D completion. However, complex multi-object layouts limit the surrounding context visible in each view, making 2D inpainting prone to artifacts. Inconsistent completions across views also introduce conflicting supervision and blurry reconstructions. We propose 3D Gaussian Multi-Object Removal via Texture-Space Inpainting (3DTexMOR). Our key idea is to perform inpainting in a unified texture space shared by all views. By combining complementary observations, this space provides richer context for recovering missing regions and promotes cross-view appearance consistency. We aggregate multi-view observations into texture maps, inpaint the missing regions, and reproject the completed maps into camera views to supervise Gaussian scene completion. To avoid the influence of view-dependent highlights and reflections, we decompose appearance and aggregate view-independent intrinsic attributes instead of RGB colors. We further introduce geometrically regularized Gaussian completion to constrain the geometry of the completed regions. Extensive experiments demonstrate visually plausible completions and state-of-the-art multi-object removal performance, improving PSNR by 5.8 dB and reducing LPIPS by at least 22% compared with existing methods.

cs.CV↗

From the Maxwell-Schrödinger equation to the Euler-Maxwell equation in the semiclassical limit

We derive the pressureless Euler-Maxwell system as a semiclassical limit from the Maxwell-Schrödinger equation. The method relies on studying a quantum modulated energy accounting for the self-consistent magnetic effects. Crucially, a correction is incorporated in the quantum modulated energy as to enable to close the Grönwall estimate provided a smallness assumption on the initial data is imposed. This is the first result on the semiclassical limit of the Maxwell-Schrödinger equation avoiding any regularization or conditional uniform in $\hbar$ bounds on the wave

math.AP↗

Information Divergence and Endogenous Evolution of Asymmetric Information

Standard theories trace information asymmetry to privately held information, heterogeneous signals, or costly information acquisition. We identify a distinct mechanism: otherwise identical rational agents observing a common Markov state at asynchronous times can hold different predictive distributions because their information differs in age. When the transition kernel separates information ages, independent Bernoulli observation arrivals make such divergence recur infinitely often almost surely, even with identical arrival probabilities. When agents instead choose costly updating intensities, independent update realizations can regenerate heterogeneous information ages among ex ante identical agents. In a stationary Gaussian AR(1) environment, the consequences depend on common staleness as well as the age gap: holding the gap fixed, older information raises individual uncertainty while reducing expected squared forecast disagreement and the value of certifying relative freshness. This generates stale consensus: low disagreement despite high uncertainty. With multiple ordered ages, costly certification supports partial-unraveling equilibria in which fresher types certify and sufficiently stale types pool. Under geometric information ages, more frequent updating weakly lowers the smallest and largest equilibrium disclosure cutoffs wherever positive-certification equilibria exist. The framework has implications for credit, insurance, financial markets, and multi-agent AI.

econ.TH↗