Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Characterizing Statistical Separability in TP-CRIV for Probabilistic AI Models

Third-party challenge-response identity verification (TP-CRIV) enables an independent verifier to assess whether a claimant possesses a model identical to a remotely deployed model without directly accessing the reference model. However, for probabilistic AI models, repeated executions of the same query may produce different outputs and therefore different verification observations. This raises the question of how such stochastic evidence should be accumulated and how much evidence is required for reliable verification. In this work, we characterize statistical separability in TP-CRIV of probabilistic AI models. Specifically, we relate challenge-wise behavior of matching and non-matching provers to verification-level separability. The characterization explicitly describes how the numbers of independent challenges and repeated responses affect detection performance and enables the verification budget required for a target AUC to be estimated. We instantiate the proposed characterization for LLMs using open-ended challenges. The experiments demonstrate matching-non-matching separation, close agreement between theoretical and empirical AUCs, and consistent estimates of the minimum verification budgets. These results provide a statistical basis for relating probabilistic model behavior to verification-level separability and the evidence required for third-party verification.

cs.CR↗

RideBench: A Large-Scale Exogenous-Aware Benchmark for Ride-Hailing Time Series Forecasting

We release Ride-Hailing, a large-scale ride-hailing time series dataset synthesized from DiDi's marketplace data across 200 spatial areas. Ride-Hailing spans four consecutive years at half-hourly granularity and covers three representative exogenous scenarios: Weather Disturbance, Holiday Effect, and Large-scale Event Impact. Built upon Ride-Hailing, we introduce RideBench, a comprehensive benchmark for exogenous-aware ride-hailing forecasting, covering both regular week-ahead forecasting and long-horizon 8-week-ahead forecasting with up to 2,688 prediction steps. RideBench evaluates over 30 representative forecasting methods, including endogenous-only models, exogenous-aware models, and time series foundation models. Our results show that future-known exogenous variables provide clear benefits in regular week-ahead forecasting, especially under weather, holiday, and large-scale event (e.g., major sporting events and concerts) scenarios. However, current exogenous-aware models still struggle to fully capture disturbance-induced pattern changes under complex external contexts. For long-horizon forecasting, existing models cannot simultaneously achieve low pointwise errors, accurate broad trends, and reliable near-term forecasts. These findings reveal a clear mismatch between existing forecasting models and real-world ride-hailing requirements, highlighting the need for models that can better exploit future-known exogenous information, scale across heterogeneous areas, and support long-horizon planning. By introducing Ride-Hailing and RideBench, we aim to encourage the community to study these practical challenges in real-world ride-hailing forecasting.

cs.LG↗

CARE: A Lightweight Plug-in Gated Correction and Uncertainty-aware Module for Long-term Time Series Forecasting

Multivariate long-horizon forecasting is critical to electricity load scheduling and traffic flow management, and to financial risk control. Existing deterministic backbones output a single trajectory, masking heterogeneous prediction difficulty across horizons and channels and providing no localized reliability signal. We present CARE (Corrective branch with Aligned context and Relative-error Estimation), a lightweight plug-in that enhances any deterministic forecaster without architectural redesign. Operating in parallel with the base model, CARE resamples historical context to match the forecast horizon, learns residual correction patterns from this aligned history, and applies scale-aware bounded updates modulated by per-coordinate sigmoid risk gates. A multi-objective loss jointly optimizes forecast accuracy, residual tracking, risk alignment, and base-model anchoring. Across eight benchmarks with three representative backbones, CARE improves accuracy with marginal parameter and latency overhead. Its risk gates reliably identify high-error regions: on Weather, the highest-gate tertile exhibits nearly four times the error of the lowest-gate tertile, offering planners an interpretable per-step trust signal. Code is available at https://github.com/CG-BNYC/CARE.

cs.LG↗

Analysis of p-Wave Sommerfeld Resonance Behavior in Finite-Size Dark Matter

Different stages of cosmic evolution impose different requirements on the dark matter annihilation cross section. Sommerfeld enhancement provides a possible way to accommodate them. However, in the point-like dark matter scenario, satisfying these requirements with either $s$-wave or $p$-wave Sommerfeld enhancement can strongly constrain the dark matter and mediator masses and their coupling. A finite dark matter size introduces an additional parameter and may relax these constraints. Motivated by this possibility, in this work we study the $p$-wave Sommerfeld resonances of finite-size dark matter (like proton or neutron) without specifying its internal structure. We find that the $p$-wave contribution can exceed the $s$-wave contribution. Nevertheless, the finite-size $p$-wave resonances are weaker than those in the point-like case. Despite this suppression, their velocity dependence remains similar to that of point-like dark matter. We then extend our analysis to nugget dark matter composed of a small number of constituents. In this case, the $p$-wave Sommerfeld enhancement exhibits behavior similar to that of point-like dark matter.

hep-ph↗

PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation

Vertical-domain few-shot classification remains challenging for small language models, as limited supervision makes it difficult to acquire domain-specific decision knowledge. On-Policy Distillation (OPD) can improve teacher-guided adaptation by supervising student-generated rollouts, while GRPO-based reinforcement learning can further refine downstream predictions. However, existing KD-to-RL pipelines typically rely on globally fixed transition schedules, ignoring that different samples may require different amounts of teacher-guided acquisition before reward-driven refinement. We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity. PIVOT moves low-perplexity samples to GRPO for reward-driven refinement while keeping high-perplexity samples under OPD for continued domain knowledge acquisition. Experiments on Banking77 and HWU64 show that PIVOT consistently outperforms continued OPD and globally synchronized OPD$\rightarrow$GRPO baselines under the same number of post-warm-up student optimization steps, achieving stronger downstream performance and more stable training dynamics.

cs.LG↗

PMTRM: Pseudo-Memory Temporal Re-encoding Module for Embodied Policy Learning

Robotic manipulation often contains repeated motions whose local observations look similar at different phases. When these phases require different actions, a policy that relies mainly on the current observation may repeat completed motions or switch phases at the wrong time. To address this phase ambiguity, we present the Pseudo-Memory Temporal Re-encoding Module (PMTRM), a lightweight plug-in module with only 7.61M parameters that encodes a bounded history of executed states and actions into a latent sequence for existing policies. To help distinguish phases, a temporal heterogeneity objective penalizes positive similarity between distant positions in this sequence, while anchor and reconstruction losses preserve information needed for action prediction. The reconstruction decoder is used only during training, leaving the temporal re-encoder to supply history to the policy at inference. We train the module progressively on synthetic sequences and robot data, then jointly with the policy, using temporal masking to accommodate partial histories. This integration retains the original action head and action space and adds auxiliary losses to the original policy loss. Experiments with multiple policy backbones in simulation and on a real robot show improved task success on tasks with phase ambiguity, with little additional computation.

cs.RO↗

Skill Constellations: Tracing the Supply Chain of Agent Skills on GitHub

Agent skills are SKILL.md instructions and scripts that AI coding agents such as Claude Code and Codex run with the permissions of their user. Developers share skills by copying them between repositories, which makes them a software supply chain without a registry, versions or provenance. The origin of a copied skill, the reach of a security fix and the repositories that warrant review are therefore unknown. Studies that record which repositories hold a skill at a single point in time cannot reveal who copied it from whom. We contribute the first dated copy network of agent skills, built from the git history of every SKILL.md in GitSkills and covering 2,193,119 skill adoptions across GitHub, together with an interactive viewer. A few repositories are the source of almost all copies, and GitHub stars do not identify them. Skill copies almost never change with their source, and a fix at the source therefore rarely reaches them. We fit a model of which repositories others copy from and use it to rank repositories for audit. Reviewing the 100 repositories it ranks highest prevents 14.9% of later adoptions of high-risk skills, against 0.5% for the 100 most starred, which gives security engineers a short list to check before a skill spreads. Platforms should therefore distribute versioned references rather than copies. Project Website: https://fahdseddik.github.io/Skill-Constellations/

cs.SE↗

SACQ: Structured Decoding with Memory-Conditioned Refinement for Long-Horizon Forecasting

Long-term time series forecasting (LTSF) models predominantly employ patch-based encoders terminated by a flatten readout head that maps the entire encoded historical memory to all future steps through a single shared projection. This implicit coupling of future positions obscures position-specific historical-to-future alignment and amplifies sensitivity to corrupted inputs and extreme supervision noise. We present SACQ, a plug-in structured prediction head that replaces flatten readout while keeping the encoder unchanged. SACQ adopts a two-stage decoding pipeline: it first establishes a coarse patch-grid forecast scaffold, then refines each future position through cross-attention over historical memory and merges the attention-derived correction with the coarse scaffold via a learned per-patch gate. To stabilize optimization under long horizons and noisy labels, we further propose a batch-adaptive scaled log-cosh loss that automatically calibrates robustness to the current residual scale, suppressing outlier gradients while preserving MSE-like sensitivity for typical errors. SACQ attains top-tier test MSE/MAE across PatchTST, DLinear, and patch-Mamba backbones with only modest incremental overhead in parameters and latency. Under inference-time input corruption and training-set label-noise stress tests, SACQ substantially outperforms flatten readouts, with ablation studies validating each architectural component.

cs.LG↗

VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding

Long-form video understanding often involves multiple questions about different aspects of the same recording. Yet existing video agents typically process each question through an isolated tool-use trajectory. This repeatedly restarts video exploration and memory construction, missing opportunities to acquire evidence jointly and progressively build a shared understanding that supports the complete question set. We introduce \textbf{VAMR} (\textbf{V}ideo \textbf{A}gent for \textbf{M}ulti-Question \textbf{R}easoning), which coordinates all questions about a video through one shared tool-use trajectory. At each round, a persistent policy model can invoke tools for one or more unresolved questions and submit answers for questions with sufficient evidence. Question-conditioned visual perception retrieves fine-grained clues for several questions in one call, while layered multi-question memory integrates reusable context into a shared video story and preserves separate evidence for individual questions. After supervised fine-tuning initializes this interaction protocol, we propose question-horizon policy optimization (\qhpo) to optimize shared trajectories in which questions progress and finish at different rounds. Specifically, a question-level critic estimates the value of each active question, while round alignment maps each question advantage to the rounds that directly serve it before the aligned advantages are aggregated to optimize the shared actor. Across LVBench, Video-Holmes, and LongVideoBench, VAMR achieves the highest accuracy overall and the fewest reasoning rounds among iterative methods. On LVBench, it reaches 62.1\% accuracy, exceeding VideoARM by \textbf{4.3} points while reducing reasoning rounds and processed frames by \textbf{85.9\%} and \textbf{61.4\%}.

cs.CV↗

2D Coprime Pilots for Delay-Doppler Sensing in OFDM-ISAC Systems

Integrated Sensing and Communication (ISAC) is envisioned to endow future 6G systems with seamless sensing capabilities. To support efficient sensing with minimum communication overhead, sparse pilots embedded within communication frames have emerged as a promising solution. Along this line of research, existing studies have achieved engaging results in maximizing the unambiguous sensing region. However, jointly maximizing the sensing region and sensing accuracy remains challenging due to the lack of a unified performance metric and an effective pilot design framework. This paper jointly optimizes the unambiguous sensing region and sensing accuracy for estimating delay-Doppler (DD) parameters in Orthogonal Frequency Division Multiplexing (OFDM)-ISAC systems, where sensing mutual information (SMI) is adopted as a unified performance metric to characterize the overall sensing capability. Specifically, the joint optimization is formulated as an SMI maximization problem by systematically resolving sensing ambiguity and optimizing sensing accuracy. In particular, based on the generalized Bezout identity, we derive a 2D (timefrequency) coprime condition, which, as far as the authors know, is the first necessary and sufficient condition to achieve the unique estimation of DD parameters in the literature. Under this unambiguous condition, we further propose an Adam-Guided Iterative Refinement (AGIR) algorithm to optimize the sensing accuracy. Numerical results demonstrate the advantage of the proposed framework over existing designs, owing to the freedom offered by the 2D coprime condition in optimizing the sensing accuracy.

eess.SP↗

IntactWorld: Joint World Modeling with Intact Features

While recent video generation models synthesize highly realistic visuals, they lack a genuine understanding of intrinsic real-world logic. Existing methods attempt to understand the world by internalizing diverse world knowledge, yet constrained by computational overhead or dimensionality alignment, their learning processes inevitably compress features, causing a severe loss of structural information. To address this, we propose \textbf{IntactWorld}, a \textbf{Joint World Modeling Architecture} utilizing uncompressed \textbf{Intact Features}. Since data naturally reside on a low-dimensional manifold within a high-dimensional space, predicting the flow velocity $v$ within this uncompressed high-dimensional space induces a severe manifold gap. To successfully eliminate this optimization bottleneck, our framework instead predicts the clean feature $x_0$ at intermediate layers. Furthermore, to mitigate the computational overhead of incorporating complete world knowledge, we introduce a \textit{Full-to-Compact Training Paradigm}. By replacing raw full features with highly refined CLS tokens, this paradigm enables efficient single-branch guidance, reducing spatial memory consumption by 11.4\% and cutting inference latency by 43.8\%. Extensive evaluations demonstrate the effectiveness of IntactWorld, outperforming established baselines by 2.46 points on the VBench 2.0 benchmark.

cs.CV↗

Higher-Order Action Supervision Makes A Strong Policy Class

Modern data-driven decision-making methods, such as imitation learning (IL) and reinforcement learning (RL), have achieved great success in solving many complex tasks. However, these methods often suffer from serious control instability and robustness issues when applied in real-world applications such as robotics and autonomous driving, posing notable challenges for their practical deployment. We argue that this instability issue stems largely from their limitations in solely supervising and optimizing zeroth-order actions (i.e., the action labels), failing to account for higher-order action dynamics and temporal consistency. In this paper, we show that simultaneously supervising both zeroth- and first-order actions can dramatically enhance policies' performance and control robustness. To achieve this, we introduce a novel and elegant loss scheme supported by formal theoretical guarantees that can equip any off-the-shelf policy model (e.g., deterministic, stochastic, or flow policies) with the capability for higher-order action supervision, without requiring any structural modifications. Moreover, our proposed method can serve as a lightweight plug-and-play module that seamlessly integrates with a broad spectrum of existing offline RL frameworks. Extensive evaluations on OGBench and D4RL demonstrate that our approach yields substantial performance and robustness improvements across a wide range of continuous control environments. Notably, our method can also enhance policies' out-of-distribution (OOD) generalization capability in the challenging low-data regime, making it an ideal tool in tackling many real-world control problems.

cs.RO↗

Multimodal Remote Sensing Image Registration: A Comprehensive Review, Challenges and Prospects

Multimodal remote sensing image registration is a crucial prerequisite for the collaborative processing and downstream application of remote sensing data, such as image fusion, change detection, and target recognition. However, significant variations in radiometry, geometry, scale, viewpoint, and time often exist between multimodal images. These differences, driven by varying sensor geometries, physical radiation mechanisms, imaging platforms, and environmental disturbances, pose severe challenges to achieving high-precision, robust registration. This paper systematically reviews the progress of mainstream multimodal remote sensing image registration methods. Based on their registration pipelines, existing approaches are categorized into three main types: region-based, feature-based, and deep learning-based methods. We detail the core principles, representative algorithms, advantages, and limitations of each category. Additionally, we summarize publicly available multimodal image datasets in the remote sensing domain, analyzing their specific characteristics and applicable scenarios. Finally, we highlight current bottlenecks in high-precision registration research and outline future development trends. This review aims to provide a comprehensive reference and valuable insights for researchers in related fields.

cs.CV↗

Dynamic Mechanism and Information Design in AI-Agent Environments: A General Analytical Framework

This paper develops a general framework for dynamic mechanism and information design with AI agents, encompassing optimal contracting and integrating evolving hidden information and hidden action with partial verification, technological authorization, and multi-agent interaction. It nests conventional dynamic screening, hidden-action contracting, and conditional dynamic information design, while allowing reporting, action, and information-use decisions to interact through a common continuation mechanism. We establish a dynamic revelation principle and a recursive characterization of implementability in which truthful reporting and post-report obedience are jointly determined through continuation incentives. In quasilinear optimal-contracting environments, dynamic envelope and integral-monotonicity methods yield a verification-adjusted dynamic virtual-surplus representation and identify three distinct margins: dynamic information rents, verification rents, and the implementation cost of hidden action. These margins generate a branched hierarchy separating conventional dynamic screening, hidden-action contracting, and their joint problem. We also study the joint design of monitoring, information use, and incentives: more informative monitoring capacity expands the designer's opportunity set, but full disclosure need not be optimal, and information use can itself alter hidden-action incentives. With multiple agents, correlated information, peer discipline, coupled incentives, joint feasibility, and continuation opportunities create additional strategic interactions. Although motivated by AI-agent environments, the framework applies more broadly to dynamic agency settings with similar economic features.

econ.TH↗

Optical Signals Synthesized from an Optical Lattice Clock with Uncertainty of 1.3E-17

We report an accurate optical frequency synthesizer, which can generate single-frequency laser light with high frequency stability and accuracy at desired frequencies over a wide optical region. The frequency of the output signal is divided from an 171Yb optical lattice clock via an accurate optical frequency divider based on an optical frequency comb. Therefore, the output of the optical frequency synthesizer inherits the frequency accuracy from the Yb optical clock. The frequency uncertainty of the 171Yb optical lattice clock is evaluated to be 6.8 mHz, corresponding to a fractional frequency uncertainty of 1.3E-17, mainly limited by the blackbody radiation shift and the lattice-induced light shift.

physics.optics↗

Who Pays the Review Cost? Triage, Fairness, and Accountability in AI-authored Pull Requests

AI coding agents are moving from local code assistance into pull-based workflows, where generated contributions must be reviewed, explained, and maintained within existing project norms. Although recent work has begun to characterize AI-authored pull requests (AIPRs), less is known about how reviewers govern their entry into review, how AI authorship reshapes credibility and fairness, and what intake mechanisms protect review sustainability. We report a mixed-method questionnaire survey of 239 practitioners from 31 countries with code-review experience and varying exposure to AIPRs. In the scenarios and self-reports elicited by the survey, AI authorship was not a categorical rejection signal. Instead, respondents described review effort as conditional on whether an AIPR arrived as an accountable contribution, with bounded scope, project-grounded rationale, validation beyond Continuous Integration (CI), contributor responsiveness, and identifiable post-merge ownership. This conditional logic extended to newcomer AIPRs, where respondents emphasized visible participation in the current review process over profile-level reputation alone. Qualitative responses further described shifts in mentoring, scrutiny, deferral, and routing when human stewardship was difficult to observe. We conceptualize missing rationale, validation, and ownership as reviewability debt, the work reviewers must absorb when generated code lacks sufficient human grounding. These findings reframe AIPR governance around making human judgment observable before generated contributions consume scarce reviewer attention.

cs.SE↗

Uniformly convex stalks and Kucher's $\ell_{\infty}(E)$ problem

Kucher asked whether $\ell_{\infty}(E)$ is a Grothendieck space for every super-reflexive Banach space $E$. We answer this question affirmatively by proving the stronger result that $\ell_{\infty}(E)$ has Pełczyński's property $(V)$. Combined with the natural dual representation of $\ell_{\infty}(E)$, this yields that every non-weakly compact operator on $\ell_{\infty}(E)$ fixes a copy of $\ell_{\infty}$. The main ingredient is a general property-$(V)$ theorem for function modules over zero-dimensional compact Hausdorff spaces with stalks admitting a common modulus of uniform convexity. The proof is based on Gierz's integral representation of functionals and the weak compactness of the associated canonical measures. In the case of $\ell_{\infty}(E)$, the relevant stalks are ultrapowers of $E$, and super-reflexivity provides the uniform geometric control required by the theorem. We also derive several consequences for operators on $\ell_{\infty}(E)$, including weak compactness into separable spaces, reflexivity of separable quotients, and compactness of operators into Schur spaces.

math.FA↗

MATE4D: Matrix-Guided Editable 4D Generation from a Single Image

Generative models have rapidly pushed content creation be-yond 2D imagery toward dynamic 3D and 4D scene synthesis. Yet pro-ducing realistic and temporally stable 4D content from a single image is still difficult because one view provides limited structural cues and weak motion evidence. We introduce MATE4D, a framework that converts one input image into editable dynamic 4D content. Our method constructs a spatio-temporal multi-view image matrix with text-guided background manipulation, delivering coherent supervision over viewpoint, appear-ance, and motion. These synthesized observations are used to optimize 3D Gaussian primitives, which are then animated through a lightweight deformation module to form a 4D representation. The resulting scenes preserve geometry more faithfully, maintain smoother temporal behavior, and keep background edits more consistent, reducing context ambiguity and motion artifacts. Experiments on Objaverse-XL and Diffusion4D show that MATE4D outperforms strong baselines in visual quality, effi-ciency, and controllability, supporting practical AR/VR content creation.

cs.CV↗