Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

What Response Marginals Miss: Adaptive Query Complexity of Functional Backdoor Recovery

Functional backdoor recovery finds any trigger whose attack success rate is at least a given threshold rather than to recover the planted trigger. We study the minimum number of queries required for this task under label feedback which returns the predicted class label. We construct two finite families of victim models that have exactly the same attack success rate for every victim and trigger candidate. The distribution of returned labels for every query is also identical under a uniformly chosen victim. These families form an explicit counterexample that despite the matched quantities, their optimal adaptive query complexities are \(Θ(\log H)\) and \(Θ(H)\) where \(H\) is the number of possible victims. The difference arises because the same non-target labels are associated with different sets of victims, so successive queries eliminate possible victims at different rates. This separation disappears when the response is reduced to binary feedback, which reports only whether the target label is returned. The separation also persists for every fixed failure probability below one. Finally, we realize the same recovery problems with trained CIFAR-10 ResNet-18 classifiers and verify the predicted optimal query budgets. These results show that attack success rate and the distribution of returned labels for each query are insufficient to determine the query complexity of functional backdoor recovery.

cs.CR↗

Model-Based Geometry-Aware Generative Optimization for Constrained Locomotion Planning

Constrained Locomotion Planning (CLP) for quadrupeds and humanoids, where robots must satisfy collision avoidance, contact consistency, kinematic feasibility, and support constraints, is challenging under high-dimensional dynamics and highly non-convex environments. Recent Model-Based Diffusion (MBD) approaches recast trajectory optimization as posterior sampling over trajectories, using known dynamics and Monte Carlo rollouts to analytically estimate the denoising score function without demonstration learning. While constrained variants further incorporate feasibility into model-based score rollouts and show promising performance, they are still limited by (1) lacking a task-modulated active constraint geometry that shapes the score direction and reverse stochasticity, and (2) using deterministic DDPM-style reverse transport without adaptive scheduling across different generative transports. Therefore, we introduce Model-Based Geometry-Aware Generative Optimization (2GO) for constrained locomotion, which turns active constraint geometry into executable denoising operators through normal- induced metric shaping, tangent-space stochastic filtering, and CFS-based retraction. 2GO further decouples generative transport from reverse stochasticity through an adaptive diffusion and flow-like schedule. Experiments on constrained quadruped and humanoid locomotion demonstrate strong performance in discrete foothold selection and continuous posture planning, with higher success rates, fewer violations, and improved execution compatibility.

cs.RO↗

Search for rare charm decays at BESIII

The BESIII experiment searches for rare and symmetry-violating decays with the largest charmonium and threshold open-charm data samples. This contribution reviews recent searches for charmonium weak decays, charm flavour-changing neutral currents, charged-lepton flavour violation, and Bose-symmetry violation. The weak-decay searches have improved several branching-fraction limits, while studies of rare charm-meson and charmonium decays probe flavour-changing neutral currents. Searches for charged-lepton flavour violation set limits in vector and pseudoscalar meson decays and constrain effective interactions. No significant signals are found in any of these searches. The neutral-kaon limits fall below the expectations of a proposed locality scenario, and the $J/ψ$ limit approaches the scale predicted for Standard Model CP violation.

hep-ex↗

Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models

Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering. However, these methods show inconsistent improvements across models and evaluation distributions, and the intervention into model behavior or internal states introduce safety-utility tradeoffs by over-refusal. Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety. We find that router logits indeed provide highly predictive signals of whether a multimodal input is safe or not. Motivated by this observation, we introduce a lightweight router-logit safety detector that reads out routing signals during prompt prefill and identifies unsafe requests before generation, without modifying model parameters or expert routing. Across Qwen3-VL and Kimi-VL, the proposed detector substantially reduces safety errors on the HoliSafe benchmark and resoundingly generalizes to out-of-distribution safety benchmarks featuring different safety patterns, including MISHard and MM-SafetyBench. The success of the proposed router-logit detector also suggests a broader perspective on model internals: rather than focusing only on manipulating internal components to steer behavior, simply reading naturally emerging signals and linking them to an external safety mechanism can provide a simple, effective, and non-intrusive complement to existing safety interventions.

cs.CL↗

Convergence Without Finite Exactness in the Eigenstate Quantum Bootstrap

The eigenstate quantum-mechanical bootstrap can converge to the physical moment set while remaining nonexact at every finite order. We prove this for a one-dimensional quasi-exactly solvable sextic oscillator at explicitly known eigenvalues, including excited levels. At every finite order, the lower and upper bounds on $u^2$ remain strictly below and above its physical value, even after adjoining arbitrary finite polynomial Gram tests and all eigenstate-annihilator relations visible in the retained moment space. Hence neither side of the centered observable admits a finite sum-of-squares certificate modulo the eigenstate annihilator. The gap persists under any finite set of additional localizing constraints whose weights are strictly positive on polynomial vectors. We then prove convergence for polynomial Schrödinger operators satisfying explicit sum-of-squares confinement and derivative bounds. The eigenstate relations give moment estimates uniform in the truncation order; these yield a regular Schrödinger representation supported on the prescribed eigenspace. Finally, a four-mode extension combines this fixed-observable obstruction with feasible position moments having no positive representing measure at every finite order. At level three, this construction yields nonphysical feasible points with strictly positive definite moment matrices.

math-ph↗

SCSM: A Traffic-Native Foundation Model for Transferable Website Fingerprinting

Website fingerprinting infers the websites visited by users from encrypted traffic metadata. However, models trained under fixed collection conditions often degrade as website sets, collection times, network paths, browsers, or defenses change. Existing transferable attacks either rely on handcrafted perturbations of individual traces or adapt language-oriented architectures to traffic, limiting their ability to capture traffic-native semantics. To address these limitations, we propose SCSM, a traffic-native foundation model for transferable website fingerprinting. Specifically, SCSM constructs pairs of pretraining views from the same group of unlabeled traces through Segmentation, Combination, Scaling, and Masking. These operations produce diverse observable patterns while preserving the underlying packet events and local traffic dynamics of real trace fragments. The corresponding windowed traffic counting matrices serve as inputs for contrastive pretraining of a state-space encoder without website annotations. The pretrained encoder is then fine-tuned on a small labeled support set, and predictions are aggregated across temporal scales at inference. Experimental results demonstrate that SCSM surpasses the strongest baselines by 16.4\% in average top-3 accuracy over six temporal-drift tasks on GTT23 and by 13.8\% in average macro-F1 over four cross-domain datasets. The code and datasets will be made available at https://github.com/SJTU-dxw/WF-SCSM.

cs.CR↗

Stochastic Subgradient Descent at Sharply Repulsive Points: A Bounded Noise Counterexample and Gaussian Noise Avoidance

Motivated by the open question of Bianchi, Hachem, and Schechtman (Bianchi et al., 2024, Remark 4), we show that in the non-weakly convex definable setting, omnidirectional noise with conditional fourth-moment control does not suffice for universal almost-sure avoidance of sharply repulsive critical points. We demonstrate this failure by constructing a globally Lipschitz, coercive, definable objective that is not weakly convex. For this objective and its sufficiently small linear perturbations, stochastic subgradient descent with independent noise uniformly distributed on a ball in $\mathbb{R}^2$ converges with probability $1$ to a nonminimal sharply repulsive critical point from every initial point in a specified ball. A two-cycle argument establishes uniform bounds on the rescaled iterates, yielding convergence to the sharply repulsive critical point. In contrast, for locally Lipschitz definable objectives, Gaussian noise ensures simultaneous almost-sure avoidance of any finite collection of sharply repulsive points. Through a construction of difference quotient functions, we show that convergence to any such point with positive probability would produce a stationary distribution with expected descent $0$, whereas a global positive lower bound on the Fréchet subgradient norms of the limiting functions forces strictly positive expected descent under the same distribution. This contradiction proves avoidance, with the everywhere positive Gaussian density playing a key role.

math.OC↗

Towards One-for-All Foundation Model for Attributed Graph Clustering

Attributed graph clustering aims to discover node groups by jointly exploiting node attributes and graph topology, yet its unsupervised nature makes model selection and adaptation inherently difficult. Existing methods typically train and tune a separate model for each input graph, leading to costly and fragile pipelines that often fail to transfer across graphs with different feature spaces, structural patterns, and attribute-structure correlations. In this paper, we study a one-for-all alternative: can a single model be trained once and directly applied to diverse attributed graphs without graph-specific training, fine-tuning, or hyperparameter search? We propose OFAG, a foundation model for attributed graph clustering. Building upon Prior-data Fitted Networks, OFAG learns a reusable clustering inference strategy from synthetic attributed graphs generated under broad priors over latent clusters, node attributes, and graph structures. To handle incompatible feature spaces across graphs, OFAG adopts a dimension-agnostic signal-wise graph encoder that treats each feature channel as a graph signal and models its response to shared graph filters. The model is trained with a hyperspherical clustering objective, producing clustering-friendly node representations in a single forward pass at inference time. On ten datasets, one frozen OFAG model achieves the best mean performance and average rank across NMI, ACC, ARI, and F1, while completing all ten datasets in 12.43 minutes total---over 6* faster than the second-fastest baseline and nearly 28* faster than the second-best on clustering quality. Our code and pretrained checkpoint are available at https://github.com/Cloudy1225/OFAG, allowing practitioners to directly apply OFAG to their own attributed graph datasets without additional training or tuning.

cs.LG↗

Fractal phase structure of QCD under imaginary rotation

We show that in hot QCD rotating at an imaginary angular velocity $Ω_I$, the angular velocity selects not only the temperature of the bulk but whether charge conjugation $C$ is broken there. For QCD at temperature $T$ and imaginary quark chemical potential $θ$, with $Ω_I/2π=p/q$ in lowest terms, the bulk far from the rotation axis is in the same thermal equilibrium state as nonrotating QCD at temperature $T/q$ and imaginary quark chemical potential $θ'=qθ+π(p+q+1)$; this follows from the Euclidean boundary conditions and locality alone, and holds nonperturbatively. Consequently, even at zero quark chemical potential, the bulk is placed at the Roberge-Weiss (RW) point whenever $p$ and $q$ are both odd, and $C$ is spontaneously broken there for $T>qT_{\rm RW}$, with $T_{\rm RW}$ the RW endpoint temperature. The bulks at $Ω_I=2π/3$ and $4π/3$, for example, are both at temperature $T/3$, yet only the former can break $C$. The set of imaginary angular velocities at which $C$ is broken has a fractal structure: raising $T$ adds rationals of ever larger denominator. At irrational $Ω_I/2π$ the bulk corresponds to zero-temperature QCD at any $T$ and remains confined.

hep-ph↗

APEX: Speculate smarter, not deeper

Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth. Fixed configurations cannot respond to changes in predictability, repetition, and acceptance during generation, so deeper drafting can increase wasted computation without proportional speedup. We introduce APEX, a learned controller that balances decoding speed and draft-token waste through request-level expert selection and block-level depth adaptation. APEX-Router selects among EAGLE-3, n-gram, and draft-model speculation for each request, while APEX-Depth adjusts draft length at each verification block using causal decoding signals and recent verifier feedback. APEX models accepted draft length as censored survival feedback, learning position-wise rejection hazards, block execution costs, and an action utility that balances throughput, accepted progress, and wasted tokens. This allows the controller to adapt speculation while retaining the target model's verification procedure. We integrate APEX into vLLM and evaluate it with Qwen3-8B across six workloads, achieving up to 5.24X speedup over autoregressive decoding. Across the aggregate evaluation, APEX-S achieves 4.27X speedup, while APEX-B achieves 3.27X speedup with a 41.0% relative reduction in wasted-token percentage compared with fixed n-gram speculation at k=16, providing distinct operating points for balancing acceleration and draft-token utilization.

cs.CL↗

Quantization Effects on Tool-Failure Recovery Vary Across Prompts and Evaluation Designs

Post-training quantization reduces the cost of deploying language-model agents, but its effect on recovery from temporary tool failures can depend on how recovery is evaluated. We compare 8-bit and 4-bit variants of Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct on twenty deterministic tool-use tasks and five prompts. The 8-bit-4-bit recovery comparison changes direction across prompts and evaluation targets. On tasks that both variants complete without faults under the same prompt, the difference ranges from 0 to +20.2 percentage points for Llama and from -50.0 to +35.0 points for Qwen. Full-pipeline point estimates favor 8-bit Llama under all five prompts, whereas the Qwen comparison changes direction across prompts. The evaluation target can also reverse the result. For Llama under one prompt, scoring each variant only on its own clean-passing tasks favors 4-bit by 17.5 points; scoring the same tasks for both variants gives no difference, while scoring the full pipeline favors 8-bit by 28.3 points. Executor leniency is a third such choice. Rescoring the same logs with strict output parsing, which 8-bit Llama violates far more often than 4-bit Llama under that prompt, turns that +28.3 into -15.0 while leaving Qwen essentially unchanged. These findings show that one prompt, one screened task set, and one scoring policy do not establish a stable conclusion about quantized-agent robustness. Evaluations should compare variants on matched tasks, report full-pipeline success for deployment decisions, state the scoring policy, and quantify uncertainty across tasks rather than injected fault sites.

cs.AI↗

Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell

Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds. Many such systems add a persistent tier storing and recalling reasoning traces, usually validated by an ablation reporting an accuracy gain. We measure both on one three-tier agent architecture. Decomposition delivers: peak KV working set of 14.3 MiB per query against 35.5 and 35.3 MiB for single-pass and retrieval-augmented baselines. The persistent tier does not: across eight controlled dataset pairs at n=100 per arm it costs +0.368 MiB [+0.167, +0.590] of peak cache and produces no detectable accuracy change (+0.015, 95% CI [-0.011, +0.046]). We argue the null is structural: single-question benchmarks supply each item with its own evidence and score it independently, and correctness requires resetting stored traces between conditions, so recall has nothing informative to retrieve. Reaching it took four measurement corrections -- three inflating the apparent benefit, the fourth making an effect that size look resolvable -- none visible in the results table. We give the conditions an agent-memory ablation must satisfy and detection procedures that need no knowledge of the specific defect.

cs.AI↗

From Laboratory to Road: Evaluating Wearable Gaze Accuracy for Driving

Bird's-eye-view (BEV) representations have become a widely used interface between perception and planning in autonomous driving, but they encode what is in a scene, not what is behaviorally relevant to a human driver. Gaze offers a compelling behavioral signal for this gap, yet wearable eye trackers are routinely deployed as if their spatial output were ground truth, despite known sensitivity to head motion, illumination, and calibration drift. We present, to our knowledge, the first unified framework for quantifying wearable gaze accuracy under real driving conditions. Our on-road study contains 41 validated scenes in which one driver fixated a vehicle's license plate. Gaze error is measured as the angular difference between the plate center and the gaze direction estimated by the glasses. Separate indoor studies with the same driver and device systematically analyze how distance, illumination, head motion, target motion, and gaze eccentricity affect both systematic bias and gaze precision. The mean on-road error was 4.58 degrees. Applying an offset estimated from the indoor recordings reduced it to 1.10 degrees and improved all 41 scenes. Because this offset varied between sessions, reliable BEV supervision may require online recalibration and condition-dependent estimates of gaze uncertainty.

cs.CV↗

Multi-Robot Multi-Goal Motion Planning with Stochastic Skills

As robots are increasingly deployed in groups and share workspaces to execute real-world tasks, planning their concurrent motions around complex manipulation skills becomes essential. These skills involve continuous physical execution and may exhibit stochastic behavior, resulting in variable execution times and uncertain continuous trajectories. Existing planners either limit execution to single-robot scenarios, rely on open-loop paths, or use post-hoc scheduling that prevents dynamic coordination. In this paper, we address this gap by integrating stochastic skills into sampling-based multi-robot planning by formulating the problem as a Markov Decision Process (MDP) over a multi-modal composite roadmap. For stochastic skills, solving the MDP yields a reactive policy that allows controllable robots to dynamically adapt their motions in response to other robots' execution of manipulation skills. By resolving skill uncertainty directly at planning time, this approach avoids the pessimism of conservative baselines and unlocks robust, dynamic multi-robot coordination. Code for the planners is available at https://www.vhartmann.com/stochastic-skills.

cs.RO↗

Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents

A central capability of embodied agents is to accomplish complex objectives through sequences of interdependent tasks. Yet existing visual goal-conditioned policies underlying these agents are typically evaluated on isolated interactions where the target is already visible, and thus do not capture the conditions that arise during continuous long-horizon task execution. In such settings, each task begins from the state left by the previous one: the agent may end at a different position and orientation, the world may have been modified, and the next interaction target may lie outside the current field of view. As a result, agents relying on such policies may struggle to proceed to the next task when they cannot ground their target in the current observation. To address this challenge, we propose Attacca, a new approach that trains visual goal-conditioned policies on complete search-to-interact trajectories using goal images decoupled from the execution environment. Attacca uses context-decoupled goal sampling to pair each demonstration with a class-compatible masked goal image from another world, removing direct scene and pose correspondence. It learns dense current-view grounding through a target-mask prediction head, providing auxiliary supervision beyond action imitation. We further introduce behavioral-phase conditioning that teaches the policy to distinguish Search, Approach, and Interact stages and adapt its control as execution progresses. We evaluate Attacca on multiple short- and long-horizon embodied tasks in Minecraft. Our method achieves 39.0-47.5% clean success, improving over the strongest baseline by 1.7-2.4x. On long-horizon tasks, it attains 54%, 30%, and 28% completion, yielding up to a 7x improvement.

cs.AI↗

Extending Pathwise Gradients to Discrete Random Variables via Finite-Order Relaxation

Pathwise gradients are preferred for continuous random variables because they are unbiased, low variance, and work with a single sample. For discrete variables, however, the pathwise identity cannot generally be exact for every differentiable function. We propose a general framework to construct finite-order exact pathwise gradient estimators for a range of common discrete variables such as Poisson. The estimator is the least-norm solution among all solutions that are unbiased for polynomials of degree at most. The resulting estimators preserve the hard forward sample, require no temperature tuning, and can be implemented in a few lines of codes. Against other admissible solutions, our estimator is unique and minimizes weight variance; in contrast, prior works use categorical variables or augmented representations to approximate non-categorical variables that induces excess variance and computations. To understand approximation bias for functions beyond the prescribed class, we also derive a non-asymptotic bias bound. In experiments our low order methods match or improve tuned baselines across linear, nonlinear and hierarchical latent-variable models, while out-speeding competitors in every runtime benchmark.

cs.LG↗

OOPMAS: Object-Oriented Multi-Agent Systems for Query-Level Workflow Generation

Multi-agent systems (MAS) powered by large language models have shown strong performance across code generation, mathematical reasoning, and question answering. However, existing methods for automating MAS design mostly operate at the task level, producing a single fixed workflow per benchmark that is applied uniformly to all queries. This assumption fails under realistic conditions. Query difficulty varies widely within a task, and real-world workloads mix heterogeneous task types. We introduce OOPMAS, a training-free framework that generates both the agent set and the coordination workflow at the granularity of individual queries. Agents are represented as object-oriented class definitions with dedicated roles, tools, and persistent state, and workflows are expressed as executable main functions over these agent objects. A dynamic skill library accumulates structured lessons from execution feedback across optimization rounds, enabling in-context improvement without any gradient updates or fine-tuning. On a mixed-task benchmark of queries spanning code, math, and QA, OOPMAS achieves 89.6% accuracy, outperforming the strongest baseline by 18.1 percentage points. A model-swap study across four LLM backbones shows consistent scaling, reaching 92.4% with the strongest model.

cs.AI↗

Image-Space Refraction Correction for Underwater 3D Reconstruction: Warping Flat-Port Views into Pinhole Perspective

Consumer-grade cameras in flat-port housings are widely used for underwater exploration and mapping of coral reefs and seafloor habitats due to their low cost and accessibility. However, refraction at flat-port interfaces causes bowl-shaped deformation in reconstructed scenes and camera trajectories, compromising the metric accuracy required for mapping and navigation. To remove the dominant refractive distortion before reconstruction, we introduce a physics-based refraction correction in image space. Our method is downstream-agnostic: the refraction-corrected images can be directly used as input to existing reconstruction and SLAM algorithms. We characterize the refractive distortion through ray-tracing simulations and validate our correction on two real underwater datasets with differing scene structures. Compared with conventional and refractive Structure-from-Motion (SfM), our approach removes reconstruction deformation while registering more frames and maintaining low reprojection error. The correction further generalizes across diverse reconstruction and VSLAM backends, demonstrating its broad applicability to downstream vision pipelines.

cs.CV↗