Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 631 records · Page 35Linked to original sources

Parameter-Free Zeroth-Order Optimization with Ellipsoidal Sampling

Zeroth-order optimization methods are essential for solving black-box problems where gradient information is unavailable or expensive to compute. This paper presents POEM-ES, a novel parameter-free stochastic zeroth-order algorithm that extends the recent POEM method by integrating subspace preconditioning with ellipsoidal randomized sampling. In contrast to traditional zeroth-order approaches that rely on isotropic random directions, POEM-ES performs anisotropic sampling guided by a fixed structural symmetric positive semi-definite (SPSD) preconditioner $\hatΣ$ that encodes the underlying low-dimensional geometry. Under a standard structural spectral normalization where $λ_{\max}(\hatΣ) = 1$, we introduce the use of the empirical effective dimension $d^* = \operatorname{tr}(\hatΣ)$, which reflects the intrinsic dimensionality of the problem and guides both the sampling and randomized smoothing parameter schedules. In practice, such a preconditioner can be effectively obtained via pilot sampling, historical trajectories, or domain-specific expert knowledge. We prove that POEM-ES achieves a dimension-reduced convergence rate under low-rank structural assumptions, requiring only $\tilde{\mathcal{O}}\left( \frac{\left( r^2 κ(\hatΣ) + d^* \right) L^2 D_{\mathcal{X}}^2}{\varepsilon^2} \right)$ stochastic zeroth-order oracle queries. The method remains fully parameter-free and demonstrates significant improvements over the original POEM in problems with low-rank structure where $d^* \ll d$. Numerical experiments on hinge-loss binary classification tasks using LibSVM datasets confirm the practical superiority of the proposed approach.

math.OC↗

Marking Contour Tones in Yorùbá: A Typographic and Computational Proposal

Yorùbá is a tonal language in which contour tones pose persistent orthographic challenges. These are especially notable for personal names and lexical items whose conventional spellings avoid vowel lengthening that would otherwise provide a host syllable for the second tone. A particular concern is a class of names in which the conventional spelling does not just omit tonal information but inverts the meaning of said name, sometimes asserting the opposite of what the name intends. This paper describes the problem, illustrates the inadequacy of current solutions, and proposes the adoption of the caron and circumflex marks. These are symbols with precedent in Yorùbá phonological scholarship since Olmsted (1951), used as orthographic conventions on single vowels to encode rising and falling contour tones, making them accessible for the first time through standard keyboard input and computational text processing. The proposal is supported by an implementation in the WriteYoruba keyboard and the TTSYoruba speech synthesizer, whose architecture and listener evaluation are reported separately (Tubosun et al., 2026).

cs.CL↗

Hear the World in Stereo: Learning Dynamic Spatial Correspondence for Immersive Joint Video-Audio Generation

Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sound location to evolve consistently with the motion of its corresponding visual source. We refer to this property as Dynamic Spatial Correspondence and propose StereoBind, a framework that binds visual source motion to stereo sound generation. StereoBind uses motion tracks to coordinate visual motion and stereo audio through three complementary mechanisms. Visual Motion Binding establishes source-aware audiovisual correspondence, the Spatial Track Encoder captures absolute source positions, and Residual Track RoPE models relative motion. For supervision and evaluation, we construct StereoWorld-29K, a large-scale stereo audio-video dataset with paired motion tracks, and StereoWorldBench for measuring audiovisual spatial consistency. Experiments show that StereoBind substantially improves spatial alignment in stereo audio generation over existing models while preserving overall audiovisual quality.

cs.CV↗

Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds

Persistent spatial memory enables embodied agents to navigate familiar environments across repeated visits. However, targets may move while unobserved, including during navigation, making remembered locations unreliable by the time an agent arrives. Despite advances in memory retrieval and state prediction, accounting for continued hidden world evolution and revising beliefs under limited visibility remain challenging. We study Evolving-World Navigation, where agents infer target locations from intermittent observations, predict their states at inspection time, and revise beliefs using visual evidence. We propose EvolvingNav, which constructs a time-indexed belief from timestamped 3D object histories through a structured persistence-relocation model. The belief distinguishes persistence at the last observed location from relocation to alternative locations and retains probability mass outside the known candidate set. An event-driven filter propagates the current belief as time elapses, forecasts target occupancy at candidate inspection times, and incorporates new RGB-D evidence. Negative observations downweight location hypotheses according to calibrated, visibility-conditioned detection probabilities, while evidence tracking prevents repeated use of the same observations. A frozen, zero-shot vision-language controller uses the updated belief to choose actions and replan. We further introduce EvoWorld-Bench, a benchmark grounded in human activity traces, comprising 54 scenes and 803,680 tasks with controlled changes before and during navigation. In simulation and real-robot experiments, EvolvingNav improves navigation success and search efficiency over the evaluated baselines. Paired experiments show the clearest gains under learnable temporal patterns, while ablations demonstrate the value of preserving uncertainty and incorporating visibility-aware evidence.

cs.AI↗

Surface-Dependent Phonon Dynamics in 9-Armchair Graphene Nanoribbon Arrays

Atomically precise graphene nanoribbons (GNRs) have electronic properties tunable through their width and edge structure, making them promising building blocks for nanoscale electronics and optoelectronics. Their integration into devices also requires understanding how the supporting surface and neighbouring ribbons affect their vibrations. Here, we use temperature-dependent Raman spectroscopy to compare five configurations of 9-armchair GNR arrays differing in substrate, alignment and coverage. Transferring the same unaligned film from Au to a Raman-optimised Al$_2$O$_3$-coated support reduces the measured $D$- and $G$-mode redshift rates by about 80\%. This reduction is consistent with a substantial thermoelastic contribution. The fitted strain and anharmonic contributions depend on the assumed thermal expansion and strain transfer. Compared with the low-coverage array, the dense aligned Au array has higher extrapolated reference frequencies ($ω_0$). On first heating, its $G$-mode linewidth decreases and then increases, forming an unusual intermediate-temperature minimum. The $D$-mode linewidth shows a similar, less pronounced variation. These findings establish substrate choice and array arrangement as routes to engineer nanoribbon vibrations, extending control beyond the atomic structure defined during synthesis.

cond-mat.mtrl-sci↗

Fair Variable Selection

Algorithms are increasingly being used to help automate and improve data-driven decisions, but care must be taken to prevent such algorithms from learning discriminatory patterns from historical data and perpetuating their biases. Statistical notions of fairness aim to mitigate either a model's disparate impact on disadvantaged groups (thus ensuring group fairness) or the resulting disparate treatment of individuals with similar features (thus ensuring individual fairness). Simultaneously mitigating disparate impact and disparate treatment is generally impossible for non-trivial models, necessitating a compromise. In this paper, we introduce the Fair Lasso and Fair Posterior as methods for selecting fair covariates in generalised linear models. By targeting variables that are simultaneously strong predictors of the response and weakly dependent on the sensitive group memberships, we aim to achieve favourable trade-offs between disparate treatment and disparate impact. Additionally, our selected set of fair features can be used as the conditioning set of legitimate features in the paradigm of Conditional Demographic parity (CDP) when no prescriptive legal framework exists.

stat.ME↗

Finite-slope p-adic L-functions for $\mathrm{GL}_{3}$

Let $Π$ be a regular algebraic cuspidal automorphic representation of $\mathrm{GL}_{3}(\mathbb{A}_{\mathbb{Q}})$ of weight $(a, 0, -a)$, and let $p$ be a prime. Generalising work of Loeffler--Williams in the nearly-ordinary case, we construct $p$-adic $L$-functions for each small slope regular $P_{1}$-refinement of $Π$ at $p$ consistent with the conjectures of Coates--Perrin-Riou and Panchishkin.

math.NT↗

Growing an Agent/Prover Interface: Evolutionary Tool Design for Cost-Efficient Theorem Proving in Rocq and Lean

Recent achievements in AI-assisted mathematics require intensive interaction of agents with proof assistants to generate machine-checked proof certificates. Agents interact with proof assistants such as Rocq or Lean through an interface that controls what the agent receives from the prover and the cost of these interactions. Today, these interfaces are adapted from tools designed for humans and not optimized for agents. We propose an evolutionary method where a frontier model incrementally proposes new features and only keeps the ones that improve the overall performance of smaller models. We demonstrate the effectiveness of our method by growing, on a curated set of mathematical problems, ROCQ-MCP-EVOLVE, a new MCP server for the Rocq prover. On the held-out test split of miniF2F-Rocq, an agent equipped with ROCQ-MCP-EVOLVE outperforms both the baseline that only exposes the Rocq compiler and an established MCP server, across four models from two families, in success rate, cost per solve, and time per solve. Although evolved for Rocq, the resulting server transfers to Lean, improving cost and time per solve on a subset of PutnamBench. We release ROCQ-MCP-EVOLVE and its port to Lean.

cs.AI↗

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box$^2$-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box$^2$-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two complementary training strategies: counterfactual supervised fine-tuning improves robustness, while outcome-based reinforcement learning can shift the balance toward greater use of helpful workflows. We further find that this behavior extends beyond workflows to other forms of external information, improving peer correction and robustness to corrupted memory. Together, our results identify selective reliance on fallible external information as a dimension of agent reliability not captured by task performance alone.

cs.CL↗

Structural Limits of the Information-Theoretic Uncertainty Decomposition

Uncertainty estimation in machine learning typically decomposes uncertainty into aleatoric uncertainty (AU) and epistemic uncertainty (EU) using the standard information-theoretic framework. However, in practice, two critical issues arise: entanglement (AU and EU are highly correlated) and epistemic collapse (EU magnitude shrinks with increasing model capacity). We analyze this framework on a functional level and discover that significant portions of the assumed AU, EU range are infeasible in finite settings, and cannot be attained with any class probabilities. We characterize how this infeasible region scales with the number of classes and Monte Carlo samples $N$ (e.g., from ensembles with $N$ members), revealing it is bounded by $\text{AU} \leq \log(2)/N$. Crucially, the infeasible region's boundary helps explain epistemic collapse: when model confidence is high, $\text{AU} > \text{EU}$ is guaranteed by this fundamental structural limitation. Our findings show that increasing ensemble size mitigates epistemic collapse by reducing the infeasible area. Lastly, we caution against interpreting AU and EU as independent quantities in low AU regimes, since we show they are coupled when $\text{AU} \leq \log(2)/N$.

cs.CV↗

Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents

Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext, iteratively crafts skills that evade detection while still delivering the payload and performing the benign task: moving the payload from code into natural language leaves static analysis inert, while framing it as the skill's legitimate purpose and splitting instructions across files keeps the LLM stage below its blocking threshold. Across three open-source models, Pretext achieves up to 97\% and 77\% against a frozen detector and a co-adaptive one, respectively, revealing major gaps in current skill scanners.

cs.CR↗

LLM Persona Unlearning

Pre-training equips large language models (LLMs) with a broad repertoire of behavioral patterns associated with roles, styles, values, and goals. Post-training teaches conditional enactment and makes a helpful Assistant the default, but it does not erase alternative modes from the weights; explicit prompts can therefore elicit personas that repeatedly shape judgment, language, and action. In open-weight settings, runtime controls can be removed, motivating persona unlearning: a weight-level edit that makes a designated persona difficult to elicit and enact on unseen contexts. We introduce PersonaUnlearnBench, a model-specific paired benchmark spanning six LLMs from three families and five personas, with aligned forget/retain sets, held-out instruction paraphrases, and four-axis evaluation. The benchmark shows that standard unlearning methods cannot reliably erase the target persona without sacrificing meaningful generation or general utility. We therefore propose PaCE, which compares target and desirable responses to the same questions to locate an internal behavior direction, then trains target-prompt states away from the target mode and toward the matched desirable response. Experiments show that PaCE consistently suppresses target personas with high response quality and useful counterpart behavior, at moderate utility cost. These results establish persona unlearning as a distinct behavior-level editing problem and a practical route toward persistent control of latent LLM response policies.

cs.CL↗

Grounding with Confidence: Controllable Generative Video Temporal Grounding

Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explicit interval-level confidence scores to guide candidate selection. We separate candidate generation from acceptance by scoring individual intervals within the original decoding pass. A lightweight confidence head reads pooled decoder states, providing an explicit score trained for interval selection. Offline verifier scores supervise the head on fixed candidate sequences, and temporal-overlap labels adapt it to current rollouts during reinforcement learning. GT-anchored candidate-pool supervision and set-level optimization train the generator. The resulting scores support ranking, threshold-based selection, and rejection without invoking an external verifier at inference. On a fixed OMTG-Bench candidate pool, confidence raises query-macro Recall@0.5 from 9.95% to 14.42% over generation order at a 10% global return budget, and from 26.48% to 31.12% at a 25% budget. The continuous scores let downstream applications adjust return budgets or acceptance thresholds to match their precision-recall preferences, without regenerating candidate intervals.

cs.CV↗

Proximal Empirical Bayes for Sparse Regression with Posterior Decision Support

In this work, we revisit estimation and variable selection in sparse Gaussian regression, possibly under affine constraints. We adopt a middle-ground approach between optimisation and fully hierarchical Bayesian methods, preserving much of the computational efficiency of the former while adding posterior uncertainty quantification. In particular, we develop an empirical Bayes framework that separates shrinkage calibration from posterior-informed selection while using the same log-concave model throughout. The resulting workflow is modular: different methods can be used for noise-variance estimation, shrinkage calibration and posterior sampling. At the same time, these tasks are linked, and numerical choices made at one stage can propagate through the workflow and ultimately affect the decision about which variables to retain. We investigate in detail one instance of the framework based on a Laplace prior, with the amount of shrinkage determined by stochastic approximation proximal gradient. The proximal machinery used for this calibration also supports posterior simulation. We further show how the empirical Bayes calibration should be modified when homogeneous constraints reduce the number of free parameters. Nonhomogeneous affine information is instead incorporated softly; when this information is strong, posterior sampling can become poorly conditioned, and we develop a preconditioning strategy to improve mixing and posterior exploration. Synthetic experiments examine how scale calibration, posterior run length and sampler conditioning affect estimation, uncertainty quantification and the final sparse decision. On a diabetes dataset, posterior uncertainty agrees closely with established Bayesian analyses, while the final decision provides a sparser practical summary.

stat.ME↗

Gestalt: Large Multimodal Interplay Model

In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal models are reaching a bottleneck: existing approaches focus primarily on accommodating additional modalities while overlooking the distinct characteristics of each modality and the relations among them. Motivated by the multistage property of human multisensory perception, we propose a multimodal interplay pyramid that organizes multimodal modeling as a progression from modality-specific processing, through cross-modal alignment, to deeper multimodal integration. Guided by this pyramid, Gestalt adopts a unified discrete diffusion framework and an interplay-partitioned architecture, with learnable interplay tokens mediating cross-modal exchange and integration. The pyramid also structures its data organization and training strategy. Strong performance across image generation, multimodal understanding, and text-only evaluation shows that Gestalt significantly improves cross-modal integration while preserving modality-specific information, effectively harnessing the strengths of diffusion-based multimodal models and offering a promising path toward unified multimodal intelligence. Project page: https://GeWu-Lab.github.io/Gestalt.

cs.CV↗

Absence of Continuous Spin Particles in Superstring Theory

The study of continuous spin particles, as general low-energy implications of relativity and quantum mechanics, has intensified in recent years with potential experimental signatures. In this note we show that these particles are absent in perturbative super-string theories, generalising a result in arXiv:1302.4771 to the supersymmetric case. This stands out as a general low-energy prediction of all perturbative superstring constructions.

hep-th↗

Complexity study of the Hartle-Hawking state in JT gravity

The Wigner function, in general, takes on negative values, and the amount of negativity in the Wigner function gives an operationally meaningful measure of the complexity of simulating the quantum state on a classical computer. In this paper, we study the growth of Wigner negativity of the Hartle-Hawking state of Jackiw-Teitelboim gravity under time evolution. We work in the gravitational length basis, at genus zero, and at high temperature, $β\ll 1$. Our main analytic result is simple: to leading order in $β$ the state is a Gaussian wavepacket of minimum uncertainty, sitting at rest at the turning point of the Liouville wall. Its Wigner function is therefore positive, and the negativity is $1+δ(β,t)$, where the correction $δ$ is smaller than any power of $β$. We show that $δ$ is an even function of time, so there is no linear growth at $t=0$, and that once the packet has reflected off the wall its leading-order evolution is a Clifford shear, which cannot change the negativity. We also numerically study the time evolution of the Wigner function and its negativity. We observe that $δ$ stays below $10^{-11}$ at $t=0$ and through the reflection, then rises slowly over a few tens of $β/π$, and then approaches a late-time value consistent with $δ_\infty(β) \simeq 0.08\, e^{-4.4/β}$ for $0.25 \le β\le 1$: the lower the temperature, the larger the late-time negativity. The exact survival amplitude $Z(β+it)/Z(β)$ fixes the spread complexity to order $t^4$, and it also fixes the seed-normalised Wigner diagnostic of our earlier work (arXiv:2607.04065, arXiv:2607.17346). We take this as evidence that the length basis is ideally suited for a dual, semi-classical description of the dynamics at genus zero. Beyond genus zero the length basis is overcomplete, and we make no statement about finite $e^{S_0}$.

hep-th↗

Empty Commitments: When Agents Promise What They Cannot Deliver

A chatbot that says "I will remind you tomorrow" will not run again until the user writes. We call such a promise an empty commitment: a promise of action after the current turn that nothing in the agent's tools or runtime can carry out. Unlike a broken promise, its emptiness is decided by the agent's configuration at the moment of speaking, so it can be detected from a single turn, before deployment or at run time. We define empty commitments on top of commitment semantics, with three failure types, an anchoring condition for promises that a tool could make real, and an outcome taxonomy that separates these failures from honest deferrals and from over-refusal. We build a checker (a setup-blind detector, deterministic feasibility rules, and a response judge) validated against 400 human labels. On a controlled benchmark of 293 follow-up requests across five setups that add one persistence affordance at a time, four open-weight models of 8-14B parameters fail on 45.9% of responses when no tool exists and nothing is stated. A frontier model fails on 4.4%, but it gets there by deferring and asking, not by using the tools it has: promises made without the enabling call remain in every model. Telling the model its runtime, the cheapest fix, cuts open-weight failures nearly in half where nothing is doable and changes nothing where a scheduler exists; a directive capability card removes most failures at the largest cost in over-refusal; running the checker in the loop and rewriting flagged replies removes more at a smaller cost. Code, prompts, model outputs, and human labels are released.

cs.AI↗