Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Optimal scenario design for climate emulation

As deep learning for physical systems continues to grow in popularity, efforts to improve generalizability have primarily focused on designing architectures that embed physical constraints. However, for machine-learning surrogate climate models (emulators), we show that the low structural diversity in existing scenarios commonly used to generate training data places a ceiling on predictive skill. Here, we examine whether training datasets themselves can be optimized to improve generalization. We introduce a method to create datasets that produce emulators capable of generalizing to new, structurally different scenarios absent from the training data. We use a differentiable Simple Climate Model (SCM) to calculate the sensitivity of emulator loss to perturbations in the training data, iteratively updating the training data to maximize emulator skill. For an SCM, training on one scenario optimized in this fashion outperforms an emulator trained on six standard ScenarioMIP pathways. We achieve this higher predictive skill despite training on a smaller dataset, finding that our emulator successfully isolates distinct physical behaviors of different climate forcing agents (e.g., greenhouse gases vs. aerosols) without single-forcing runs. We then demonstrate that scenarios optimized using an SCM, when used to drive an intermediate-complexity climate model, produce a training dataset that yields a more skillful emulator than training on ScenarioMIP outputs. Our results suggest that, in the compute-constrained environment of running full-scale climate models, generating a small number of dynamically rich scenarios provides greater marginal value for emulation and characterizing system responses than expanding the suite of traditional emissions pathways.

physics.ao-ph↗

Machine-learning-driven kinetic discovery of carbon interstitial color centers in diamond

Diamond hosts optically active point defects central to quantum technologies, yet the carbon self-interstitials introduced during growth and irradiation compete with them and form new defects whose configurational landscape is poorly charted, as subtle energy differences govern the competing minima and pathways. Here we build an interstitial-focused dataset by active learning and benchmark three machine-learning interatomic potentials -- GAP, NEP and the equivariant MACE -- against density functional theory for energies, forces and migration barriers. MACE reproduces the reference energetics and relative stabilities, whereas the others can misorder the ground states. Annealing molecular dynamics with the validated potentials uncovers a series of previously unreported carbon interstitial clusters, from di- to octa-interstitials -- several introducing in-gap states of interest as colour centres -- and shows that their metastability is governed by kinetically accessible pathways rather than energetic ordering. These results chart the interstitial defect landscape and accelerate defect discovery for quantum technologies.

physics.comp-ph↗

Temporal Self-Imitation Learning

Long-horizon policies trained with reinforcement learning can still achieve high return through inefficient interactions, while rare efficient behaviors discovered during training may be forgotten. We argue that temporal efficiency itself provides a source of self-supervision for reinforcement learning. We introduce Temporal Self-Imitation Learning (TSIL), a reinforcement learning framework that mines temporally efficient successful trajectories generated during learning and converts them into reusable supervision for future policy improvement. TSIL progressively refines learning using configuration-conditioned adaptive temporal targets derived from fast successful trajectories, while preserving and replaying efficient behaviors through efficiency-weighted self-imitation learning. Across 30 long-horizon tasks spanning robot manipulation and interactive navigation, TSIL consistently improves task success rates, learning efficiency, behavioral efficiency, and robustness to unstable training conditions.

cs.RO↗

EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies

Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-relevant cues become occluded or unobservable over time. While existing memory-augmented methods utilize historical context, they either suffer from severe information bottlenecks, incur high latency via decoupled dual systems, or rely on unselective buffers that accumulate massive visual redundancies. To address these limitations, we introduce EventVLA, an end-to-end framework founded on the concept of sparse visual evidence memory that comprises two core components: foundational visual anchors to retain initial and short-term contexts, and a dynamic Keyframe Evidence Memory (KEM) module. Specifically, KEM directly predicts future keyframe probabilities from the VLA's latent embeddings to autonomously capture and store sparse, task-critical visual events. This foresight-driven mechanism empowers the policy to dynamically evaluate the future causal utility of current observations, preserving transient visual evidence before it becomes unobservable. Furthermore, we propose RoboTwin-MeM, a diagnostic benchmark specifically designed to evaluate non-Markovian manipulation tasks with interactive visual evidence. Extensive evaluations show that across 17 memory-requiring simulation tasks and 4 real-world bimanual tasks, EventVLA achieves an average success rate improvement of +40% over state-of-the-art memory-augmented VLAs.

cs.CV↗

SoftSkill: Behavioral Compression for Contextual Adaptation

Natural-language skills let agents reuse task knowledge, yet deploying a long Markdown document makes the model interpret that knowledge anew on every call. We ask whether the behavior induced by a skill can be carried by a compact, trainable context. SoftSkill initializes virtual token embeddings from a skill document and optimizes a soft skill with next-token prediction while keeping the language model frozen. The resulting conditioning sequence can occupy the skill section or another supported prompt location. On Qwen3.5-4B, a 32-token soft skill improves over no skill by 7.6, 42.1, and 1.3 points on SearchQA, LiveMath, and DocVQA. It also exceeds the optimized textual skill on SearchQA and LiveMath by 4.5 and 12.5 points, respectively, while replacing skill documents of hundreds to thousands of tokens. The same approach improves multi-step agent execution: on Qwen3.6-35B-A3B, OfficeQA and ALFWorld rise by 8.2 and 14.1 points over no skill; on Qwen3.5-4B with aligned action decoding, ALFWorld success rises from 44/134 with the untrained initialization to 91/134 after training. These findings show that a skill document can serve as the starting point for a continuous control that improves both answers and actions without updating the backbone.

cs.AI↗

CRAX: Fast Safe Reinforcement Learning Benchmarking

Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving. While benchmarks have been central to progress in RL, existing 3D physics-based safety benchmarks remain computationally slow, limiting large-scale experimentation and rapid prototyping. To address this gap, we propose CRAX (Constrained RL Accelerated with JAX). Built on top of the MuJoCo XLA (MJX) physics engine, CRAX leverages vectorized operations and hardware acceleration, yielding up to 200x faster training over comparable CPU-based safety benchmarks. The benchmark features eight tasks spanning three difficulty levels and multiple agent morphologies. Evaluating seven popular safe RL methods, we find that none dominates across tasks, and that learning safe policies from pixels remains largely unsolved.

cs.LG↗

Gen2-IC: Bridging Generative Models and Image Codecs through Latent Transport

Diffusion-based image compression has achieved strong perceptual quality at ultra-low bitrates. However, existing codecs are often tied to specific backbones and specialized components, making diverse, rapidly evolving generative models difficult to reuse. This raises a natural question: Can modern generative foundation models be connected to image compression through a simple and extensible interface? Two insights guide our design: stronger generative priors make a simpler codec interface viable, and generation and compression can be intrinsically linked through latent transport. We therefore propose Gen2-IC with two stages: (1) Latent Compression maps clean image latents to entropy-constrained latents; and (2) Latent Transport refines them with one near-terminal update based on the pretrained model. Gen2-IC requires neither auxiliary conditioning signals nor task-specific backbone modifications. With lightweight adaptation and no distillation, it supports fast encoding and one-step decoding across multiple bitrates. We validate Gen2-IC on SD-2.1, SANA-1.5, FLUX.1-dev, and Qwen-Image-2512, spanning U-Net and Transformer architectures as well as diffusion and flow-matching formulations. With stronger priors, Gen2-IC delivers gains below 0.05 bpp: the Qwen variant leads diffusion-based generative codecs in reconstruction fidelity (PSNR), perceptual similarity (LPIPS and DISTS), and recognizer-based semantic fidelity (OCR CER/WER and face-ROI similarity) across four benchmarks.

eess.IV↗

OVIG: Optimistic Verification of AI Training Integrity via Gradient Signals

The rapid growth of AI has increased the demand for domain-specific models. Post-training of open-source models offers a more economical way to meet this growing demand, but the cost of accelerator infrastructure often pushes organizations to outsource the process to third-party providers. An untrusted provider may deviate from the declared training procedure to save computation or inject malicious behavior. A potential solution is to audit the training by having an independent verifier replay the training and compare the results. This faces two key challenges: benign numerical drift from floating-point computation across heterogeneous accelerators is difficult to distinguish from malicious deviations, and the checkpoints and metadata required for verification introduce substantial storage and transmission costs. Existing approaches either have security limitations or high deployment costs. We present OVIG, an optimistic verification framework that verifies training using an empirical boundary on gradient differences calibrated from honest heterogeneous replay. Gradient differences that exceed this boundary are treated as malicious deviations. OVIG further uses optimistic verification and partitions training into stride-$s$ intervals, retaining evidence only at interval endpoints to reduce storage and transmission costs. Across shortcut training attacks and adaptive targeted manipulation attacks, OVIG maintains \(0\%\) ASR on language, vision, and diffusion workloads. On Qwen3, increasing the stride from \(s=1\) to \(s=2000\) reduces off-chain storage and evidence transmission by \(1996\times\) while preserving \(0\%\) ASR and only adds \(14.3\%\) system computation overhead relative to unverified training.

cs.CR↗

Resolved Ages and Stellar Metallicities in Progenitors of Milky Way Analogs: A Closer Look at their Star Formation Histories since $z=5$

We present the evolution of the resolved mass-weighted age, stellar metallicity, and sSFR of 872 Milky Way Analog (MWA) progenitors up to redshift $z=5$ from the Canadian Unbiased Cluster Survey (CANUCS). The metallicity and mass-weighted ages were obtained via spatially resolved SED-fitting with the non-parametric code Dense Basis. We split the sample into mergers versus non-mergers using the merger parameter from the Gini-$M_{20}$ plane obtained through Gini-$M_{20}$ analysis of the morphology of the stellar mass maps with Statmorph. Across our redshift range, non-mergers have negative or flat average age gradients from $-0.022$ to 0.005 dex/kpc, and positive or flat sSFR gradients from $-0.089$ to 0.092 dex/kpc, consistent with inside-out assembly. The average $\log(Z/\Zsun)$ gradients for non-mergers range from $-0.029$ to 0.044 dex/kpc, however, positive gradients only appear between $2 < z < 3$. At every redshift epoch, mergers typically have flatter age gradients, more negative sSFR gradients, and similar metallicity gradients compared to non-mergers. We divide the property maps of ongoing mergers into separate regions based on their component galaxies, and find little to no difference between the components' average ages or metallicities, but the less massive of the merging system is on average $0.1-0.4$ dex higher in sSFR. Our results point to major mergers contributing some momentary disruption to the general trend of inside-out mass assembly, but does not upend the overall picture of MWA disks growing inside-out over cosmic time.

astro-ph.GA↗

Congruences for Overcubic Partition $k$-Tuples

In the last few years, a number of authors have proved divisibility properties satisfied by various functions which count the number of overcubic partition $k$--tuples of weight $n$ for small values of $k$. In this work, we use generating functions to prove some of their results as well as multiple infinite families of new congruences for overcubic partition $k$-tuples which do not yet appear in the literature. In particular, we focus on a new perspective which provides insights as to why these functions are often divisible by powers of 2, and we also prove families of congruences whose moduli are odd. For example, we prove that, for all $m\geq 0$, $\OL{b}_{4}(22m + 11) \equiv 0 \pmod{11}$ and we also prove infinite families such as $\OL{b}_{9l+2}(9m+3) \equiv 0 \pmod{3}$ for all $m,l\geq0$.

math.NT↗

Is Agent Code Less Maintainable Than Human Code?

Maintainability is a core dimension of software engineering, shaping how code is written, reviewed, and developed over time. While coding agents have demonstrated strong performance on single-issue tasks, it remains unclear how maintainable their code is when future agents build on top of it, potentially leading to compounding downstream effects. We investigate how agent code compares to human code in these maintenance settings, presenting CodeThread, a framework to construct controlled experiments from repository-level coding benchmarks. Applying CodeThread to four frontier coding agents and four benchmarks, we find that agents are less effective at resolving tasks when building on agent code compared to human code, with task resolve rate drops of up to 13.1%. Regression analysis reveals that many traditional software engineering maintainability metrics do not explain this difference. Instead, the clearest signals are subtler behavioral differences in agent code, such as changes to input validation and error handling, along with differences in downstream code size and task difficulty. These findings highlight the need to evaluate these systems not only by immediate task resolution but also by code maintainability, and point to potential sources of downstream errors introduced by agent code.

cs.SE↗

Probabilistic Storage and Retrieval of Quantum Superchannels for "Retrospective" Intervention

Storing an unknown quantum computation in a quantum state and retrieving it at a desired later time is a challenging task, hindered by the no-programming theorem of quantum computations. In the previous studies on the task of probabilistic storage-and-retrieval (pSAR) of quantum channels, the maximum probability of exactly retrieving a single unknown unitary channel from a quantum state in which the unknown unitary has been encoded via multiple calls to the unknown unitary channel is derived. In this work, we consider a higher-order version of pSAR, the probabilistic storage-and-retrieval of definite-causal unitary superchannels, which are physically modeled by sequences of unitary channels with open slots where arbitrary channels can be inserted between the unitary channels for intervention. This task requires activating the "retrospective'' intervention functionality on the superchannel, beyond its normal intervention functionality. We propose two protocols: partial teleportation, which is supported by our numerical results for the small queries, and staircase backstitch, which achieves unit success probability asymptotically as the number of queries increases. We also derive a universal inversion protocol for unitary superchannels.

quant-ph↗

Multiplicative Colombeau algebras and the Nyman--Beurling criterion for the Riemann hypothesis

This paper establishes an equivalence between the Riemann hypothesis and the association, together with uniform $L^2$-boundedness, of a moderate net in a Colombeau-type algebra built from polynomially damped Báez-Duarte sums. The regularization is performed by multiplicative (Mellin) convolution, which respects the dilation symmetry of the Beurling functions and guarantees that every approximant lies in the $L^2$-closure of the Beurling space. The equivalence is unconditional under the Riemann hypothesis: it uses only the qualitative convergence of Báez-Duarte, Mazur's theorem, and the classical Nyman--Beurling criterion. As a separate quantitative refinement, we prove that under two additional hypotheses on the non-trivial zeros of $ζ$ (simplicity and separation), the damping error admits a power-law bound with an explicit constant. This refinement is independent of the equivalence and is not used in its proof. The exponential damping $e^{-k\eps^2}$ is discussed as an open problem.

math.CV↗

A bilinear approach to the finite field restriction problem, II

Let $P_3$ denote the three-dimensional paraboloid over a finite field of prime order in which $-1$ is not a square. We prove that the Fourier extension operator associated with $P_3$ maps $L^2$ to $L^r$ for $r>\frac{176}{51}=3.45098\ldots$. The argument combines the author's bilinear approach to the problem with point-line incidence estimates. We also prove that the extension operator associated with the paraboloid $P_6$ in six dimensions maps $L^2$ to $L^{8/3}$. This was previously known up to but not including the endpoint, and is the sharp $L^2$ estimate in six dimensions. Finally we observe that the endpoint restriction conjecture for $P_3$ in finite fields implies that the integer lattice points on the $3$-d Euclidean paraboloid are a $Λ(3)$ set.

math.CA↗

Flow Games with Public Arcs: the Least Core and the Nucleolus

We study flow games with public arcs, an extension of classical cooperative flow games that allows players to use public resources. In these games, a coalition corresponds to a set of arcs, while certain arcs, called public arcs, can be used freely by any coalition. The value of a coalition is the maximum flow value achievable using the arcs controlled by the coalition along with the public arcs. We investigate two solution concepts, the least core and the nucleolus. Both solution concepts provide fair ways to allocate the value of the grand coalition among individual players. We provide polynomial-size formulations of the least core of these games. Besides, we give a polynomial-time algorithm for computing the nucleolus, whether or not the core is empty. This resolves a long-standing gap left by Potters et al. [GEB'06], whose algorithm for the nucleolus of simple flow games with public arcs assumes a non-empty core.

econ.TH↗

A new boundary mass for asymptotically flat half-manifolds

We introduce a boundary analogue of the Gauss--Bonnet--Chern mass for asymptotically flat half-manifolds with non-compact boundary. We prove that this mass is well defined and establish the corresponding positive mass theorems for graphical and conformally flat graphs. Also we provide a Penrose-type inequality for the mass $\mathfrak{m}_{a,B}(g)$.

math.DG↗

Low-power analogue neural networks with trainable nonlinear connections for continuous control

Physical neural networks promise low-power machine learning by computing directly with analogue device physics, but most architectures force nonlinear device responses to act as scalar weights. Inspired by Kolmogorov-Arnold networks, we place trainable nonlinear functions on the connections, making each physical connection a learnable computational element. Realising these functions as analogue band-pass filters on field-programmable analogue arrays, we find that the benefit is task-dependent and follows from the smoothness of the physical basis: the networks represent smooth, continuously valued targets, including robotic kinematics, continuous control, and photovoltaic maximum-power-point tracking, with far fewer nodes and connections than multilayer perceptrons, but offer no parameter-efficiency advantage on classification-like decision boundaries. Trained networks transfer to hardware across approximately 35,000 connections with quantified fidelity, and a dedicated CMOS implementation is projected to operate at approximately 30 microwatts. A memristive realisation reproduces the same behaviour in simulation, indicating that the advantage comes from placing trainable nonlinearity on connections, rather than from a particular device.

cs.LG↗

CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking

Reliable provenance for LLM outputs requires multi-bit watermarks that remain robust under editing while maintaining low false-positive rates. Existing ECC-based LLM watermarks rely on hard-decision decoding, discarding token-level reliability information and limiting robustness under post-generation edits. We propose CORE-BREW, a COnstant-hit-Rate Embedding extension of BREW for multi-bit watermarking. CORE-BREW calibrates the watermark channel by targeting a fixed hit rate $p^\star$, yielding closed-form per-token log-likelihood ratios (LLRs) for soft-decision decoding. It incorporates entropy-aware erasures to limit perturbations in low-entropy contexts and combines likelihood-based scoring with soft-decision list decoding to exploit soft evidence. Experiments on open-source LLMs under token-level edits and paraphrasing demonstrate that CORE-BREW generally improves detection robustness and payload recovery over the BREW baseline while maintaining low observed false-positive rates. Despite higher conditional perplexity, BLEU and BERTScore remain close to those of unwatermarked text, indicating comparable reference-based translation quality.

cs.CR↗