Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof

AI systems can now write and optimize production GPU kernels, but validating them remains an important challenge. Evaluating the kernel on a few random inputs and checking that its outputs match a trusted reference kernel within numeric tolerances is not sufficient: races can cause nondeterministic behavior that fails to manifest in tests, and numeric tolerances can hide bugs and cause false positives even after extensive calibration. To address this challenge, we present RESOLVE, which combines testing and formal verification to build a comprehensive kernel validation pipeline. It operates in three steps: First, it tests for nondeterminism using binary instrumentation that perturbs execution timing to expose races. Second, an agent rewrites the candidate and reference kernels to obtain "reduced-concurrency" versions that are simpler to analyze but still produce bitwise-identical outputs in all tests. Third, the reduced kernels are formally analyzed in the F*/Pulse framework and prove that they perform the same computation on real numbers. This sidesteps the need for numeric tolerances. We show that RESOLVE can validate a broad selection of kernels using KernelBench, and prove equivalence across fused GEMMs in three state-of-the-art frameworks and languages: CUTLASS, Triton, and Gluon. It also analyzes mega-kernels, notoriously difficult to validate, and finds four previously unreported issues, including two clear bugs. We show that agents can use RESOLVE to repair the issues, with minimal performance impact, highlighting that agents can optimize aggressively when they can rigorously check their results.

cs.PL↗

Two-Sample Testing via Path-based Inference

Modern deep generative models are primarily studied for their ability to generate realistic samples, yet the generative dynamics they learn can also serve as objects of statistical inference. We develop this idea for two-sample testing, the problem of deciding whether the same distribution generated two finite datasets. Using stochastic interpolants, we connect both distributions to a shared Gaussian bottleneck, so that each half of the resulting path is a Gaussian channel acting on a single population. We prove that the null hypothesis holds if and only if the population denoiser, or equivalently, the velocity fields of the two halves, coincide at any single noise level, which amounts to a reflection symmetry of the path about the bottleneck. Deviations from this symmetry yield a continuum of two-sample witnesses, which we estimate via held-out regression risks on learned denoisers and velocities and aggregate along the path; under an information-theoretic weighting, the aggregated discrepancy equals the Jeffreys divergence between the noise-smoothed distributions. Calibrating the resulting statistics by permutation yields tests that are valid in finite samples for any trained networks and consistent when the fields are learned accurately. On a synthetic benchmark and three image benchmarks, the proposed tests improve power over the strongest baseline by up to 33 percentage points at an equal total sample budget, with the best choice of regression representation and path weighting depending on the data modality. These results show that generative paths provide a principled representation for statistical testing, extending stochastic-interpolant models beyond generation.

stat.ML↗

Spend Bytes on Breadth: Precision-Count Trade-offs for Decode-Time KV Compression in Long Chain-of-Thought Reasoning

Reasoning models write most of their KV cache while decoding long chains of thought (CoT), so the cache has to be compressed online under a fixed memory budget. Decode-time methods mostly decide which tokens to evict. We ask how a fixed byte budget should be split between the number of cached tokens and their precision. BreadthKV spends the bytes on more tokens at low precision, combining quantization with eviction, and picks the bit-width for each model and budget with a 60-problem end-to-end calibration, since offline attention error does not predict it reliably. On three reasoning models and four math and science benchmarks, it scores above eviction alone in 17 of 18 settings and produces shorter outputs. Much of what eviction loses comes from derailed runs, which keep reasoning until the length cap without reaching an answer. On Qwen3-8B at our tightest budget, eviction sends 91% of AIME samples to the cap and BreadthKV 40%. Under the same protocol, BreadthKV is statistically indistinguishable from a joint rate-distortion allocator (RDKV) that uses 27% more KV memory-time, and it outperforms our re-implementation of ThinKV.

cs.CL↗

Do Time-Series QA Systems Read the Time Series? Evidence Use and Reasoning Reliability

In recent years, time-series question answering (QA) systems have made significant progress. However, generating a correct answer does not show whether retaining the supplied numerical series improves task performance, nor whether the prediction is sensitive to changes in that input. While some systems provide rationales, answer accuracy also does not show whether their numerical claims are grounded in the supplied series or whether the stated inference is valid. In this work, we focus on evaluating four time-series QA systems: TimeOmni-1, ChatTS, TimeOmni-VL, and Time-MQA. First, for three systems with released evaluation data, we reproduce their reported results and compare the performance of the systems with their backbones. Then, we introduce a benchmark named COMMON-TSQA, which collects public evaluation datasets from existing time-series benchmarks and unifies their sample representation, task definitions, and answer schemas, while evaluating each system through its own interface under common evaluation criteria. The evaluation uses the original condition and six interventions while keeping the question and target fixed. Our analysis shows that aggregate performance alone can obscure how systems use numerical evidence. Similar task-level scores can arise despite substantial changes in individual predictions. Some interventions induce simple fallback behavior rather than preserved task ability. We also evaluate rationales for factual grounding, inference validity, and consistency with the final answer. We find that rationales often contain time-series claims unsupported by the input. Moreover, the rationale audit shows that agreement between a rationale and its final answer can coexist with incorrect numerical descriptions or invalid intermediate inferences.

cs.AI↗

Target-Generated Dirichlet Problems in State-Constrained Stochastic Control

We study state-constrained stochastic control problems with a scalar surplus $R \geq 0$. Controls for which the drift and volatility of $R$ vanish at zero define a lower-dimensional Hamilton-Jacobi-Bellman equation. Starting from control-wise state-constraint viscosity inequalities, we prove that the upper and lower limits of the interior value are a subsolution and a supersolution of this boundary equation. Terminal compatibility and comparison identify their common limit. We require one-sided bounds on positive surplus drift and generator growth, together with domination by boundary generators. For compact or coercive controls, these follow from a local lower bound and continuity of the boundary control set. For a smooth stochastic target value $w$, the transformation $R = Y - w(t,X)$ flattens the viable epigraph, and, under the stated hypotheses, the boundary control value supplies the Dirichlet datum. Applications include set-valued boundary controls, unbounded drift with quadratic costs, and redundant hedging instruments with a nonconstant target boundary.

math.OC↗

Square-Root Regret for Adversarial Multiplayer Bandits without Collision Information or Shared Randomness

We study adversarial multiplayer bandits with $K$ arms and $2\le m<K$ labeled players, without collision information, shared randomness, or an external communication channel. We design a constructive communication and synchronization protocol with a Monte Carlo public constructor. With probability at least $1-CN^{-32}$ over preprocessing, where $N=2Km(T+1)$, its fixed published output satisfies \[ R_T\le C K^{5/2}\sqrt T\log^2(2Km(T+1)) \] simultaneously for every oblivious reward sequence chosen after preprocessing. Here $R_T$ is expected regret over the players' private execution randomness. Positive reward observations establish a common learning schedule and synchronize players before learning begins. The cost of delayed communication is charged to the support of positive rewards, ensuring that periods with little useful feedback incur only limited regret. A slow--fast learning procedure then maintains valid reward estimates while assignments and scores are exchanged.

cs.LG↗

Toward AI Trustworthiness: Finding Analytically Proven Forward-Invariant Sets for AI-Controlled Systems

Neural-network (NN) controllers are increasingly used in nonlinear control systems, but their highly nonlinear behavior makes them difficult to explain and verify, raising trustworthiness concerns in safety- and mission-critical applications. A key step toward certifiable trustworthiness is to find a Forward-Invariant Set (FIS): a state-space region such that any trajectory starting inside remains inside. If the FIS excludes unsafe states, safety can be guaranteed for initial states within it. Finding an analytically proven FIS for a given AI-controlled system with a fixed controller is difficult. We propose a framework that uses an Invertible Neural Network (INN) to transform the original state space into a latent space where a regular-shaped FIS is more likely to exist. We train the INN so that a preferred hyper-rectangular candidate becomes invariant in the latent space, then formally verify it. We prove that, whenever verification succeeds, both the latent-space candidate and its inverse-transformed counterpart in the original state space are analytically proven FISs. We evaluate the approach on 45 AI-controlled systems across three representative control testbeds. Our method finds certified FISs for all 45 systems, whereas an adapted state-of-the-art baseline finds none. It is also faster on 40 of the 45 systems, and the centers of the resulting FISs roughly match domain-expert preferences.

cs.AI↗

Errors of LLM-Assisted Literature Retrieval in Environmental Science: A Comparison Study of Abstract versus Full-text Based Prompts

Large language models (LLMs) are increasingly used for literature search and synthesis. However, it is unclear whether they retrieve accurate bibliographic information in environmental science. Therefore, we quantitatively compared the errors of widely used LLM platforms in retrieving references related to original articles from five leading environmental science journals (Energy and Environmental Science, Nature Sustainability, Nature Climate Change, Lancet Planetary Health, and Environmental Science and Technology) published in 2024 to 2025. Claude, ChatGPT, Grok, DeepSeek, Perplexity, and Gemini were used as the LLM platforms. LLMs retrieved 10 references for each of the 50 randomly selected original article using either the article's abstract or its full-text as prompt. The retrieved references were subject to a multimetric score ratio combining validity of bibliographic data, Google Scholar link, digital object identifier, Scopus Electronic Identifier and relevance score (cited by or being the index paper), and the proportion of complete fabrication that failed all metrics. Abstract-only prompt yielded significantly higher accuracy than full-text one. This advantage was confirmed in multilevel mixed-effect multivariable regression after adjusting for journal, platform, and output order. Source journal and the position of a reference within the output list were also independently associated with retrieval accuracy, with lower-listed references associated with lower accuracy. These findings suggest that LLM assisted literature retrieval in environmental science remains moderately accurate and overall inconsistent, varying significantly by platform, journal, prompt type, and output position. Abstract-based prompting, as task-aligned information compression, may outperform full-text one in literature retrieval. Caution should be used when generalizing our findings.

cs.DL↗

AudioGAR: Bridging Reconstruction and Generation in Latent Audio Generative Models

Latent audio generative models are typically trained in two stages: an audio codec is learned first, followed by a latent generative model. This decomposition leads to a decoder train-generation mismatch: the codec decoder is trained on encoder-induced latents but deployed on generator-produced latents at inference time. Across diverse datasets and latent generative models, we observe clear reconstruction-generation gaps under both FD and FAD, showing that strong reconstruction quality does not necessarily translate into strong end-to-end generation quality. A natural remedy is to adapt the decoder on generation-produced latents, but generated latents lack correspondence with source audio and therefore cannot directly provide the paired supervision used for decoder fine-tuning. We introduce \textbf{AudioGAR}, which constructs intermediate latents by perturbing encoder latents and denoising them through the frozen latent diffusion model. These latents form a trajectory from reconstruction toward generation, with lower-noise latents retaining source correspondence and supporting paired decoder fine-tuning. We fine-tune only the codec decoder on these latents, while keeping the codec encoder and latent generative model frozen. When applied to AudioX, AudioGAR substantially improves generative performance. It requires only 1.5\% of the original training audio hours and 0.26\% of the original training cost.

cs.SD↗

Benchmarking Generative Trajectory Models for Active-Inference Control

Learning from trajectory demonstrations offers a route to active-inference control of complex systems whose dynamics are difficult to model explicitly. We introduce generative active-inference control (GenAIF), in which one generative trajectory model learns from demonstrations and measured action interventions to supply a goal-conditioned policy distribution and a state-to-observation likelihood mapping. From this control design, we derive three model requirements: (i) useful action proposals, (ii) accurate prediction under imposed actions, and (iii) probabilistic observation evidence for belief updating and expected information gain. We benchmark diffusion, autoregressive Transformers, conditional variational autoencoders (CVAEs), and flow matching in a MuJoCo manipulation task with multiple physical conditions. Diffusion delivers the strongest control across the tested dynamics, while CVAE combines comparable short-horizon prediction with much faster inference. Correct conditioning is decisive, and trajectory reuse offers further computational savings. With the same frozen models, a hidden-dynamics experiment demonstrates prompt belief adaptation after an unannounced tilt change; subsequent instability identifies sustained inference as a remaining challenge. These findings support the use of shared generative trajectory models to connect action proposal, controlled prediction, and observation evidence within GenAIF.

cs.RO↗

Speaker Tracking: Segment-online Multi-talker Organization with a Varying Number of Speakers

Speaker tracking is the task of separating and following multiple speakers over time. It must address overlapped speech, speech onset and offset, talker identity, and time-varying speaker count. We propose a segment-online, modular framework for single- and multi-channel speaker tracking. The proposed system first performs speaker separation in each segment and computes speech activity via voice activity detection (VAD). To generate speaker tracks over time, we introduce a two-stage sequential organization strategy: Overlap-based stitching for continuous grouping and memory-based speaker verification for discontinuous grouping. For speaker separation, we employ complex spectral mapping to estimate the real and imaginary spectrograms of underlying speakers. The proposed system achieves state-of-the-art segment-online tracking performance on the LibriCSS and AMI datasets. Our framework significantly reduces diarization error rate (DER) and concatenated minimum-permutation word error rate (cpWER) compared to other methods.

eess.AS↗

6G as It Is Actually Being Built: Insights from Early 3GPP Standardization

As mobile communications cross the threshold from 5G to 6G, the 3GPP has entered a decisive phase. Following the Technical Specification Group (TSG) plenary meetings of June 2026 and the associated 6G workshop in Singapore, the Release-20 study phase is now well underway across all three TSGs: Radio Access Network (RAN), Service and System Aspects (SA), and Core Network and Terminals (CT). Concurrently, the timeline for the first normative 6G specifications in Release 21 has been established. Rather than offering another technology vision, or a chronological account of committee progress, this article asks what the first body of agreed 6G material tells us. We review service requirements, system architecture, the emerging radio interface, native AI, ISAC, non-terrestrial networks (NTN), and core-network protocols, tracing how each moves from study item toward specification. Read across the three TSGs, this material also supports seven cross-cutting lessons: AI has moved from a feature to a primitive spanning every stage, non-terrestrial access has crossed from overlay to substrate, and connectivity is now one objective among three alongside computing and sensing; the industry favors deployability over redesign, with spectrum access and site reuse, rather than waveform innovation, appearing to be the tighter constraint; and standardization targets AI workflows rather than AI models, with testability across vendors recurring as a decisive criterion for which techniques survive. Because the material spans very different stages of maturity, every item cited is tagged with its status at a single snapshot date (approved, draft, individually proposed, or vendor-demonstrated), and the lessons are offered as the authors' synthesis rather than as 3GPP positions. We close with the road toward trials and commercial launch around 2030, and with the open questions these lessons help prioritize.

cs.NI↗

Virtual model control for compliant reaching under uncertainties

Virtual Model Control (VMC) is an approach to design a controller for force-controlled robots in complex uncertain environments. While this method was primarily investigated for legged robot locomotion in the past, it can be more generally applicable to other types of robotic systems. This paper investigates the VMC framework for reaching tasks in a force-controlled robotic arm. We propose six different approaches to designing virtual models in order to achieve reaching tasks in environments with obstacles and uncertainties. A force-controlled 8 degree-of-freedom humanoid robot was used to validate the proposed approach in the real world. We conducted three experiments to test the performance of VMC controllers in terms of predictability, sensitivity to external force, and adaptability against known and unknown obstacles. Experimental analyses show that, even though the proposed approach needs to sacrifice accuracy and trajectory optimality, it enables us to design complex reaching motions under uncertainties, in an intuitive and extendable manner.

cs.RO↗

Substring Edit Correcting Codes and Optimal Single Burst-Deletion Correcting Codes

A $k$-substring edit in a sequence first deletes a substring of length at most $k$, and then inserts a sequence of length at most $k$ at the same position. A code that can correct a $k$-substring edit is called a $k$-substring edit code. In this paper, we develop a new localization method and use it to construct a $q$-ary $k$-substring edit correcting code with $\log n+8\log\log n+o(\log\log n)$ bits of redundancy for any fixed $q\ge2$ and $k\ge1$, where $n$ is the code length. For the binary alphabet, this improves upon the redundancy $\log n+16k\log\log n+o(\log\log n)$ obtained by Li \emph{et al}. When the deleted substring and the inserted sequence have different lengths, we further construct codes with redundancy $\log n+O_{q,k}(1)$, which is optimal up to an additive constant. As corollaries, for all fixed $q\ge2$ and $1\le t\le T$, we obtain $q$-ary $(\le t)$-burst-deletion correcting codes and $(t,T)$-localized deletion correcting codes with redundancies $\log n+O_{q,t}(1)$ and $\log n+O_{q,T}(1)$, respectively. To the best of our knowledge, these are the first constructions attaining optimal redundancy up to an additive constant for these two deletion models over the full range of fixed parameters. For $(\le t)$-burst-deletion correction with $t\ge2$, such redundancy had previously been achieved only for $q=t=2$ by Levenshtein in 1967.

cs.IT↗

From Token-Max to Outcome-Max: How You Use AI Determines Its Productivity

Generative artificial intelligence (AI) models can perform increasingly complex tasks, yet greater AI usage does not necessarily translate into proportional productivity gains. We identify token-max as one source of this inefficiency: when token consumption is treated as productive effort, agents are encouraged to over-exert and expend computation beyond what is necessary. We instead propose outcome-max, which rewards independently verified task completion per unit cost and induces a principled stopping rule. Then, to study these objectives, we develop a three-level simulation framework spanning immediate interaction, long-run behavioral adaptation, and organizational collaboration. Across all three levels, outcome-max improves the efficiency of AI-assisted production while largely preserving verified task performance. To further align these incentives with outcome-max, we introduce OutcomeShare, an incentive mechanism. Theory and simulation show that OutcomeShare can induce participation while generating shared gains for employees, firms, and LLM providers. Together, our results suggest that AI productivity not only depends on model capability, but also on how to construct the objectives governing AI use.

cs.AI↗

Large time behavior of Lévy processes and their nonlocal Schrödinger semigroups

We study the large time asymptotics of the Feynman-Kac semigroups of the symmetric Lévy process on unbounded open sets. Our main result proves the exact exponential asymptotic decay rate for the survival probability given in terms of the bottom of the spectrum of the associated nonlocal Schrödinger operator. We also prove quantitative upper bounds with an explicit polynomial correction. The proof is probabilistic and is done by decomposing the Lévy process into a finite-range jump process collecting the small jumps and an independent compound Poisson process describing the large jumps. Our approach gives a direct link between the spectral properties of nonlocal Schrödinger operators and pointwise decay of survival probabilities, which was previously unknown for jump processes.

math.PR↗

Planetary Geospatial Foundation Models: A New Paradigm for Global Public Health

The efficacy of traditional disease prediction is limited by spatial gaps and temporal lags, which impact the timing and targets of resource deployments. Outbreaks escalate undetected, chronic disease burdens are quantified years later, and at-risk populations in data-sparse regions remain unaddressed. Planetary geospatial foundation models complement existing epidemiological workflows to provide operational improvements, encoding multimodal search, mobility, and environmental signals into generalizable place representations. As illustrations of this complementarity, we present independent global health case studies of Google Earth AI's Population Dynamics Foundation Model (PDFM) -- a foundation model for geospatial inference -- across four domains (vaccine-preventable, communicable, noncommunicable, maternal mental health), five tasks (spatial extrapolation, interpolation/nowcasting, probabilistic forecasting, prospective forecasting, risk stratification), and four countries (USA, Canada, Mexico, and the Democratic Republic of the Congo). Across these case studies, PDFM addresses critical surveillance gaps across domains: improving US-Canada border MMR vaccination coverage predictions by capturing cross-border behavioral spillovers domestic models miss; nowcasting cardiovascular disease to accelerate data availability; enhancing short-term municipal Mexican dengue forecasts for timely outbreak vector control; improving forecasts of cholera hotspots; and adding a transferable signal to individual-level postpartum-depression risk prediction in US states the model had never seen, while not replacing individual socioeconomic data or closing demographic screening gaps. Together, these results showcase capabilities of geospatial foundation models for public health surveillance.

cs.LG↗

Voltic: Distinguishing Volatility from Stochasticity in Recurrent Memory

Recurrent sequence models must decide how strongly to overwrite their memory at each token. Read as Bayesian filtering, this write is the gain of a Kalman update, set by uncertainty from two sources that pull it in opposite directions: volatility, how quickly the underlying associations change, and stochasticity, how noisy each observation of them is. First, we show that the update of gated delta-rule memories is the form this filter takes under isotropic uncertainty. Next, we introduce Voltic, a recurrent memory that keeps the covariance anisotropic and makes both noise variances input-dependent, so the write is vector-valued and carries uncertainty accumulated over the sequence. A dense covariance would have to be propagated token by token, ruling out the parallel training these models depend on. We therefore give two assumed-density approximations, diagonal and quasi-diagonal, both of which leave the memory update in delta-rule form and reuse its chunked kernels. On controlled recall tasks in which associations change and observations are corrupted, Voltic leads all baselines. On the task combining volatility and stochasticity, its margin over the strongest baseline is larger at both extrapolation sizes than at the training sizes. In 45M-parameter language models it leads an eight-task reasoning average and achieves higher retrieval accuracy beyond the training context length than gated baselines, at throughput close to those baselines. Deriving the write from an uncertainty recursion therefore makes memory more responsive to change.

cs.LG↗