Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 307 records · Page 17Linked to original sources

Toward the Goldilocks Blind Compression of Quantum States

Quantum autoencoders (QAEs) are learning architectures that compress quantum data into a low-dimensional latent state while preserving the information needed for reconstruction. We study blind single-copy compression of quantum states through a $k$-qubit bottleneck and investigate the minimal circuit width required to attain the information-theoretic optimum under average infidelity. Between the conventional architecture, which is narrow but nonuniversal, and fully general \emph{completely positive and trace preserving} (CPTP) realizations, which are universal but overparameterized, we identify a balanced regime. We prove that for every distribution of pure $n$-qubit states, there exists a QAE with $k$ encoder ancillas and $n$ decoder ancillas that achieves the optimal fidelity over all CPTP encoder--decoder pairs. The encoder-side statement is sharp in that we construct source families for which every optimal scheme necessarily uses at least $k$ encoder ancillas, thereby determining the universal encoder threshold exactly. On the decoder side, we show that isometric decoders are optimal for several analytically tractable source families, but we also exhibit an explicit counterexample demonstrating that decoder isometry is not universally sufficient. Nevertheless, numerical experiments indicate that the performance gap is practically negligible.

quant-ph↗

VIDA: A Dataset for Visually Dependent Ambiguity in Multimodal Machine Translation

Ambiguity resolution is a key challenge in multimodal machine translation (MMT), where models must genuinely leverage visual input to map an ambiguous expression to its intended meaning. Although prior work has proposed disambiguation-oriented benchmarks probing the role of vision, we observe that existing benchmarks remain limited by task-format mismatch, narrow ambiguity coverage, or insufficient visual-dependency validation. Moreover, existing ambiguity evaluations are not well suited to diverse ambiguity types in open-ended translation. To address these limitations, we present VIDA (Visually-Dependent Ambiguity), a dataset of 2,500 carefully curated instances in which resolving an annotated source span requires visual evidence. We further propose Disambiguation-Centric Metrics that use an LLM-as-a-judge classifier to verify whether annotated ambiguous expressions are resolved correctly at the span level. Evaluations with stronger recent LVLMs show that visual disambiguation remains challenging. Using chain-of-thought supervised fine-tuning as a diagnostic setting, we observe stronger out-of-distribution disambiguation than with SFT, with robust gains on collective-noun ambiguities and model-dependent gains on sentence-level ambiguities.

cs.CL↗

Scalable generative modeling of non-Gaussian spatio-temporal fields via autoregressive Gaussian processes

Generative modeling of spatio-temporal fields is crucial for a variety of applications, including stochastic weather generators and climate-model surrogates. However, many such fields exhibit complex dependence structures that vary across space and time and are nonlinear, resulting in nonstationary and non-Gaussian joint distributions. Our approach represents the joint density of a spatio-temporal field as a product of univariate conditional distributions and models these conditionals using Gaussian processes within an autoregressive transport-map construction. This prior distribution provides regularization, making our method suitable for a small number of training samples. Data-dependent sparsity in the conditioning sets ensures scalability to high-dimensional distributions. We also propose a variant of the method designed to sample or predict forward in time from a given incomplete space-time trajectory. We demonstrate the accuracy and scalability of our approach on non-Gaussian climate-model output with tens of millions of data points.

stat.ME↗

Uncertainty Quantification in Forecast Comparisons

Skill scores, which measure the relative improvement of a forecasting method over a benchmark via consistent scoring functions and proper scoring rules, are a standard tool in forecast evaluation, yet their sampling uncertainty is rarely rigorously quantified. With modern forecasting applications being increasingly multivariate and involving evaluations across multiple horizons, variables, spatial locations, and forecasting methods, standard tools like the pairwise Diebold-Mariano forecast accuracy test or pointwise confidence intervals fail to account for the multiple comparison problem, leading to inflated Type I error rates and invalid joint inference. To address the lack of a coherent, statistically rigorous framework for quantifying uncertainty across these multi-dimensional evaluation problems, we introduce simultaneous confidence bands for expected scores and skill scores. Our framework provides a versatile tool for joint inference that is applicable to any forecast type from mean and quantile to full distributional forecasts. We develop a bootstrap implementation and show that our bands are valid under multivariate extensions of the classical Diebold-Mariano assumptions. We demonstrate the practical utility of the approach in two case studies by quantifying the benefits of time-varying parameter models for macroeconomic forecasting, and by comparing data-driven and physics-based models in probabilistic weather forecasting.

stat.ME↗

Real-Time Estimation of High-Resolution Flow Fields and Reduced-Order Coordinates from Event-Based Imaging Velocimetry

We propose a data-driven framework to estimate high-resolution (HR) velocity fields and reduced-order flow coordinates from real-time Event-Based Imaging Velocimetry (rt-EBIV). Fast event analysis first provides low-resolution (LR) velocity snapshots on a coarse grid. Offline, paired LR/HR fields are used to identify the LR-to-HR mapping and a linear dynamical model in a POD-based latent space. Online, each LR snapshot is projected onto the LR basis, the corresponding HR coordinates are estimated and temporally regularized, and the HR field is reconstructed from the retained POD modes. Three estimators are compared: a direct Kalman filter (KF), a linear stochastic estimator followed by Kalman filtering (LSE), and a variance-rescaled variant (LSE+VR). The method is tested on two turbulent flows acquired with pulsed EBIV: a submerged water jet and a channel flow over a square rib. All estimators outperform direct cubic interpolation of the LR fields, yielding more consistent HR reconstructions of instantaneous flow states, turbulent kinetic energy, spectra, reduced-order dynamics, and temporal coherence. LSE gives the lowest overall reconstruction error, while LSE+VR achieves similar errors with improved recovery of fluctuation energy and higher-order content. The direct KF is the most computationally efficient and provides the closest agreement with the HR reference in spectral analyses. Since most of the cost is associated with full-field HR reconstruction, the latent-coordinate estimation is negligible compared with LR processing. The framework allows deliberately coarse rt-EBIV processing to be combined with reduced-order refinement, extending real-time operation toward higher update rates while preserving richer and dynamically consistent HR flow representations for diagnostics and future observer-based flow-control applications.

physics.flu-dyn↗

Exact entanglement trade-offs in qutrit and composite-dimensional stabilizer states

When absolute maximal entanglement is unavailable in a state family, different objectives can select different distributions of residual entanglement. We determine exact trade-offs for qutrit stabilizer states and binary--ternary product states. For eight qutrits, evaluation of the complete published classification gives three nondominated balanced-cut entanglement profiles, characterized by the probabilities of obtaining three or four maximally entangled qutrit pairs. We relate the squared entropy-deficit objective to code weight enumerators and use a uniformity-forcing bound to prove that its eleven-qutrit minimum is six. The unique minimizing Clifford class is an already classified code, whose entropy profile also attains a sharp bound on average reference--erasure mutual information among four-uniform eleven-qutrit stabilizer states. In local dimension six, explicit sector alignments attain 56 maximally mixed four-party reductions, the maximum for arbitrary binary--ternary product states. Within the stabilizer product family, the exact cost--count Pareto boundary consists of $(44,48)$ and $(48,56)$. These results identify when aggregate optimization determines operational entanglement properties and when it does not. Published classifications supply coverage; analytical bounds, explicit witnesses, and independent arithmetic checks establish the stated consequences.

quant-ph↗

LensVLM: Selective Context Expansion for Compressed Visual Representation of Text

Vision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into long token sequences. Since VLM image encoders map fixed-size images to a fixed number of visual tokens, varying rendering resolution provides a fine-grained compression knob. However, accuracy deteriorates quickly as compression increases: characters shrink below the vision encoder's effective resolution, making them indistinguishable. To address this, we propose LensVLM, an inference framework and post-training recipe that enables VLMs to scan compressed images, then selectively expand only the relevant images to their uncompressed form via learned tools. Building on Qwen3.5-9B-Base, LensVLM maintains accuracy comparable to the full-text upper bound at 4.3$\times$ effective compression and outperforms retrieval-based, text- and visual-compression baselines up to 10.1$\times$ effective compression across seven text QA benchmarks. LensVLM also generalizes to multimodal document and code understanding tasks, with the accuracy gain over baselines growing as compression increases. Our analysis validates this approach: training makes visual compression robust to rendering choices, and as compression grows the model increasingly relies on expanded content rather than unreliable visual reading. The analysis also yields practical tool-choice guidance: text expansion is preferable for rendered text, while high-resolution image expansion suits native documents whose layout cues carry task-relevant information.

cs.CV↗

On hypercyclic and (common) upper frequently hypercyclic subspaces

Let $B$ be a unilateral weighted backward shift on $\ell_p$, $1 \leq p < \infty$, that admits a $\mathscr{U}$-frequently hypercyclic subspace. We prove that $B$ admits such a subspace free of frequently hypercyclic vectors. The proof technique we develop also allows us to prove that $B$ admits a hypercyclic subspace free of $\mathscr{U}$-frequently hypercyclic vectors, and to solve a question posed by Bès and Menet in 2015 on the existence of common $\mathscr{U}$-frequently hypercyclic subspaces.

math.FA↗

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models rests on variance reduction, which requires both a reliable baseline and high prompt diversity within each training batch. This is especially difficult in multi-domain training for general reasoning models, where prompts from different tasks induce highly diverse gradient signals. Existing approaches fall short in different ways: GRPO estimates its baseline as the group mean over rollouts from the same prompt, so an accurate baseline leaves fewer distinct prompts in the batch, while PPO avoids this trade-off by training a policy scale critic, roughly doubling the cost of training. We introduce POISE (Policy Optimization with Internal State Value Estimation), a reinforcement learning algorithm that turns the model's internal states into a value model. A lightweight probe reads the signals already computed during the forward pass to predict the baseline, and is trained online alongside the policy. To preserve gradient unbiasedness, we introduce a cross-rollout construction that predicts each rollout's value from an independent rollout's internal states. On Qwen3-4B and OLMo3-7B-Instruct-DPO across a six-domain verifiable-reward corpus, POISE outperforms other RLVR baselines while achieving more stable training. Moreover, the probe matches a separate LLM-scale value model, generalizes to various tasks, and remains accurate as the policy scales. By leveraging the model's internal representations, POISE enables stable policy optimization.

cs.LG↗

SkillLens: Adaptive Multi-Granularity Skill Reuse for Cost-Efficient LLM Agents

Skill libraries have become a practical way for LLM agents to reuse procedural experience across tasks. However, existing systems typically treat skills as flat, single-resolution prompt blocks. This creates a tension between relevance and cost: injecting coarse skills can introduce irrelevant or misleading context, while rewriting entire skills is expensive and often unnecessary. We propose SkillLens, a hierarchical skill-evolution framework that organizes skills into a four-layer graph of policies, strategies, procedures, and primitives, and retrieves them at mixed granularity. Given a task, SkillLens first retrieves semantically relevant skill seeds, expands them through degree-corrected random walk over the skill graph, and then uses a verifier to decide whether each visited unit should be accepted, decomposed, rewritten, or skipped. This enables the agent to reuse compatible subskills directly while adapting only locally mismatched components. To improve the system over time, SkillLens further refines multi-granularity skills and verifier in order to improve its routing decisions. We provide theoretical analysis showing that mixed-granularity adaptation incurs sublinear cost under sparse mismatch assumptions and that the evolutionary update rule monotonically improves the validation objective until a local optimum. Across MuLocbench and ALFWorld, SkillLens consistently improves over strong skill-based baselines, achieving up to a 6.31 percentage-point Acc@1 gain for bug localization and raising agent success rate from 45.00% to 51.31%.

cs.AI↗

dFlowGRPO: Rate-Aware Policy Optimization for Discrete Flow Models

Discrete flow models (DFMs) are a class of flexible generative models for generating discrete data, and diffusion large language models (dLLMs) can be viewed as a special case with a specific choice of a mixture path and a masked source distribution. While several recent works have explored reinforcement learning for dLLMs, its application to more general discrete flow models remains underexplored. In this work, we present discrete Flow-GRPO (dFlowGRPO), a unified reinforcement learning framework for discrete flow models that supports a broad family of probability paths and non-masked source distributions. We derive the full trajectory probability for DFMs and formulate the denoising process as a Markov decision process, enabling dFlowGRPO to incorporate information from both the associated conditional transition rates and the posterior model during reinforcement learning. We apply dFlowGRPO to FUDOKI, a recent multimodal discrete flow model, and evaluate it on both image generation and multimodal understanding tasks. Empirical results show that dFlowGRPO outperforms existing GRPO-type methods for dLLMs on text-to-image generation tasks and achieves performance competitive with continuous flow-based models trained using Flow-GRPO, while also demonstrating strong capabilities on understanding tasks.

cs.LG↗

DynGhost: Temporally-Modelled Transformer for Dynamic Ghost Imagings

Ghost imaging reconstructs spatial information from a single-pixel bucket detector by correlating structured illumination patterns with scalar intensity measurements. While deep learning approaches have achieved promising results on static scenes, two critical limitations remain unaddressed: existing architectures fail to exploit temporal coherence across frames, leaving dynamic ghost imaging largely unsolved, and they assume additive Gaussian noise models that do not reflect the true Poissonian statistics of real single-photon hardware. We present DynGhost (Dynamic Ghost Imaging Transformer), a transformer architecture that addresses both limitations through alternating spatial and temporal attention blocks. Our quantum-aware training framework, based on physically accurate detector simulations (SNSPDs, SPADs, SiPMs) and Anscombe variance-stabilizing normalization, resolves the distribution shift that causes classical models to fail under realistic hardware constraints. Experiments across multiple benchmarks demonstrate that DynGhost outperforms both traditional reconstruction methods and existing deep learning architectures, with particular gains in dynamic and photon-starved settings.

cs.CV↗

Principled Design of Diffusion-based Optimizers for Inverse Problems

Score-based diffusion models achieve state-of-the-art performance for inverse problems, but their practical deployment is hindered by long inference times and cumbersome hyperparameter tuning. While pretrained diffusion models can be reused across tasks without retraining, inference-time hyperparameters such as the noise schedule and posterior sampling weights typically require ad-hoc adjustment for each problem setup. We propose principled reparameterizations that induce invariances, allowing the same hyperparameters to be reused across multiple problems without re-tuning. In addition, building on the RED-diff framework, which reformulates posterior sampling as an optimization problem, we further develop the OptDiff pipeline. OptDiff provides a simplified tuning framework that facilitates the integration of convex optimization tools to accelerate inference. Experiments on image reconstruction, deblurring, and super-resolution show substantial speedups and improved image quality.

cs.CV↗

On the Interpretability of Whisper Encodings Using Sparse Autoencoders

While deep transformer-based models have advanced rapidly, their internal mechanisms remain largely a mystery. Recent work has prioritized understanding text-based transformer models, leaving ASR systems largely unexplored. In order to address this gap, we examine the internal representations of Whisper's encoder using a sparse autoencoder. We find diverse monosemantic features across linguistic and non-linguistic boundaries, spanning a hierarchy from phonetic to semantic representations, and conduct a causal feature-steering campaign across this hierarchy, including cross-lingual steering. We further find that steering is more reliable for higher-level features than lower-level ones, an asymmetry that may reflect redundant encoding of lower-level information. Altogether, this work demonstrates that Whisper's encoder represents a surprisingly rich hierarchy of linguistic information that extends well beyond what is strictly necessary for transcription.

cs.CL↗

Beyond Localization: A Comprehensive Benchmark of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. We study this challenge as Perspective-Conditioned Spatial Reasoning (PCSR) in 360 degree omnidirectional images, where broad scene coverage reduces ambiguity from partial observations without eliminating the need for viewpoint-dependent inference. To assess this capability, we introduce PCSR-Bench, a diagnostic benchmark of 84,373 QA pairs from 2,600 omnidirectional images across 26 indoor environments, organized into eight tasks under three cognitive groups--- Perception, Spatial, and advanced PCSR. We evaluate 14 representative MLLMs and observe a substantial perception--reasoning gap: accuracy reaches 57.59% on Limited Field-of-View Reasoning (T7) but drops to 13.49%, 7.13%, and 0.64% on Relative Direction (T2), Egocentric Rotation (T4), and open-ended Compositional Directional Chains (T3), respectively. To probe the plasticity of this gap, we conduct an RL-based diagnostic study on a 7B-scale model. Reward shaping improves a matched 7B baseline from 31.10% to 60.06% under a controlled setting, suggesting that PCSR exhibits partial plasticity rather than being fully immutable. Still, these gains are task-selective, sensitive to reward design, and partially dependent on the evaluation protocol. These results position PCSR as a key bottleneck in current MLLMs and highlight meaningful yet bounded room for recovery under targeted optimization. Details and access are available at https://github.com/Caleb-ychen/PCSR-Benchmark.

cs.CV↗

ReForge: Refining Merged Models with Anchor-Regularized Regression

Model merging aims to combine multiple task-specific expert models into a single model without joint retraining, offering a practical alternative to multi-task learning when data access or computational budget is limited. Existing model merging methods rarely exploit strong merged models as priors for further improvement. To address this limitation, we propose ReForge, a bilevel optimization framework that formulates module-wise refinement as Bayesian linear regression with an anchor-centered prior. The inner level yields a closed-form MAP estimate from unlabeled calibration activations. The outer level uses Bayesian optimization to jointly select heterogeneous regularization strengths and assembly scales using held-out validation data. Furthermore, we develop a data-free variant of ReForge that replaces activation statistics with task-vector Grams, eliminating the need for calibration examples. Across extensive benchmarks, including up to 20-task merging in vision and 5-task merging in language, ReForge consistently outperforms all evaluated plug-and-play anchor baselines (e.g., TA, WUDI-Merging, and TSV). On 20-task ViT-B/32, ReForge improves the strongest evaluated baseline, ISO-CTS, from 77.6% to 82.8% in the data-assisted setting and to 81.5% in the data-free setting. On eight-task ViT-L/14, the data-assisted variant achieves 95.1% mean accuracy, compared with 95.8% for the individual task experts. Our source code will be released soon.

cs.LG↗

Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models

Diffusion Language Models(DLMs) provide a promising alternative to autoregressive language models through iterative denoising and bidirectional generation. However, their iterative generation process introduces distinct safety vulnerabilities because harmful content can emerge at arbitrary positions and persist across subsequent denoising steps. Existing defenses rely on fixed interventions or aggressive remasking, which limits adaptive control over denoising trajectories and can degrade generation quality. We propose an inference-time defense framework that combines adaptive safety steering with safety-aware remasking. Our method uses a gating direction to continuously adjust steering strength from the current denoising state and applies a steering direction to masked positions to guide subsequent predictions toward safer trajectories. Our method further employs a lightweight response detector after the first generation block to identify unsafe trajectories at an early stage. The detector triggers targeted remasking over generated content and part of the conditioning prompt, and the model regenerates the selected positions under adaptive safety steering. This design combines continuous trajectory control with explicit correction of unsafe content while requiring no modification of model parameters. Experiments on LLaDA and Dream demonstrate that our method improves robustness against diverse jailbreak attacks while preserving benign generation quality and general model capability. Our code is available at https://anonymous.4open.science/r/DLM_Steering-C32B/.

cs.CL↗

Minimum energy paths for defect annihilation in finite dodecagonal quasicrystal clusters

Defect repair in quasicrystals involves local structural rearrangements on a complex potential energy landscape. Here, we combine the string method with the spring pair method to compute a minimum energy path connecting a locally defective finite dodecagonal quasicrystal cluster to a locally repaired reference configuration in the Lennard-Jones--Gauss particle model. The computed path contains seven local minima connected through six transition states. A detailed analysis of one defective region resolves three successive stages: a phason flip, aggregation of neighboring shield-like defects, and decomposition of the aggregated motif. The phason flip restores the local outer-ring structure and relocates a shield-like defect into the interior. Subsequent collective rearrangements remove the defective motifs and yield a net decrease in the local interaction-energy measure together with an increase in local bond-orientational order. These results provide a microscopic description of a specific defect-repair pathway on the potential energy surface of a finite dodecagonal quasicrystal cluster.

cond-mat.mtrl-sci↗