Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

PageWeaver: KV-Guided Query Unions for Sparse Attention

Dynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work. Query unions share page loads and populate Tensor Core tiles; their cost depends on which queries are grouped together. We present PageWeaver, an execution design that uses selected-page affinity to assemble query groups while preserving each query's original support and complete output ownership. A bounded GPU search produces query IDs, and an ID-aware two-CTA kernel consumes them without materializing reordered Q tensors or cross-page partial outputs. A direct KV-page union implementation provides a complementary design study of nonlocal reuse and reduction cost. With FP8 KV throughout, the H200 Union8 implementation achieves a 1.70x geometric-mean complete-call speedup over the measured FlashInfer path on six captures. Online regrouping further lowers latency by 3.26-7.66% on five selected 64K-context captures. Whole-model prefill throughput is 7.88-14.36% above the tested native path; the incremental regrouping benefit is smaller, with observed median gains of 0.47-0.73% at 32K/64K and regressions at 8K. A B300 comparison identifies cases where preparation cost and a stronger native kernel remove the advantage. These results separate execution-group reuse from the complete cost of exploiting it online.

cs.DC↗

The invariant subspace problem and Rosenblum operators II

Let $\mathcal{H}$ be the separable, infinite-dimensional complex Hilbert space. In the first paper of this series, we introduced, by means of the Rosenblum operators, the shift representation operators $K_x=\sum_{n=0}^{\infty}T^nx\otimes e_n$, which relate an operator $T$ with $r(T)<1$ to the unilateral shift $S$ and reveal its hidden analytic structure. In this paper we replace the pair $(H^2,S)$ by the Sobolev disk algebra $R(\mathbb{D})$ and the multiplication operator $M_z$, whose canonical left inverse $B$ plays the role of $S^*$, and study the two-parameter family $K_{x,y}=\sum_{n=0}^{\infty}T^nx\otimes B^{*n}y$. We prove that $T\in B(R(\mathbb{D}))$ with $r(T)<1$ is intransitive whenever there exist a nonzero vector $g$ and a nontrivial $M_z$-invariant manifold $\mathcal{M}$ such that $f(T)g$ has a zero in the closed unit disk for every $f\in\mathcal{M}$; the proof explicitly constructs a nontrivial invariant subspace of $T$. We further introduce the ideal property and show that every intransitive operator on $\mathcal{H}$ is unitarily equivalent to an operator on $R(\mathbb{D})$ with this property, so that the Invariant Subspace Problem for arbitrary operators reduces to a problem about a concrete function algebra. Applications include a characterization of intransitivity through the rationality of the generating function $\sum_{n=0}^{\infty}\langle T^nξ,η\rangle z^n$ and new results on Pearcy's problem. Our main application is Halmos's third problem, which has remained open for more than fifty years: we reduce it to invertible operators with conjugate bi-geometric form, and we prove that $T^{-1}$ is intransitive for every invertible $T$ admitting a nontrivial projection $P$ with $PTP=TP$ and $PT(I-P)$ of rank one. Finally, we exhibit tridiagonal operators that are intransitive and admit infinite decreasing chains of invariant subspaces.

math.FA↗

Pinpointing the cosmic web between massive galaxy clusters I. The incidence of HI and OVI as WHIM tracers

At redshift z<1, the $Λ$CDM paradigm predicts that $\sim$40-50\% of baryons reside in the warm-hot intergalactic medium (WHIM) within cosmic web filaments. Characterizing the physical state and distribution of this dominant baryonic component is crucial for a complete picture of baryon evolution in the low-redshift Universe. However, directly observing the WHIM has been challenging due to its low density and high temperature and ionization state. Absorption-line spectroscopy offers a promising approach, particularly through analyses of broad Ly-$α$ absorbers (BLAs; Doppler parameter b>40 km/s) and OVI absorbers. This work focuses on probing WHIM signatures by analyzing the incidence rates of these tracers in and around putative cosmic web filaments connecting massive galaxy clusters. We use far-ultraviolet (FUV) absorption spectroscopy from HST/COS along ten QSO sightlines at z<1, selected for their proximity in projected distance and velocity to cluster pairs. We focus on absorption features of WHIM tracers (BLAs and OVI) and cold-gas tracers (narrow Ly-$α$ absorbers, or NLAs; b<40 km/s), which are used as a control sample for colder gas. We measure excess incidence rates for total HI, NLAs, BLAs, and OVI within a rest-frame $Δv=\pm1000$ km/s and an impact parameter $Δd<3$ Mpc relative to the field expectation, and find that the excesses of BLAs and OVI are about twice those of NLAs. Additionally, the covering fractions of BLAs and OVI appear larger by factors of $\sim$1.7 and 1.8, respectively, than in a random field control sample, whereas the NLA covering fraction is consistent with the random expectation. The larger relative excess of WHIM tracers compared to cold-gas tracers near independent cluster pairs indicates the prevalence of warm-hot gas in these environments, representing a potential signature of WHIM detection in inter-cluster filaments.

astro-ph.CO↗

Sharp stability near sums of ground states for fractional Schrödinger equations

We study quantitative stability for the fractional Schrödinger equation $(-Δ)^s u+u-|u|^αu=0$ near finite sums of widely separated positive ground states. For every $n\ge1$, $0 0$, we estimate the $H^s$ distance to the family of such sums in terms of the $H^{-s}$ norm of the equation's residual. The optimal rate changes at $α=n/[2(n+2s)]$, with a logarithmic correction at the threshold. Above the threshold the rate is $t^{(n+2s)/(n+2s+1)}$; below it the exponent is $[(1+α)(n+2s)-n/2]/(n+2s+1)$. The proof combines uniform invertibility away from the translation modes with precise interaction estimates and a weighted correction of the approximate configuration. This correction resolves the translation interactions even when the nonlinearity has a small exponent. We also construct nonnegative configurations that attain these rates. The estimates quantify the effect of the algebraic decay of fractional ground states on stability.

math.AP↗

Dissecting Representation Structure in Vision Transformers: A Rigorous Architectural Study

Representation structure is crucial for understanding Vision Transformer (ViT) architectures and their generalization behavior. However, prior studies neither isolate nor analyze module-level features nor investigate how their interactions contribute to performance estimation. In this work, we conduct the first rigorous analysis of feature information across diverse architectural scales, empirically uncover the relationship between ViT representation and generalization behavior, and leverage these insights to guide efficient ViT design. Our contributions are fivefold: Across diverse architectural scales, 1) We identify feature collapse at initialization, which leads to redundancy, and propose a reduction scheme to mitigate this issue. 2) We quantify feature information using entropy and the minimum eigenvalue, demonstrating that these metrics serve as reliable indicators for generalization prediction. 3) We show that feature in the token space provides a more faithful representation than those in embedding space. 4) We discover an unexpected finding: features produced by linear submodules within ViT layers are critical for the prediction of generalization performance. 5) Our proposed proxy improves the correlation ranking by 18-48% over prior baselines and can effectively identify ViT architectures that achieve higher accuracy at lower or comparable computational cost.

cs.CV↗

Do Flatter Minima Drive Better Generalization? An Algorithmic Separation in Grokking

Flat loss landscapes have long been linked to better generalization in neural networks. However, its role as a causal mechanism for generalization is less established. Grokking provides an unique testbed to understand this distinction: models are prone to fit observed data using non-generalizing structure and remain in that regime for prolonged periods, transitioning to generalization only under particular training conditions. In this work, we study whether flat loss landscapes can act as a driving mechanism in this transition. While recent work has argued for flatness as a necessary geometric condition for this transition, we find that biasing training toward flatter solutions using sharpness-aware minimization (SAM) is insufficient to reliably induce this transition, despite producing flatter solutions. However, when SAM is paired with mechanisms that drive generalization such as weight decay, an interesting property emerges: SAM can accelerate the transition to generalizing solutions by up to 4x at the epoch-level. We theoretically untangle this relationship between SAM and weight decay using a minimal interpolating two-layer ReLU model with both memorizing and generalizing solutions. We show that even in this simple setup, flatness alone cannot distinguish a memorizing solution from a generalizing one, while weight decay favors generalizing solutions. However, under a local stability analysis, there exists a window where a memorizing interpolant is locally stable under gradient descent but unstable under SAM in the low-norm regime, which can explain SAM's ability to accelerate this transition. Overall, our results provide a more interpretable account of the role of flatness in driving generalization, especially in settings where models are vulnerable to minimizing loss through learning non-generalizing structure.

cs.LG↗

Chameleon Gravity with an Environmental Frequency Dependent on Density

Scalar field theories of modified gravity are strongly constrained by the absence of detectable long range fifth forces in laboratory and Solar System environments. The chameleon mechanism addresses this problem by making the equilibrium value and effective mass of the scalar field dependent on the ambient matter density. In this work, we investigate an effective phenomenological extension of the chameleon framework in which the scalar potential contains an additional quadratic environmental contribution characterized by a density dependent background frequency. This background frequency is distinct from the frequency of scalar perturbations and parametrizes unresolved environmental effects beyond the conventional matter coupling. For an inverse power law scalar potential, we derive the modified equilibrium configuration and identify a frequency dominated regime in which the environmental contribution controls the density dependence of the field. Linear perturbations about the environmental minimum obey a Klein Gordon dispersion relation with a density-dependent effective mass, and the absence of tachyonic instability requires a positive environmental coupling. We further derive analytical consistency conditions associated with frequency dominance, thin shell screening, fifth-force suppression, and the weak field approximation. Numerical illustrations for representative parameter choices demonstrate how the modified density scaling can enhance fifth-force suppression relative to the conventional chameleon case.

gr-qc↗

A Semiparametric Functional Generalized Linear Model for Multi-Population Data

Functional generalized linear models provide a flexible framework for relating a scalar response to functional and scalar predictors, but their conventional formulation requires specification of the response distribution and is typically developed for a single population. We propose a semiparametric functional generalized linear model for multi-population data that leaves the baseline response distributions unspecified and links them across populations through a density ratio model. The proposed framework accommodates both functional and scalar predictors while borrowing information across related populations. We develop a maximum empirical likelihood estimation procedure and use penalized B-spline approximations to estimate the functional coefficients. A cross-validation procedure based on a kernel-smoothed empirical likelihood density estimator is introduced to select the smoothing parameters. Simulation studies show that borrowing information through the density ratio structure can improve estimation efficiency relative to fitting each population separately, and that the proposed semiparametric method remains competitive when a parametric Gaussian model is correctly specified while providing substantial robustness when it is misspecified. We illustrate the method using county-level soybean yield data from Kansas, with daily temperature curves and irrigation levels as predictors.

stat.ME↗

On Ricci solitons whose level hypersurfaces have parallel second fundamental form

We study gradient Ricci solitons whose potential functions have level hypersurfaces with parallel second fundamental form. We show that such a soliton is locally a multiply warped product of a one-dimensional base and Einstein fibers, with potential function depending only on the base. Thus, locally, this condition characterizes the multiply warped structure that occurs in many classical constructions of Ricci solitons. We further prove that the same local structure follows for any gradient Ricci soliton if just one regular level hypersurface has parallel second fundamental form and parallel intrinsic Ricci tensor. As an application, we prove that a complete nonsteady gradient Ricci soliton with constant scalar curvature is rigid as soon as one regular level hypersurface has this property. In particular, the conjecture of Cao holds for gradient shrinking Ricci solitons with this property. We also prove that a complete gradient shrinking Ricci soliton whose regular level hypersurfaces have parallel second fundamental form and at most one nonzero principal curvature is rigid, without any assumption on the scalar curvature.

math.DG↗

Automorphic symbols and automorphic \emph{L}-values of GL(2) over imaginary quadratic fields

We study non-vanishing modulo primes of critical values of $L$-functions over imaginary quadratic fields twisted by Coates--Wiles characters. For a parallel weight two Hecke eigenform over an imaginary quadratic field, we prove that a positive proportion of Coates--Wiles characters have nonzero integral $L$-values modulo each prime in a positive-density set. The argument constructs automorphic symbols in parabolic homology with integral coefficients and expresses the critical values through their pairings with parabolic cohomology classes. A vertical family of additive averages of these pairings recovers Fourier coefficients and force the module generated by symbols to have full rank. We transfer this full-rank property to non-vanishing modulo primes. This extends the homological strategy of Kim--Sun from classical modular curves to arithmetic orbifolds, while addressing the unit obstructions specific to the imaginary quadratic setting.

math.NT↗

Event-Aware Spatiotemporal Precipitation Forecasting with Geographic Context and Physics-Guided Regularization

Hourly precipitation forecasting involves several distinct statistical challenges, including spatially varying predictor-precipitation relationships, a strongly imbalanced precipitation distribution, and progressive degradation of precipitation event skill with increasing lead time. We develop a spatiotemporal forecasting framework that addresses these challenges through three complementary components: explicit geographic representation, event-aware learning for imbalanced precipitation, and weak asymmetric regularization derived from the atmospheric water budget. The physical information is treated as an asymmetric constraint rather than as an additional prediction target, designed to discourage physically unsupported precipitation attenuation without replacing the data-driven forecast. Experiments using ERA5 data across regional, enlarged domain, and spatial subset settings show that explicit geographic information improves spatial field prediction, while event-aware learning provides the most consistent gains in detecting moderate and heavy precipitation events. Physical regularization has a more selective effect, mainly reducing systematic underprediction over the enlarged domain while improving longer lead precipitation event prediction in the spatial subset experiment. These results indicate that the benefit of physical guidance depends on the available data regime and becomes most apparent when data-driven precipitation information deteriorates with increasing lead time.

stat.AP↗

Light-induced interlayer spacing dynamics via orbital phonon coupling

Interlayer coupling controls the electronic properties of layered van der Waals transition metal dichalcogenides. We investigate light-induced control of the interlayer spacing through orbital-selective excitation in trilayer 1T'-WSe2 and 1T'-WS2. Real-time time-dependent density functional theory simulations show that the interlayer spacing contracts or expands depending on whether chalcogen p-orbital density is depleted from or accumulated in the interlayer region. Effective Lindblad models coupled to the lattice dynamics reproduce these contrasting responses with simplified dynamics. A single effective excited state captures the cosine-like displacive motion in trilayer 1T'-WSe2, whereas the shift of the equilibrium spacing in trilayer 1T'-WS2 requires two excited states with different electron-phonon couplings and relaxation channels. Static calculations at varied interlayer spacings indicate that these spacing changes modify the electronic gaps and could access different electronic phases. These results connect orbital redistribution, carrier relaxation, and interlayer breathing motion, and establish orbital-selective optical excitation as a route to tuning the electronic properties of layered materials.

cond-mat.mtrl-sci↗

Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration

KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers. KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference. This contrast raises the question of whether per-KV compression and multi-KV aggregation can be bridged within a single mechanism for efficient attention. We identify RAM-Net as such a bridge through soft assignments over a discrete address space. These assignments determine recurrent updates to the continuous slot state associated with each address. Under a restricted RAM-Net construction, we prove that soft address assignments extend hard quantized matching to a separable read-write overlap that locally approximates full-attention similarity and supports recurrent aggregation. These connections further enable Transformer-to-RAM-Net weight migration through a new path based on a soft-quantized intermediate construction. Across nine pretrained Transformer models from 0.3B to 7B parameters, RAM-Net recovers an average of 87.1% of the teachers' accuracy gains over random guessing across six commonsense and knowledge tasks using only a 500M-token budget per model.

cs.LG↗

GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images

Reconstructing complete, scene-aligned 3D objects from casual images requires integrating sparse, uncertain observations and inferring surfaces hidden by occlusions. We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images. Our local modality mixer couples patch-aligned RGB, target-mask, and pointmap features before cross-view reasoning, preserving scene context while distinguishing the target from its surroundings. Text-guided semantic conditioning complements these spatial cues with category names and object captions through stage-specific adapters for structure, geometry, and appearance generation. The generated asset initializes a multimodal agent, providing instance-specific geometry and pose for targeted structural and texture refinement through an observation-guided edit-render-review loop. Across synthetic objects, cluttered tabletops, and indoor scenes, GATOR achieves strong geometric and appearance fidelity while recovering scene-relative pose from sparse observations. Time-budget comparisons and scene-level simulation further demonstrate the reconstruction efficiency and simulation readiness. Project page: https://research.nvidia.com/labs/lpr/gator/

cs.CV↗

The Lattice of Transition Laws

Diffusion and autoregression (AR) have long been seen as different categories of generative models, with diffusion specialising in continuous fields and AR specialising in discrete tokens. Recent work seeks to combine the advantages of the two models, and each hybrid fixes its decoding schedule by design. In this paper, we ask whether the performance of decoding schedules of one model can be predicted before decoding at a fixed number of steps. We describe diffusion, AR, and models in between as paths on one corruption lattice, and define the cost of a schedule as the dependence its parallel steps discard. The cost shows that the fewest steps of a zero-cost schedule are set by the geometry of the data, in the same way for tokens and for continuous fields. In particular, for data that are Markov on a graph and dependent along its paths, the fewest steps equal the graph's treedepth, which is logarithmic in the length of a sequence and linear in the side length of a grid. With fewer steps than the treedepth, every schedule pays a positive cost, whose ranking we predict before decoding with a kernel of pairwise dependence estimated from pretrained weights. Across text generation, image generation, and video generation, we verify most of the predictions about the rankings of different schedules under different metrics and benchmarks. This work therefore provides a design principle for decoding for future AR models, diffusion models, and anything in between. Our code is available at https://github.com/TSUITUENYUE/The-Lattice-of-Transition-Laws.

cs.LG↗

BRACE: Differential Privacy for Dense Associative Memory with LSR Energy

Dense associative memory (DAM) provides an energy-based framework for memory retrieval with close connections to attention mechanisms in modern artificial intelligence. Despite growing interest in differential privacy for AI, the privacy of DAM retrieval dynamics remains relatively unexplored. In this paper, we develop a differential privacy framework for log-sum-ReLU (LSR) dense associative memory, whose finite-support retrieval dynamics pose distinctive challenges for privacy-preserving computation. We propose the Boundary-Responsive Adaptive Correction Evolution (BRACE) algorithm, a differentially private retrieval mechanism for LSR-DAM that adaptively corrects boundary-sensitive perturbations to control their cumulative effect over the retrieval trajectory. In theory, we prove that our method is minimax optimal by deriving dimension-independent terminal and full-trajectory retrieval error rates, with optimal dependence on the inverse temperature and, in the growing-horizon regime, the retrieval horizon. We further establish central limit theorems that enable uncertainty quantification for private retrieval by characterizing its asymptotic distribution and the additional variability introduced by privacy. Numerical experiments compare our proposed method with baseline differential privacy approaches and evaluate its retrieval accuracy. Together, our results provide a theoretical foundation for optimal privacy-preserving retrieval and uncertainty quantification in energy-based associative memory systems.

cs.CR↗

A computational framework for the acoustic characterization of Suikinkutsu (water harp cave)

Suikinkutsu are partially concealed acoustic devices traditionally installed adjacent to stone water basins (tsukubai) in Japanese temple gardens. While previous research has examined their physical acoustics and cultural significance, computational methods for documenting and comparing their acoustic characteristics remain underexplored. This exploratory study introduces a computational framework for analyzing suikinkutsu using spectral and temporal audio features extracted from field recordings. Rather than estimating the physical geometry of individual suikinkutsu, the framework characterizes their acoustic profiles using a multidimensional set of computational descriptors, enabling systematic comparisons between recordings across different sites and installations. Audio recordings from nine suikinkutsu located at temples, shrines, and gardens in the Kansai region of Japan were analyzed using a feature-extraction pipeline that incorporates Fast Fourier Transform (FFT), Mel-Frequency Cepstral Coefficients (MFCCs), Root Mean Square (RMS) energy, spectral centroid, spectral bandwidth, and spectral contrast. The extracted features indicate that, although the suikinkutsu share broadly similar spectral and timbral characteristics, each exhibits a measurably distinct acoustic profile; a supplementary analysis of within-site and between-site variability suggests these profiles reflect a bounded range of characteristic resonant behaviour rather than a fixed acoustic signature. The proposed framework provides one of the first computational approaches for systematically characterizing and comparing suikinkutsu, contributing a practical foundation for future research and the digital preservation of this distinctive form of Japanese environmental design.

cs.SD↗

How Firm Should a Grasp Be?

An ideal robot grasp is firm enough to securely handle an object, yet gentle enough to avoid damaging it. Achieving this balance requires knowledge of the object's material properties, such as its mass, elasticity, and surface friction. These properties, however, are seldom precisely known a priori. In this work, we propose a visuotactile approach to estimating material properties in real time, during the process of grasping. Our method uses these estimated properties to determine the minimum grasp force required to handle the object. We contribute a new dataset of real-world objects (fruits and vegetables) with measured physical properties (shape, mass, elasticity, and friction), which we use to construct our force estimation model via simulations. We experimentally validate our approach to grasp force control using a robot with a parallel-jaw gripper. We demonstrate our system's ability to gently grasp a wide variety of objects, in each case adapting to their unique physical properties.

cs.RO↗