Search arXiv⌕ Search

arXiv · 2609.37538

AVSG: Accelerated Vectorized Sparse Gather for Efficient KV Cache Offload in Sparse-Attention LLM Serving

Abstract

Dynamic sparse attention reduces long-context attention computation by selecting only a subset of tokens, but still requires access to the full KV cache, leaving serving memory-bound. Offloading the KV cache to host memory reduces device memory pressure but places H2D transfers on the decoding critical path. In DSA, substantial overlap in selected KV entries across decoding steps creates an opportunity for device-resident reuse. Exploiting this reuse efficiently, however, presents three critical challenges: costly matching of selected tokens against entries retained in the HBM buffer, uneven distribution of the remaining H2D transfers across accelerator cores, and retaining frequently accessed KV entries within limited HBM capacity. We present AVSG, an operator that addresses these challenges for DSA while preserving its exact token selections. AVSG uses vectorized hash matching to identify reusable HBM-buffer slots and reserve slots for entries requiring transfer, then evenly partitions these H2D transfers across accelerator cores. Lifetime-based buffer management retains frequently accessed KV entries in the device buffer across decoding steps, and shared slot reservations extend reuse across tokens within a multi-token prediction iteration. On a single NPU, vectorized hash matching is 2.80x faster than scalar dual-pointer matching, miss-only transfer raises effective H2D bandwidth by up to 34.44x over request-level assignment, and an 8K-entry buffer reaches a 94.83% hit rate on real requests. These gains translate into end-to-end improvement: on a serving stack processing a real-world production dataset, AVSG reduces time per output token by 39% and increases output throughput by 1.27x relative to the same KV offload layout without HBM-buffer reuse, demonstrating the benefit of efficient matching, balanced transfer, and effective HBM residency in production-scale serving.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wenwei Kuang, Xiangyu Wang, Chong Wu, Jun Wang, Weijie Zhang, Brian K Chen, Longwen Lan, Ken Zhang. 2026-09-29. AVSG: Accelerated Vectorized Sparse Gather for Efficient KV Cache Offload in Sparse-Attention LLM Serving. https://arxiv.org/abs/2609.37538

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Memory in Behavioral Models as Motion on a Slow Invariant Manifold

A single-tone large-signal operating point of a nonlinear two-port is a periodic orbit of a periodically forced circuit. When the device has memory (self-heating, trapping), the Floquet exponents of that orbit separate into fast (electrical) and slow (thermal and trapping) modes, and long-term memory is motion on the invariant manifold attached to the slow modes. An existence and uniqueness theorem for that manifold follows from the parameterization method of Cabré, Fontich and de la Llave, applied to the stroboscopic map at the orbit; the manifold is the spectral submanifold of Haller and Ponsioen, without a small-forcing parameter. The manifold is a bundle over the circle of drive phase, its fiber dimension the number of slow Floquet exponents, and the dynamic X-parameter kernel of Verspecht et al. identifies its reduced dynamics from step changes of the drive amplitude. Consequently, an exact reduced model has as many memory states as slow exponents, the memoryless X-parameter surface is the fixed-point family of the reduced dynamics, and the envelope-domain model is the reduced dynamics driven by the envelope. The hypotheses are verified and the manifold constructed for a GaN HEMT compact model with a three-pole thermal network and a drain-lag trap: the trap contributes a $14\,μ$s time constant set by the linearization and not by its $6$ ms emission time, the thermal submanifolds are nearly flat with linear reduced dynamics, and the expansion in the trap direction is valid only within a few thermal voltages ($nV_T\approx26$ mV), so trap memory needs a global representation of the manifold.

eess.SP↗

Blind Interference Suppression in IRS-Aided Wireless Systems: A Statistical Channel Ratio Estimation Approach

This paper addresses the problem of suppressing non-cooperative interference in intelligent reflecting surface (IRS)-aided wireless links without any channel state information (CSI) or cooperation from the interferer. We propose a fully blind framework that relies solely on received signal power measurements. A key insight is that nulling the aggregate interference channel requires only the complex ratios between the IRS-reflected paths and the direct interference link, rather than absolute CSI. We develop a novel estimation algorithm that obtains unbiased estimates of these channel ratios using only power samples collected under random IRS configurations. Theoretically, we prove that unbiased estimation is feasible when the number of discrete phase levels $K\geq 3$, and establish the Cramer-Rao lower bounds (CRLBs) for both the phase offset and amplitude ratio estimates, thus providing design guidance. Based on the estimated ratios, we propose two low-complexity IRS phase optimization algorithms: a one-shot greedy method and an iterative variant that mitigates error propagation from weakly reflecting elements. Simulations demonstrate that the proposed schemes can suppress strong interference to within a few dB of the interference-free upper bound, offering a practical, CSI-free solution for robust wireless communications in contested spectral environments.

eess.SP↗

Blind Interference Suppression for IRS-Aided Robust Wireless Communications

The application of intelligent reflecting surfaces (IRSs) to suppress interference in wireless communication systems has recently attracted significant research attention. Most existing approaches rely on complete or partial channel state information (CSI) to configure the IRS. However, acquiring accurate CSI in IRS-assisted systems involves considerable pilot overhead and introduces non-negligible delays. This issue is further exacerbated under strong interference conditions, where interfering sources are typically non-cooperative, making CSI acquisition even more challenging. As a result, existing CSI-dependent interference suppression methods become difficult to deploy in practice. To address these limitations, we propose a novel blind interference suppression strategy that combines a proportional phase-inversion (PPI) algorithm with the conditional sample mean (CSM) method. The proposed approach determines the IRS configuration using only the received signal power, without requiring any prior CSI. We conduct a comprehensive performance evaluation by deriving the theoretical performance of the proposed scheme, which is subsequently verified through numerical simulations. Furthermore, simulation results across various parameter settings demonstrate that the proposed blind interference suppression scheme reduces the interference power to the level of noise, thereby achieving a marked signal-to-interference-plus-noise ratio (SINR) improvement and outperforms existing benchmark schemes.

eess.SP↗