Search arXiv⌕ Search

arXiv subjects

Amar Kanakamedala

Publications and source records attributed to Amar Kanakamedala.

2 recordsLinked to original sources

SOCKET: SOft Collision Kernel EsTimator for Sparse Attention

Exploiting sparsity is key to efficient long-context inference, as attention dominates the cost of autoregressive decoding. Sparse attention reduces this cost by restricting computation to a subset of tokens, but its effectiveness hinges on fast and accurate token scoring and selection at inference time. Data-agnostic approaches offer an attractive way to perform this selection, but often incur substantial memory overhead to maintain high recall. We revisit Locality-Sensitive Hashing (LSH) and introduce SOCKET, a SOft Collision Kernel EsTimator that replaces hard bucket matches with probabilistic, similarity-aware aggregation. Traditional LSH relies on binary collision signals, providing limited information for ranking tokens and necessitating many hash tables for accurate retrieval. In contrast, soft LSH accumulates graded collision evidence across hash tables, closely preserving the true top-$k$ ordering with significantly less memory. This reframes LSH from a candidate-generation mechanism into a principled scoring kernel for sparse attention. Building on this insight, SOCKET enables efficient token selection without ad hoc voting and matches or outperforms existing sparse attention methods across multiple long-context benchmarks and diverse language models. With a custom set of CUDA/Triton kernels for scoring, selection, and attention, SOCKET achieves up to approximately $1.5\times$ higher throughput than FlashAttention. Code is open-sourced at https://github.com/amarka8/SOCKET.

cs.LG↗

RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts

Softmax Attention has a quadratic time complexity in sequence length, which becomes prohibitive to run at long contexts, even with highly optimized GPU kernels. For example, FlashAttention-2/3 (exact, GPU-optimized implementations of Softmax Attention) cannot complete a single forward-backward pass of a single attention layer once the context exceeds ~4 million tokens on an NVIDIA GH200 (96 GB). We introduce Repeated Arrays-of-Count Estimators (RACE) Attention, a kernel-inspired alternative to Softmax Attention that is strictly linear in sequence length and embedding size. RACE Attention replaces the exponential kernel with a sharpened angular similarity, and approximates attention outputs via Gaussian random projections and soft Locality-Sensitive Hashing (LSH), avoiding construction of the full attention matrix. Across language modeling, masked language modeling, and text/image classification, RACE Attention matches or outperforms strong baselines up to 64K seqeuence length while reducing wall-clock time and memory usage. In addition, we conduct a controlled scaling study on a single attention layer and demonstrate processing of up to 12 million tokens on an NVIDIA GH200 GPU and 75 million tokens on an Intel Xeon Gold 5220R CPU in a single forward-backward pass, which is well beyond the capabilities of current state-of-the-art attention implementations. RACE Attention thus offers a practical and theoretically grounded mechanism for long-context training on today's hardware. We release our code at https://github.com/sahiljoshi515/RACE_Attention.

cs.LG↗