arXiv · 2609.39361
LampAttention: Look-Ahead Mixed-Precision FlashAttention for Dedicated Accelerators
Abstract
While most attention logits can be computed in low precision without degrading numerical stability, current attention kernels fail to exploit this phenomenon. We introduce a novel hardware-algorithm co-design in the form of mixed-precision FlashAttention. Our method accumulates key-query products and evaluates their exponentials in 8-bit formats, then adaptively identifies sensitive sub-blocks and recomputes them in 16-bit formats. We propose the specifications for a dedicated accelerator capable of executing this pipeline efficiently. Simulated experiments with Qwen3 and Gemma 3 show that rerouting a selective minority of sub-blocks to high precision is sufficient to recover the baseline model performance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Stanislav Budzinskiy, Marian Gloser, Tolunay Yilmaz, Ying Hong Tham, Yuanyi Lin, Wenyi Fang, Fan Wu, Philipp Petersen. 2026-09-30. LampAttention: Look-Ahead Mixed-Precision FlashAttention for Dedicated Accelerators. https://arxiv.org/abs/2609.39361
Cite the original work for its findings. Save a collection to share your selection of sources.