arXiv · 2609.27437
Mamba-Family State-Space Model Kernels on a Programmable CGLA
Abstract
Edge and embedded inference is constrained by power and data movement. Mamba-family state-space models replace attention with sequence-linear recurrence, but their inference path combines dense projections, short-reduction SSD kernels, and recurrent-state updates. This paper maps these kernel groups onto IMAX, a programmable CPU-Grounded Linear Array (CGLA), and measures them from kernel execution to token-level integration. Projection kernels match the long-reduction IMAX pipeline, whereas SSD Step-1 is limited by short reductions and kernel-boundary overheads. Mamba-130M token-level integration identifies projection GEMV as the decode bottleneck. These results show that programmable CGLAs fit long-reduction projection kernels, while SSD and decode-time projection support require boundary reduction and persistent-weight execution.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Takuto Ando, Yasuhiko Nakashima. 2026-09-23. Mamba-Family State-Space Model Kernels on a Programmable CGLA. https://arxiv.org/abs/2609.27437
Cite the original work for its findings. Save a collection to share your selection of sources.