Search arXiv⌕ Search

arXiv · 2608.19046

APEX: A Dual-Sparsity Accelerator for Precise and Efficient SNN Inference

Abstract

Spiking Neural Networks (SNNs) have emerged as an energy-efficient alternative to Artificial Neural Networks (ANNs), leveraging sparse accumulate operations in the place of power-hungry multiply-and-accumulate operations. ANN-SNN conversion is a widely adopted approach to realize deep SNNs with accuracy comparable to that of ANNs. The Quantization-Clip-Floor-Shift (QCFS) activation minimizes conversion error, yet requires a large number of inference timesteps to match the source ANN accuracy on real-world vision datasets. PASCAL addresses this by proposing the Precise ANN-SNN Conversion Integrate-and-Fire (PASC-IF) neuron, which guarantees mathematical equivalence between the converted SNN and the source ANN, thereby achieving ANN-equivalent accuracy at significantly reduced timesteps. Despite this algorithmic advancement, the hardware implications of deploying the PASC-IF neuron remain unexplored. In this work, we present APEX, a dual-sparsity SNN inference accelerator that integrates the PASC-IF neuron into the LoAS hardware framework. The three-stage PASC-IF datapath is realized as a fully combinational circuit with no additional latency cost. APEX exploits dual sparsity in both input spikes and weights through a fully temporal-parallel dataflow, enabling efficient sparse computation and reduced memory traffic. Across all evaluated models, the PASC-IF neuron on average achieves up to 3% higher accuracy than the standard IF neuron, with a power overhead of only 1.3%-5.4%, an area overhead of 2.1%-2.7%, and 40% energy reduction for best accuracy configurations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Devgokul Bawa Venkatesh, Sreeram Radhakrishnan, Rajshekhar Rakshit, Gopalakrishnan Srinivasan. 2026-08-19. APEX: A Dual-Sparsity Accelerator for Precise and Efficient SNN Inference. https://arxiv.org/abs/2608.19046

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Exploiting Decompression Latency for Covert Channels in Inter-Line-Compressed LLCs

The recently proposed XOR cache is an inter-line-compressed last-level cache (LLC) that leverages the data-inclusion relationship between the private caches and the LLC, compressing two cache lines into one by XORing them. The architecture relies on the cache coherence protocol for data decompression. In this paper, we demonstrate that this mechanism - specifically the latency asymmetry between a cache hit on an uncompressed vs. compressed line - introduces microarchitectural vulnerabilities. Based on this observation, we propose a covert channel attack targeting the XOR cache. A colluding sender controls the receiver's access latency by triggering decompression through targeted write requests to partner cache lines. By exploiting the data-dependent compression behavior of the XOR cache, the sender and receiver establish the channel using pre-agreed data values. The channel achieves higher bandwidth than the Prime+Probe baseline for two reasons: first, each bit is encoded in the compression state of an individual line rather than the occupancy of a cache set, so a single set carries multiple bits; second, each bit is resolved by manipulating coherence-protocol state rather than forcing shared-cache evictions, so it costs fewer LLC accesses and demand misses than Prime+Probe. Full-system simulations show a bandwidth of 2.9 Mbps at an observed 0.98% bit-error rate (BER) over 50,000 transmitted bits, 13.1 times the bandwidth of Prime+Probe under the same sub-1%-BER selection rule.

cs.AR↗

Mamba-Family State-Space Model Kernels on a Programmable CGLA

Edge and embedded inference is constrained by power and data movement. Mamba-family state-space models replace attention with sequence-linear recurrence, but their inference path combines dense projections, short-reduction SSD kernels, and recurrent-state updates. This paper maps these kernel groups onto IMAX, a programmable CPU-Grounded Linear Array (CGLA), and measures them from kernel execution to token-level integration. Projection kernels match the long-reduction IMAX pipeline, whereas SSD Step-1 is limited by short reductions and kernel-boundary overheads. Mamba-130M token-level integration identifies projection GEMV as the decode bottleneck. These results show that programmable CGLAs fit long-reduction projection kernels, while SSD and decode-time projection support require boundary reduction and persistent-weight execution.

cs.AR↗

Energy-Oriented CGLA Mapping of a Memory-Polynomial Digital Predistortion Kernel

Memory-polynomial digital predistortion (DPD) evaluates a small fixed coefficient set over a sliding input history, so its reduction step is a complex-MAC workload with local reuse. We map this DPD reduction kernel onto In-Memory Accelerator eXtension (IMAX), a programmable CPU-Grounded Linear Array (CGLA) composed of a one-dimensional processing-element/local-memory pipeline. For a (P,M)=(5,5) odd-order memory-polynomial instance, the mapping keeps the 120 B coefficient set in local memory, advances the five-tap history over 1024-sample tiles, and realizes the 15 order-delay terms as a 33-stage streaming complex-MAC reduction. The evaluation measures kernel latency and modeled energy. All measured paths use the same single-precision complex workload of 32 sequences, each with 2048 complex samples, across an IMAX FPGA prototype, a CUDA implementation on an RTX 4090 system, and an ARM-NEON implementation on Jetson AGX Orin. With this 1024-sample tile configuration, the IMAX FPGA prototype reports 20.201 ms end-to-end latency and 1.948 ms kernel-only latency. Using the previously reported 28 nm IMAX frequency and power model, the projected IMAX configuration gives 3.14 ms end-to-end latency and 0.34 ms kernel-only latency. The RTX 4090 baseline has the lowest end-to-end latency at 0.484 ms. Under model-based platform power accounting and the stated power assumptions, the projected IMAX configuration gives 169.1 times smaller modeled end-to-end energy per batch than the RTX 4090 baseline. This value uses platform power assumptions rather than workload-dependent runtime power or a direct silicon power measurement. A controlled synthetic PA-model validation checks that the same 15-term form improves test-set NMSE by 26.1 dB and ACLR by 26.0 dB. These results characterize the mapped memory-polynomial DPD reduction on IMAX for the evaluated tile configuration and power model.

cs.AR↗