Search arXiv⌕ Search

arXiv · 2609.33207

MorphAtt: A Neuromorphic Accelerator for Efficient Multi-Head Attention Processing in Spiking Vision Transformers

Abstract

Spiking Vision Transformers (SViTs) are developed as an energy-efficient alternative to conventional ViTs for computer vision tasks at the edge. However, huge parameter counts and complex multi-head self-attention (MHSA) operations make it challenging to achieve high energy efficiency in SViT inference, especially in tightly constrained applications. To maximize efficiency gains of SViT processing, we propose MorphAtt, a novel digital accelerator that expedites SViT inference through streamlined processing. Specifically, it processes MHSA operations using cascaded hardware modules: a Spiking Query-Key-Value generator (SpikeQKV), a low-complexity Spiking Multi-Head Self-Attention engine (SpikeAtten), and Reparameterization Convolution (RepConv) modules. To mitigate traffic congestion in on-chip memory accesses and data reuse, specialized inter-module buffers are integrated within the dataflow. Under synthesis using 32nm CMOS technology, MorphAtt achieves 792-1605 GOPS of throughput, while incurring ~39-55 mW of power consumption and 1.5 mm^2 of area, which lead to 20.3-29.1 TOPS/W of energy efficiency. These results also demonstrate that our MorphAtt offers better performance and efficiency trade-offs than state-of-the-art, thereby enabling highly energy-efficient vision-based AI systems at the edge.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rachmad Vidya Wicaksana Putra, Amirhesam Jafari Rad, Muhammad Shafique. 2026-09-27. MorphAtt: A Neuromorphic Accelerator for Efficient Multi-Head Attention Processing in Spiking Vision Transformers. https://arxiv.org/abs/2609.33207

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Edge GPU Aware Multiple AI Model Pipeline for Accelerated MRI Reconstruction and Analysis

Advancements in AI have greatly enhanced the medical imaging process, making it quicker to diagnose patients. However, very few have investigated the optimization of a multi-model system with hardware acceleration. As specialized edge devices emerge, the efficient use of their accelerators is becoming increasingly crucial. This paper proposes a hardware-accelerated method for simultaneous reconstruction and diagnosis of \ac{MRI} from \ac{CT} images. Real-time performance of achieving a throughput of nearly 150 frames per second was achieved by leveraging hardware engines available in modern NVIDIA edge GPU, along with scheduling techniques. This includes the GPU and the \ac{DLA} available in both Jetson AGX Xavier and Jetson AGX Orin, which were considered in this paper. The hardware allocation of different layers of the multiple AI models was done in such a way that the ideal time between the hardware engines is reduced. In addition, the AI models corresponding to the \ac{GAN} model were fine-tuned in such a way that no fallback execution into the GPU engine is required without compromising accuracy. Indeed, the accuracy corresponding to the fine-tuned edge GPU-aware AI models exhibited an accuracy enhancement of 5\%. A further hardware allocation of two fine-tuned GPU-aware GAN models proves they can double the performance over the original model, leveraging adequate partitioning on the NVIDIA Jetson AGX Xavier and Orin devices. The results prove the effectiveness of employing hardware-aware models in parallel for medical image analysis and diagnosis.

cs.AR↗

Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy

Modern AI accelerators rely on matrix multiply-accumulate units (MMAUs), such as NVIDIA Tensor Cores and AMD Matrix Cores, to accelerate deep neural network workloads. MMAUs expose only instruction-level or API-level interfaces of matrix multiply-accumulate (MMA) operations, while leaving internal floating-point arithmetic behavior undocumented. Consequently, MMAUs across vendors and architectural generations often produce numerical discrepancies for identical inputs, and sometimes exhibit reduced numerical accuracy that can cause training instability. Diagnosing and understanding the root causes of these effects is challenging without white-box models of their arithmetic behavior. This paper proposes closed-loop feature probing (CLFP), a generic and systematic framework for constructing bit-accurate arithmetic behavior models of MMA operations. Based on this framework, we analyze all MMA instructions on ten GPU architectures spanning NVIDIA Volta through RTX Blackwell and AMD CDNA1 through CDNA3, and derive the first bit-accurate arithmetic models for these MMAUs. Our models explain previously observed cross-platform numerical discrepancies and accuracy issues, enable white-box numerical error analysis, reveal four types of precision bottlenecks and one type of numerical asymmetry, and inform software workarounds as well as design suggestions for future MMAUs. This work is open-source at https://github.com/microsoft/MMA-Sim

cs.AR↗

Coarse-to-Fine Macro Placement via Evolutionary Search and Critical Macro Tuning

Macro placement is a critical stage in chip physical design that substantially affects downstream implementation quality. Recent search-based methods improve existing layouts through partial reconstruction, but quality-biased or spatially restricted macro selection can limit the diversity of reconstruction proposals, potentially hindering escape from local optima. Moreover, coarse-grid representations restrict placement precision. To address these challenges, we propose C2FPlace, a \textbf{C}oarse-to-\textbf{F}ine macro \textbf{Place}ment framework that integrates population-based evolutionary search with fine-grained refinement. During coarse-grained optimization, tournament selection chooses promising parents from randomly sampled groups of layouts, and stochastic partial rip-up and re-place generates offspring by sampling macro subsets across the entire layout. A two-phase schedule samples reconstruction ratios from a higher range early in the search and a lower range later, supporting broad exploration followed by more conservative refinement. During fine-grained optimization, critical macro tuning enables positional adjustments beyond the coarse grid to obtain additional half-perimeter wirelength (HPWL) reduction. Experiments on the ISPD2005 benchmark show that C2FPlace reduces HPWL by 17.82\% over EGPlace and 17.86\% over RollPlace on average. On the ICCAD2025 benchmark, C2FPlace achieves the best average ranking among the compared methods under the evaluated power, performance, and area (PPA) metrics. Our codes are available in \href{https://github.com/lxxxxb/C2FPlace}{https://github.com/lxxxxb/C2FPlace}.

cs.AR↗