Search arXivSearch

arXiv · 2605.00555

Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis

Abstract

To efficiently support Large Language Models (LLMs), modern GPGPU architectures have introduced new features and programming paradigms, such as warp specialization. These features enable temporal overlap between the producer and consumer, as well as between matrix multiplication and activation function operations, substantially improving performance. To conduct effective AI infrastructure and computer architecture research, cycle-accurate simulators that support these new features, together with analytical models that faithfully capture workload characteristics, are essential. However, existing academic tools provide limited support for these emerging requirements. Existing cycle-accurate simulators do not incorporate new NVIDIA GPU features, such as the Tensor Memory Accelerator (TMA), in a timely manner. Moreover, existing analytical models can misestimate DRAM traffic under certain configurations. In this paper, we build Sim-FA, a cycle-accurate simulation framework for Hopper TMA/WGMMA pipelines. We first develop an operator-agnostic trace frontend that instruments kernels at the Triton TTGIR level and validates it on 23 GEMM shapes, achieving 5.49\% MAPE against H800, confirming that the simulator core is not tied to any single operator. Because FlashAttention-3 introduces additional complexity beyond standard TMA/WGMMA kernels (asymmetric producer-consumer pipelines, softmax, ping-pong synchronization), we further build an FA3-specialized frontend that achieves 5.7\% MAPE with a maximum error of 12.7\%. Within the same framework, SimFA-python serves as an analytical fast path for large-scale design-space exploration where cycle-accurate simulation is prohibitively slow; validated against cuTile kernels on Blackwell (GB10), it explains why existing analytical models can produce inaccurate traffic estimates.

Explore related subjects

Keep this discovery

BibTeXRIS

Zhongchun Zhou, Yuhang Gu, Chengtao Lai, Ya Wang, Zeyu Han, Wei Zhang, Jun Liu. 2026-09-02. Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis. https://arxiv.org/abs/2605.00555

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Deep belief networks are exact

We prove that every strictly positive probability distribution on \(\{-1,1\}^n\) is represented exactly by a sigmoid belief network with finite parameters. This answers a question of Sutskever and Hinton. The proof upgrades their probability-sharing approximation to exact representation using Brouwer's fixed-point theorem.

cs.AI

Are Widely Known Findings Easier to Retract?

Failures of retraction are common in science. Why do they occur? And what determines whether a retraction is successful? We use data from citation records and Altmetrics to test proposed answers to these questions. LaCroix et al. employ network models to argue the social spread of information helps explain failures of retraction. One prediction is that widely known results, surprisingly, should be easier to retract, since their retraction is more relevant. Our results support this conclusion. We find highly cited papers show more significant reductions in citation after retraction and garner more attention to their retractions as they occur.

cs.DL

Hardware-conscious Software Training for Deep Neural Network Inference Accelerator Chips to Recover Accuracy Degradation due to Hardware Variabilities

Deep neural network (DNN) has been widely applied in various industries. Specialized chips are being discussed for the purpose of achieving lower power consumption with higher throughput. Hardware variations introduced during the process of chip manufacturing are the main reason for affecting the inference accuracies. In this paper, we propose hardware-conscious software training (HCST) method which enables high inference accuracies even under the influence of hardware variations.

cs.AR