arXiv · 2501.16103
Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference
Abstract
It has long been a problem to arrange and execute irregular workloads on massively parallel devices. We propose a general framework for statically batching irregular workloads into a single kernel with a runtime task mapping mechanism on GPUs. We further apply this framework to Mixture-of-Experts (MoE) model inference and implement an optimized and efficient CUDA kernel. Our MoE kernel achieves up to 91% of the peak Tensor Core throughput on NVIDIA H800 GPU and 95% on NVIDIA H20 GPU.
Explore related subjects
Keep this discovery
Yinghan Li, Yifei Li, Jiejing Zhang, Bujiao Chen, Xiaotong Chen, Lian Duan, Yejun Jin, Zheng Li, Xuanyu Liu, Haoyu Wang, Wente Wang, Yajie Wang, Jiacheng Yang, Peiyang Zhang, Laiwen Zheng, Wenyuan Yu. 2025-01-27. Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference. https://arxiv.org/abs/2501.16103
Cite the original work for its findings. Save a collection to share your selection of sources.