arXiv · 2610.02882
DyRA: Dynamic Residual Approximation for Efficient Matrix Multiplication in DNNs
Abstract
Large-scale foundation models achieve strong performance across diverse tasks, but their size makes inference costly, largely due to dense matrix multiplications. Prior work reduces this cost by replacing dense weight matrices with efficient structured forms such as low-rank factorizations. However, these methods approximate weights rather than the output activations that determine inference accuracy. Consequently, small weight-space errors can be amplified by input activations, producing large output errors. In this work, we propose DyRA, an input-adaptive method that improves structured matrix multiplication approximation by correcting residual output errors during inference. We show that matrix multiplication can be approximated more effectively by directly optimizing low-rank factors of the output. DyRA builds on this insight by dynamically approximating and correcting the output error introduced by structured weight approximations. This combines efficient structured computation with input-dependent correction, yielding a more faithful approximation of full matrix multiplication under the same computational budget. Across vision, speech, and language models, DyRA consistently improves the accuracy-efficiency trade-off over structured weight approximations alone. Notably, DyRA achieves a 1.5$\times$ end-to-end GPU speedup for DINOv3 while reducing accuracy degradation by more than 3$\times$ relative to weight-only baselines.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Daewon Chae, Hyunwon Chung, Changwoo Lee, Hun-Seok Kim. 2026-10-02. DyRA: Dynamic Residual Approximation for Efficient Matrix Multiplication in DNNs. https://arxiv.org/abs/2610.02882
Cite the original work for its findings. Save a collection to share your selection of sources.