PRQuant: Permutation Residual Quantization for Low-Overhead Inference
Low-bit quantization of linear layers is often dominated by a small number of outlier channels. Existing smoothing, rotation, and residual-based methods can mitigate this issue, but may shift the quantization bottleneck to weights or introduce costly online operations. To address these limitations, we propose PRQuant (Permutation Residual Quantization), a training-free framework that combines channel permutation with offline weight residual compensation. PRQuant identifies the scaled-weight columns with the largest quantization errors and permutes them into contiguous tail blocks. This structure allows the corresponding weight residuals to be precomputed entirely offline, while replacing scattered activation gathering with simple contiguous access during inference, yielding a single regular MXFP4 GEMM for compensated computation. Experiments show that PRQuant substantially reduces down-projection reconstruction error, with scaling and residual compensation providing the main numerical gains while permutation enables a hardware-friendly contiguous layout. Comprehensive experiment results on Qwen3-4B-Instruct-2507 and Qwen3-30B-A3B-Instruct-2507 illustrate that PRQuant achieves up to averagelly 2.6x and 1.8x operator speedup over BF16 respectively, while preserving near plain MXFP4 end-to-end decoding efficiency. Across five downstream benchmarks, PRQuant achieves the best average accuracy among the quantized methods, improving accuracy over MXFP4 by 1.24 and 0.55, respectively.