arXiv · 2610.03218
VisionMX: Unlocking Microscaling Post-Training Quantization for Vision Models
Abstract
Microscaling (MX) formats are emerging as a hardware-supported approach to efficient training and inference. They combine low-precision elements with shared block scales, but their impact on vision models remains underexplored. We systematically investigate post-training MX quantization across vision models and tasks. An analysis of direct conversion identifies three sources of error: block-scale representation, the poor alignment of some small convolutional weight tensors with nonuniform element grids, and the underuse of signed codes by nonnegative activations. These findings motivate VisionMX, a post-training MX quantization method that optimizes bounded weight rounding and applies a foldable affine correction to activations. We evaluate VisionMX across image classification, object detection, semantic segmentation, and low-light image enhancement using several MX-style formats. It improves on direct conversion and the evaluated post-training quantization baselines, with the largest performance recoveries in architectures most sensitive to MX conversion
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Elad Dror Cohen, Ofir Gordon, Lior Dikstein, Idan Achituve, Hai Victor Habi. 2026-10-02. VisionMX: Unlocking Microscaling Post-Training Quantization for Vision Models. https://arxiv.org/abs/2610.03218
Cite the original work for its findings. Save a collection to share your selection of sources.