arXiv · 2512.16282
CALM: A CKA-Guided Adaptive Layer-Wise Modularization Framework for LLM Quantization
Abstract
Current mainstream post-training quantization methods for large language models typically apply a uniform quantization strategy across all network layers, overlooking the substantial differences in algorithmic suitability among layers. To address this limitation, we propose CALM (A CKA-guided Adaptive Layer-wise Modularization)a fine-tuning-free, plug-and-play framework for algorithmic heterogeneous quantization. CALM independently evaluates multiple PTQ algorithms on each layer and employs Linear Centered Kernel Alignment (CKA) as a metric to automatically select the optimal quantization strategy per layer. The individually optimized strategies are then integrated to construct a hybrid quantized model. Experiments demonstrate that our approach consistently outperforms both uniform quantization baselines and state-of-the-art mixed-precision methods across mainstream LLMsincluding LLaMA and Qwenin terms of perplexity (PPL) and downstream task performance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jinhao Zhang, Yunquan Zhang, Daning Chen, JunSun, Zicheng Yan. 2026-01-08. CALM: A CKA-Guided Adaptive Layer-Wise Modularization Framework for LLM Quantization. https://arxiv.org/abs/2512.16282
Cite the original work for its findings. Save a collection to share your selection of sources.