Search arXivSearch

arXiv · 2410.09615

SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression

Abstract

Conventional model compression techniques for LLMs address high memory consumption and slow inference challenges but typically require computationally expensive retraining to preserve accuracy. In contrast, one-shot compression methods eliminate retraining cost, but struggle to achieve accuracy comparable to dense models. This paper presents SLIM, a new one-shot compression framework that holistically integrates hardware-friendly quantization, sparsity, and low-rank approximation into a unified process. First, we formulate the quantization process using a probabilistic approach (SLIM-Quant) that enables us to apply uniform quantization. Then, we use an existing one-shot pruning method to apply semi-structured sparsity on top of the quantized weights. Finally, to compensate for the introduced aggregated quantization and sparsity error, we use a novel saliency function with unique invertible and additive features that enables us to mathematically compute the value of low-rank adapters. SLIM improves model accuracy by up to 5.66% (LLaMA-2-7B) for 2:4 sparsity with 4-bit weight quantization, outperforming prior methods. Models compressed with SLIM achieve up to 4.3x and 3.8x on Nvidia RTX3060 and A100 GPUs, respectively. Additionally, they achieve up to 0.23x end-to-end memory reduction in comparison to their dense counterparts. We also propose an optional PEFT recipe that further improves accuracy by up to 1.66% (LLaMA-2-13B) compared to SLIM without fine-tuning.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mohammad Mozaffari, Amir Yazdanbakhsh, Maryam Mehri Dehnavi. 2025-08-14. SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression. https://arxiv.org/abs/2410.09615

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Online Regularized Statistical Learning in Reproducing Kernel Hilbert Space With Non-Stationary Data

We study recursive regularized learning algorithms in the reproducing kernel Hilbert space (RKHS) with non-stationary online data streams. We introduce the concept of a random Tikhonov regularization path and decompose the tracking error of the algorithm's output for the regularization path into random difference equations in RKHS. We show that the tracking error vanishes in mean square and almost surely if the regularization path is slowly time-varying. Then, leveraging the monotonicity of inverse operators and the spectral decomposition of compact operators, and introducing the RKHS persistence of excitation condition, we develop a dominated convergence method to prove the mean square and almost sure consistency between the regularization path and the unknown function to be learned. Especially, for independent and non-identically distributed data streams, the mean square and almost sure consistency between the algorithm's output and the unknown function is achieved if the input data's marginal probability measures are slowly time-varying and the average measure over each fixed-length time period is uniformly above a strictly positive finite Borel measure.

cs.LG

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models

Direct Preference Optimization (DPO) improves relative preference by increasing the margin between chosen and rejected responses, but this objective does not specify how probability mass should be redistributed to achieve that margin. Consequently, preference separation can improve even as probability mass collapses away from both preferred and unrelated responses, a phenomenon associated with the squeezing effect. We show that this pathology is localized: continued negative updates become destructive when rejected responses enter low-probability valleys, where further suppression induces harmful redistribution through the coupled softmax geometry. We introduce Gradient-Gated Preference Optimization (Gate-DPO), which uses the current probability geometry of each rejected response to selectively attenuate its contribution, limiting excessive rejected suppression while delaying preference-loss saturation. Gate-DPO is objective-agnostic and composes with existing preference methods. Across four architectures, two preference datasets, and multiple objectives, gating consistently reduces squeezing and improves chosen-response likelihood, while exhibiting robust reductions in rejected-response collapse across architectures and objectives. Mass-dynamics analysis further shows that large preference margins can conceal substantial suppression of unrelated responses, while gating achieves healthier redistribution. Together, our results show that preference margin alone is insufficient to characterize optimization health and that selectively controlling probability dynamics provides a simple, general mechanism for stabilizing preference optimization.

cs.LG

Noise-Debiased Thermodynamic Variance for Local Learning Coefficient Probes

Local learning coefficient (LLC) probes offer a singularity-aware view of neural-network training, but mean-energy methods require a local loss baseline that is ambiguous at transient checkpoints. Thermodynamic variance avoids this input; under mini-batch evaluation, however, direct variance mixes cross-state loss fluctuations with same-state noise. We operationalize this route with the \emph{Shift-Invariant Variance Estimator} (SIVE), which estimates and subtracts the latter component using repeated evaluations. Conditional on any fixed retained path, unclipped SIVE is unbiased for noiseless path variance without requiring MCMC stationarity. The finite-scale diagnostic remains indexed by localization scale $h$---even a locally linear loss has tether-dependent variance---while interpretation as a Real Log Canonical Threshold (RLCT) requires additional stationary low-temperature conditions. Toy experiments recover calibrated finite-scale targets. At the primary localization scale, all five MNIST MLP trajectories exhibit a mid-training trough followed by a rebound in SIVE, while Raw Variance decreases from Epoch 40 to 100 in every trajectory. Across four localization scales, the joint early-drop/late-rise criterion is met in 19 of 20 trajectory--scale pairs. At Epoch 40, the estimated observation-noise correction accounts for $77.5\%$ of Raw Variance. Same-state debiasing thus reveals a reproducible turning structure masked by time-varying observation noise.

cs.LG