Abstract
Compressing large-scale neural networks is essential for deploying models on resource-constrained devices. Most existing methods adopt weight pruning or low-bit quantization individually, often resulting in suboptimal compression rates to preserve acceptable performance drops. We introduce a unified framework for simultaneous pruning and low-bit quantization via Bayesian variational learning (SQS), which achieves higher compression rates than prior baselines while maintaining comparable performance. The key idea is to employ a spike-and-slab prior to induce sparsity and model quantized weights using Gaussian Mixture Models (GMMs) to enable low-bit precision. Due to the intractability of the objective involving spike-and-slab priors with GMMs, we derive an efficient approximation that facilitates effective compression with minimal accuracy loss. In theory, we provide a consistent result for our proposed variational approach to a sparse and quantized deep neural network. Extensive experiments on compressing ResNet, BERT-base, Llama3.2, and Qwen2.5 models show that our method achieves higher compression rates than a line of existing methods with comparable performance drops.
Project page: https://comeusr.github.io/SQS_Webpage/.
1 Introduction
Deep Neural Networks (DNNs) have achieved state-of-the-art performance across a wide range of tasks but at the cost of significantly increased computational and memory requirements (Radford et al., 2018; Xu et al., 2020; Touvron et al., 2023; Kumar et al., 2025), making deployment on resource-constrained devices challenging. Model compression methods have therefore been proposed to reduce the size and computational complexity of DNNs while maintaining predictive accuracy, including pruning (LeCun et al., 1989; Han et al., 2016), weight quantization (Courbariaux et al., 2015; Rastegari et al., 2016; Frantar et al., 2023; Lin et al., 2024), knowledge distillation (Park et al., 2019; Gou et al., 2021), and neural architecture search (Liu et al., 2018; Wang et al., 2020b).
Among these, weight pruning and low-bit quantization are particularly effective and widely adopted for compressing DNNs (Buciluǎ et al., 2006; Choudhary et al., 2020; Liu et al., 2025a). Weight pruning eliminates redundant or unimportant weights by setting selected weights to zero, thereby reducing the number of active parameters without significantly altering the model architecture (You et al., 2019; Guo et al., 2016; Dong et al., 2017). On the other hand, quantization reduces the bit-width of numerical representations for inputs, outputs, and weights by converting high-precision formats (e.g., FP32) to lower-precision alternatives, such as FP8 or INT8. This quantization coarsens the model representation and yields significant reductions in memory footprint and computational overhead. It enhances efficiency in both training and inference across diverse architectures, including ResNet (Banner et al., 2018), Transformers (Sun et al., 2019), Large language models (Dettmers et al., 2023; Wang et al., 2025), and vision-language models (Wortsman et al., 2023).
However, quantization and pruning inevitably introduce distributional shifts from the original DNNs, often leading to performance degradation (Dong et al., 2022). To mitigate this, existing methods adopt conservative compression rates, limiting their applicability to resource-constrained environments (Wang et al., 2020c; Wang et al., 2020b; Bai et al., 2022; Frantar et al., 2022; Bai et al., 2023). Achieving high compression rates while maintaining acceptable performance remains an open question to explore.
To tackle the above problem, we introduce a unified framework: Sparse Quantized Sub-distribution compression (SQS), which unifies pruning and quantization within a single variational learning process. Instead of applying pruning and quantization separately, the key idea of SQS is joint pruning and quantization that learns a sparse, quantized sub-distribution over network weights through variational learning. To model the variational posterior, we adopt a spike-and-slab prior combined with a Gaussian Mixture Model (GMM): the spike component encourages sparsity for pruning, while the GMM component models a quantized weight distribution, effectively mitigating performance degradation. The training pipeline of SQS is in Figure 1. Theoretically, we show that under mild conditions, our SQS method finds a sparse and quantized neural network that converges to the true underlying target neural network with high probability.
In our experiments, we compare several recent state-of-the-art compression methods across a range of widely used neural networks, including ResNet, BERT-base, Llama3.2, and Qwen2.5. Our findings show that (1) under the same bit-width setting, our SQS achieves the highest compression rate, requiring fewer parameters than baselines. (2) At the same compression rate, our SQS achieves the smallest accuracy drop or F1 score drop among all approaches, with particularly strong performance at 2-bit and 4-bit precision.
Further ablation studies highlight the contributions of individual components: (1) The spike-and-slab distribution is more effective in promoting sparsity than Gaussian alternatives. (2) Bayesian averaging during inference outperforms greedy weight selection. (3) An outlier-aware window strategy better preserves informative weight outliers compared to uniform windowing, further improving performance.
We propose SQS, a unified Bayesian framework for compressing full-precision DNNs into sparse, low-bit models. Unlike methods that perform pruning and quantization separately, SQS jointly learns which weights to remove and how to quantize the remaining weights using a spike-and-GMM variational distribution.
We derive a tractable approximate objective for training SQS and provide theoretical guarantees showing that, under mild conditions, the learned sparse and quantized network converges to the target regression function.
We conduct experiments on ResNet, BERT-base, Llama3.2, and Qwen2.5. Across these architectures, SQS achieves higher compression rates with comparable or smaller accuracy degradation than existing baselines. Ablation studies further validate the effectiveness of the spike-and-slab formulation, Bayesian averaging at inference time, and the outlier-aware windowing strategy.
2 Preliminaries
Low-bit Quantization uses discrete low-bit values to approximate full-precision floating-point values, primarily to reduce precision for more efficient storage and computation while preserving essential information (Gholami et al., 2022). Formally, it is defined as a mapping , where the input is the full-precision weight and denotes the set of low-bit discrete values. Representative quantization methods include deterministic quantization (Jacob et al., 2018), stochastic quantization (Courbariaux et al., 2015), and end-to-end learnable quantization (Dong et al., 2022).
Specifically, let represent the pre-trained full-precision weights of a deep neural network, with denoting the -th weight. Given a quantization set , a general stochastic quantization is a map from the real numbers to the space of probability distributions over with finite support of cardinality . For each weight ,
for . Here is the learnable parameter and is the corresponding probability that weight is quantized to weight . A key challenge is the distribution divergence between the quantized weights and the original weights, leading to significant performance degradation (Dong et al., 2022). To mitigate this, Dong et al. (2022) propose to approximate the quantized weight distribution using a Gaussian Mixture Model (GMM):
| (1) |
where denotes a Gaussian distribution, and is the weight for the -th Gaussian component . To control the sharpness of this mixture, a temperature-scaled softmax is applied to obtain , that is:
| (2) |
where the temperature parameter controls the concentration of the distribution. As , the GMM in Equation (1) approaches a single dominant Gaussian component. Given a prior distribution over the quantization set , the posterior component weight is:
Additionally, with sufficiently small , the GMM approximates a multinomial distribution over , effectively bridging continuous and discrete quantization. For simplicity, we denote as a shorthand for throughout the remainder of this paper.
In our experiments, we find that the GMM-based compression method (Dong et al., 2022) still cannot achieve a high compression rate while maintaining a small performance drop, as it cannot efficiently encourage sparsity during training.
Variational learning. Given an observed dataset , the goal of a Bayesian framework is to infer the true posterior distribution , where denotes the prior and the likelihood. Since the posterior is generally intractable, variational learning (Jordan et al., 1999) is proposed to approximate it by selecting the closest distribution from a variational family in terms of the Kullback–Leibler (KL) divergence (Csiszar, 1975):
| (3) |
Following (Blei et al., 2017), this optimization is equivalent to minimizing the negative Evidence Lower Bound (ELBO), defined as:
| (4) |
where the first term measures how well the variational distribution aligns with the log-likelihood of the observed data, and the second term regularizes to stay close to the prior .
Our SQS method employs a variational family based on a spike-and-GMM distribution to approximate the sparse and quantized posterior. The first term in Equation (4) allows the spike-and-GMM to learn the posterior distribution given the data. For the second term, we adopt a spike-and-slab prior distribution to promote sparsity in the network weights.
3 Methodology
The objective is to approximate a full-precision neural network with a Bayesian model that is both sparse and low-precision, while minimizing performance degradation. To achieve this, we employ a spike-and-slab distribution combined with a GMM to parameterize the variational posterior.
3.1 SQS: Variational learning for sparse and quantized sub-distribution
The spike-and-slab prior consists of a point mass at zero (spike) and a continuous distribution (slab) (Bai et al., 2020; Ishwaran and Rao, 2005). Formally, let be a binary indicator vector, where each determines whether the corresponding weight is preserved () or pruned (). The prior for each weight is defined as:
where is the prior probability of retaining a weight, and is the prior variance of the Gaussian slab. Marginalizing out the binary variable , the prior distribution over becomes:
| (5) |
where corresponds to the prior pruning probability. For example, in a DNN with a target sparsity of , setting implies that each weight has a prior probability of being pruned.
3.1.1 Training Procedure
To incorporate quantization into the variational family, we extend the spike-and-slab formulation by modeling the slab using a -component GMM. Each variational distribution is then defined as:
where is the mixture weight for component , and is the variational probability of retaining weight . The marginal variational distribution is:
| (6) |
Given this variational family, we define the learning objective based on the ELBO:
| (7) |
Yet, computing Equation (7) is intractable, as no closed-form solution exists for the KL divergence between and the spike-and-slab prior . To overcome this challenge, we propose the following approximation: For each coordinate , the posterior mean is . Collecting these coordinate-wise means, define
The approximate objective becomes
| (8) |
where . The first term uses the plug-in approximation Thus, the likelihood is evaluated using the complete parameter vector . The second and the third terms provide an upper bound on the term by applying Lemma 3. Please refer to Appendix A for a detailed derivation of Equation (8).
3.1.2 Inference Procedure
In the inference stage, we first sample the sparse and quantized weights given the learned parameters and predict the output for each testing input . Let denote the optimization solution of the above variational learning, associated with the optimal parameter estimations , and the corresponding ’s (for each ) are obtained from Equation (2). Then, the -th quantized weight is sampled from the set of quantization levels according to
| (9) |
Compared to sampling from , posterior sampling in Equation (9) reduces memory consumption.
To enforce sparsity, we introduce a user-specified pruning parameter, the Non-zero rate. Each weight is associated with a score , which reflects the likelihood of being retained. We deterministically prune by setting the -th weight to zero if is smaller than the Non-zero-quantile of all values; otherwise, the weight is kept unchanged. Formally,
This deterministic rule provides exact control over the sparsity level, in contrast to stochastic pruning via posterior sampling (Bai et al., 2020; Sun et al., 2022), which does not guarantee a fixed sparsity rate and often requires an additional pruning step.
Bayesian averaging. Given a test input , the predicted output is computed using Bayesian averaging:
| (10) |
where ’s are many samples from the sparse quantized sub-distribution. In the following experiments, we set the default to 4. Our ablation study (in Figure 3) shows that Bayesian averaging consistently yields smaller accuracy degradation than the greedy alternative (detailed in Equation 11).
Greedy approach is to greedily select the most likely weight for making predictions on the test set. Specifically, for each quantized weight , we choose the index corresponding to the highest posterior probability . The quantized weight is then set to the mean of the selected component, and the predicted output is computed using these selected means. Formally, this greedy inference strategy is given by:
| (11) |
We empirically compare the greedy inference approach with Bayesian averaging in the ablation study shown in Figure 3.
Outlier-aware windowing. Recent studies show that the weight distribution of large language models (LLMs) often contains significant outliers (Wei et al., 2022). To address this, we use an outlier-aware windowing strategy to enhance the performance of SQS. Specifically, the full-precision weights are partitioned into four groups using window sizes determined by a modified interquartile range (IQR) rule (Dekking et al., 2006), which helps preserve large-magnitude weights during quantization. Each group is then quantized to representative values. As shown in the ablation study (Figure 2), this strategy outperforms the approach using equal-sized windows. Implementation details are provided in Appendix C, and the full procedure is summarized in Algorithm 1.
Remarks. DGMS (Dong et al., 2022) adopts Gaussian mixtures, but uses them primarily as a clustering mechanism. In contrast, our method leverages a principled Bayesian framework that supports posterior inference and enables Bayesian model averaging, enhancing robustness to quantization noise. Furthermore, by unifying pruning and quantization within a spike-and-GMM variational family, our approach creates a joint optimization space that encourages globally optimal solutions across both pruning and quantization.
3.1.3 Windowing strategy in quantization
We observe that weight distributions vary significantly across layers, including Gaussian and long-tailed forms. In particular, long-tailed distributions contain a small subset of weights with large magnitudes. Previous works (Nagel et al., 2020; Hubara et al., 2021; Frantar et al., 2022) have demonstrated that layer-wise compression methods lead to better performance.
To address performance degradation arising from such heterogeneous distributions, we extend our proposed method to support layer-wise quantization, where each group of weight parameters within a layer is assigned its own quantization set. This enables each layer to learn and utilize a distinct, trainable quantization set tailored to its distribution.
Equal-size windowing. For the equal window strategy, given a layer of weights , we group the weights into 4 windows where each one has an equal window size . Within each window, a -component GMM is applied to approximate the weight distribution.
Outlier-aware windowing. For layers with long-tailed distributions, we further introduce an outlier-aware windowing strategy. Specifically, the weights in each layer are partitioned into four windows, with two dedicated to capturing the lower and upper tails of the distribution. To identify these tail regions, we apply a standard outlier detection rule based on the inter quartile range (IQR): let and denote the first and third quartiles of the weights, and define . The outlier-aware windows are then defined as
| (12) |
Within each of the four windows in every layer, we fit a -component GMM to approximate the local weight distribution.
We adopt the layer-wise quantization scheme with outlier-aware windowing in all our experiments. This approach improves the preservation of extreme values during quantization and enhances robustness across layers. An ablation study evaluating the effectiveness of outlier-aware windowing is presented in Figure 2.
3.2 Theoretical Justification of SQS
For clarity, this section focuses on regression tasks with fully connected neural networks. We analyze the variational posterior of sparse and quantized neural networks, i.e., the optimization of Equation (7). We show that this variational posterior converges to a true regression function under some mild conditions.
Consider a regression problem with random covariates,
| (13) |
where is the underlying unknown true function, is sampled from a -dimensional uniform distribution, is the noise term from a Gaussian distribution of zero mean and variance . Let denote the true underlying probability measure of the data, and denote the corresponding density function. An -hidden-layer fully connected NN with constant layer width and parameters , and activation function can be defined as:
| (14) |
For simplicity, is assumed to be known. Let be the “oracle” sparsity level (see Equation 19 in Appendix B for formal definition) and be the set of network weight parameters such that the network has a sparsity of and shares at most distinct values. Let and be the true data distribution and the distribution under parameter , respectively.
Theorem 1 carries a proof sketch stating the two-step structure: Lemma 1 upper-bounds the variational objective with high probability; Lemma 2 converts that bound into convergence of the variational posterior in squared Hellinger distance.
Lemma 1.
Under Conditions 1-3, with high probability,
where is either some positive constant if , or any diverging sequence if . And is defined as:
For any , let .
Lemma 2.
Under Conditions 1-5, if is set to be constant and for any positive diverging sequence , then with high probability, then we have
| (15) |
where is some constant, and
We refer to Appendix B.1 for the proof of Lemma 1 and Appendix B.2 for the proof of Lemma 2.
Theorem 1.
Let , for any from Lemma 2, and . Then, under mild conditions specified in the supplementary material, with high probability:
| (16) |
where denotes the Hellinger distance, and and are some constants.
Sketch of Proof.
Based on prior work (Bai et al., 2020), the proof proceeds in two steps. Lemma 1 establishes a high-probability bound on the ELBO in Equation (7). Lemma 2 connects this bound to the convergence of the variational distribution toward the true full-precision posterior. Together, these results show that the variational posterior induced by our method converges to the true regression function with high probability. The full proof is in Appendix B. ∎
Remark. Similar to previous Bayesian sparse DNN results (Bai et al., 2020; Chérief-Abdellatif, 2020), the convergence rate of variational Bayes is determined by the deep neural network structure via 1) statistical estimation error , 2) variational error , and 3) approximation error . The first two are positively related to the network capacity, while the third one is negatively related to the network capacity. The estimation error and variational error vanish as . Prior work (Beknazaryan, 2022) shows that under , , and -Hölder smoothness of , the approximation error also vanishes.
While the theoretical analysis mainly considers an -hidden-layer fully connected NN with constant layer width , our method SQS is empirically validated on a variety of models such as ResNets, BERT-based models, and LLMs (refer to Section 5).
4 Related Work
Weight pruning was initially introduced by LeCun et al. (1989), with further development by Hassibi et al. (1993) through a mathematical method known as the Optimal Brain Surgeon (OBS). This approach selects weights for removal from a trained neural network using second-order information. Subsequent improvements, as indicated by studies (Dong et al., 2017; Wang et al., 2019; Singh and Alistarh, 2020), have adapted OBS for large-scale DNNs by employing numerical techniques to estimate the second-order information required by OBS. Meanwhile, Louizos et al. (2018) introduced an -regularized method to promote sparsity in DNNs. Frankle and Carbin (2019) established a critical insight that within a randomly initialized DNN, an optimal sub-network can be identified and extracted. Recently, Xia et al. (2024) showed that structured pruning combined with targeted retraining can significantly reduce computational costs while preserving robust performance for large language models. Concurrently, spike-and-slab distributions have been employed to promote sparsity in DNNs using Bayesian Neural Networks formulation (Deng et al., 2019; Blundell et al., 2015; Bai et al., 2020).
Low-bit quantization. Quantization improves DNN efficiency, particularly in resource-constrained environments (Sze et al., 2017; Frantar et al., 2023; Lin et al., 2024; Lin et al., 2025). Research in this field typically follows two paradigms: discontinuous-mapping and continuous-mapping quantization. Discontinuous-mapping methods project full-precision weights onto a low-bit grid using rounding operations (Gupta et al., 2015; Hubara et al., 2018; Wu et al., 2018; Louizos et al., 2019; Courbariaux et al., 2015; De Sa et al., 2018; Marchesi et al., 1993). The non-differentiability of these mappings necessitates the use of the straight-through estimator (STE) for gradient approximation (Courbariaux and Bengio, 2016; Rastegari et al., 2016). However, STE-based training may introduce pseudo-gradients, leading to training instability (Yin et al., 2019). Meanwhile, many researchers propose post-training quantization methods that have limited access to the training dataset (Wang et al., 2020a; Hubara et al., 2021; Li et al., 2021; Frantar et al., 2022; Frantar et al., 2023; Lin et al., 2024).
Continuous-mapping quantization offers an alternative that avoids pseudo-gradients, leading to more stable training (Yin et al., 2019; Nielsen et al., 2025). These methods often use variational learning (Ullrich et al., 2017; Louizos et al., 2017; Shayer et al., 2018) or Markov Chain Monte Carlo techniques (Roth and Pernkopf, 2018) to approximate discrete weight distributions. However, variational methods often require manual prior specification (Ullrich et al., 2017; Louizos et al., 2017; Shayer et al., 2018), while MCMC approaches can be memory-intensive (Roth and Pernkopf, 2018). DGMS (Dong et al., 2022) addresses these limitations through automated quantization using GMMs. Our work extends DGMS by integrating pruning and quantization into a unified framework, thereby achieving higher compression rates.
Joint pruning and quantization. A growing line of work optimizes sparsity and precision jointly: Bayesian formulations derive both from a single prior or posterior (Ullrich et al., 2017; Louizos et al., 2017; Achterhold et al., 2018), while non-Bayesian approaches rely on differentiable gates, second-order saliency, or joint policy search (Wang et al., 2020c; Frantar et al., 2022; Wang et al., 2020b; Bai et al., 2023). Bayesian Bits (Van Baalen et al., 2020) unifies pruning and quantization by gating a chain of residual terms that double the bit-width, with pruning as the -bit case. Unlike its gates on a uniform grid, SQS learns the levels themselves as GMM means with a per-weight retention probability, giving a non-uniform codebook, exact sparsity control, and a posterior for Bayesian averaging.
Large language model compression. Recent work on compressing large models has pursued several complementary directions. SpinQuant (Liu et al., 2025b) applies learned rotations to weights and embeddings to reduce outliers, making models easier to quantize. LeanQuant (Zhang and Shrivastava, 2025) introduces a loss-aware post-training quantization method that learns adaptive affine transformations and non-uniform quantization grids, aiming to preserve outlier-sensitive weights. EfficientXpert (Zhao et al., 2025) uses LoRA-guided pruning to obtain domain-aware pruned models. Xu et al. (2025) targeted vision–language models and adaptively pruned attention heads using entropy-based effective rank and the Kolmogorov–Smirnov distance, reporting substantial FLOP reductions. In contrast, our SQS introduces a learnable codebook and a spike-and-slab posterior tailored to sparse, quantized weights in the low-precision regime.
5 Experiments
In this section, we show that our SQS achieves a much higher compression rate (see the second-to-last column in Tables 1-3) while incurring a comparable or smaller performance drop (see the last column in the same tables). Through ablation studies, we further validate that (1) under the same sparsity level, the spike-and-slab prior more effectively preserves model accuracy (see Table 4). (2) Under identical hyperparameter settings, we show that SQS with Bayesian averaging outperforms the greedy approach (see Figure 3).
5.1 Experiment settings
We evaluate all methods using two metrics: the compression rate and the performance drop (i.e., accuracy drop or F1 score drop). The compression rate is defined as the memory footprint of the compressed model over the original dense full-precision model:
| (17) |
where is the codebook size. Non-zero rate is the percentage of weights that are pruned to zero. In all our experiments, the Non-zero rate is configured as a hyperparameter to control the sparsity. Equation (17) accounts for the codebook and the quantized value indices assigned to the nonzero weights, but, following the convention adopted by the pruning baselines we compare against (e.g., L-OBS (Dong et al., 2017), PLATON (Zhang et al., 2022) and ExactOBS/OBC (Frantar et al., 2022)), it does not separately charge bits for the binary sparsity mask that records which weights are pruned. We adopt this convention so that the reported compression rates in Tables 1–3 remain directly comparable to the values reported for these baselines in their original papers, all of which similarly define compression as a function of remaining/quantized parameter count alone. We note that some deployment-oriented compression pipelines, such as (Han et al., 2016; Dettmers et al., 2024), instead charge explicit bits for the sparsity structure (e.g., via a bitmap or compressed-sparse-row index); under that stricter accounting, absolute compression rates for all pruning-based methods, including ours, would be lower, though the relative ranking among methods is unaffected. In practice, for the models we study, the product “ nonzero weight counts” is usually much larger than , because the number of nonzero weights is very large and the codebook size is small. In this regime, the memory used to store the indices dominates, and the memory used by the codebook is very small.
We compare methods of different compression types (the “Compression type” column): “P+Q” denotes combined pruning and quantization, “P” denotes pruning only, and “Q” denotes quantization only.
In our experiments, we store the indices of the weights using INT4, and we perform computation in FP32. This setting is common in prior work, such as QLoRA (Dettmers et al., 2023) and learned codebook methods (van den Oord et al., 2017; Dong et al., 2022).
To ensure a fair comparison, each method is initialized with the same full-precision pre-trained model and is run with the same set of hyperparameters for compression. All methods are constrained to a maximum runtime of 24 hours. The resulting compressed models are then evaluated on the same test sets, and the key performance metrics are summarized in the corresponding tables. Appendix D provides detailed experimental configurations and baseline settings.
5.2 Experimental analysis
Compression on ResNet models. Table 1 summarizes the result of all methods for compressing ResNet-20, ResNet-32 and ResNet-56, evaluated on the CIFAR-10 dataset. On compressing the ResNet-20 model, our SQS attains a better compression rate than the baselines. On compressing ResNet-32 and ResNet-56 models, our SQS attains substantially higher compression rates while incurring smaller accuracy drops compared to the baselines. Optimizing pruning and quantization separately overlooks redundancies in each step; by merging them into a single optimization, we effectively eliminate these inefficiencies.
| ResNet-20 | Methods | Compression | Bits | Non-zero rate | Compression | Top-1 accuracy |
| type | (%) | rate | drop | |||
| LQNets (Zhang et al., 2018) | Q | |||||
| DGMS (Dong et al., 2022) | P+Q | |||||
| SQS (Ours) | P+Q | |||||
| (a) Compressing 32Bits ResNet-20 model on CIFAR-10 dataset with Top-1 accuracy . | ||||||
| ResNet-32 | Method | Compression | Bits | Non-zero rate | Compression | Top-1 accuracy |
| type | (%) | rate | drop | |||
| TTQ (Zhu et al., 2017) | Q | |||||
| DGMS (Dong et al., 2022) | P+Q | |||||
| SQS (Ours) | P+Q | 32 | ||||
| (b) Compressing 32Bits ResNet-32 model on CIFAR-10 dataset with Top-1 accuracy . | ||||||
| ResNet-56 | Method | Compression | Bits | Non-zero rate | Compression | Top-1 accuracy |
| type | (%) | rate | drop | |||
| TTQ (Zhu et al., 2017) | Q | 2 | 16 | |||
| L1 (Li et al., 2017) | P | 32 | 10 | |||
| DGMS (Dong et al., 2022) | P+Q | |||||
| SQS (Ours) | P+Q | 2 | ||||
| (c) Compressing 32Bits ResNet-56 model on CIFAR-10 dataset with Top-1 accuracy . | ||||||
| BERT-base | Methods | Compression | Bits | Non-zero rate | Compression | F1 score |
|---|---|---|---|---|---|---|
| type | (%) | rate | drop | |||
| GMP (Zhu and Gupta, 2017) | P | |||||
| L-OBS (Dong et al., 2017) | P | |||||
| ExactOBS (Frantar et al., 2022) | P | |||||
| PLATON (Zhang et al., 2022) | P | |||||
| OBQ (Frantar et al., 2022) | Q | |||||
| GPTQ (Frantar et al., 2023) | Q | |||||
| OBC (Frantar et al., 2022) | P+Q | |||||
| SQS (Ours) | P+Q | 4 |
Compression on BERT-base model. We apply our compression method to the BERT-base model (Devlin et al., 2019) and evaluate its performance on the SQuAD v1.1 dataset (Rajpurkar et al., 2016). The evaluation metrics include the F1 score drop and the compression rate. As shown in Table 2, our method achieves the lowest F1 score drop and the highest compression rate, outperforming existing methods. This demonstrates the effectiveness of our SQS method in preserving accuracy under aggressive compression.
| Llama3.2 | Method | Compression | Bits | Non-zero rate | Compression | Top-1 accuracy |
| type | (%) | rate | drop | |||
| AWQ (Lin et al., 2024) | Q | |||||
| DGMS (Dong et al., 2022) | P+Q | |||||
| SQS (Ours) | P+Q | |||||
| (a) Compressing 32Bits Llama3.2-1B model on SST-2 dataset with Top-1 accuracy . | ||||||
| Qwen2.5 | Method | Compression | Bits | Non-zero rate | Compression | Top-1 accuracy |
| type | (%) | rate | drop | |||
| AWQ (Lin et al., 2024) | Q | 4 | 100% | |||
| DGMS (Dong et al., 2022) | P+Q | |||||
| SQS (Ours) | P+Q | |||||
| (b) Compressing 32Bits Qwen2.5-0.5B model on SST-2 dataset with Top-1 accuracy . | ||||||
Compression on Llama and Qwen models. In Table 3, we compare our SQS with others on the SST-2 task in the GLUE benchmark using Llama3.2-1B and Qwen2.5-0.5B models, as our method could preserve the weight outliers, which are crucial in maintaining the performance (Lin et al., 2024). We further observe that DGMS (Dong et al., 2022) incurs a large performance drop for compressing Llama3.2-1B and Qwen2.5-0.5B, which occurs because the weight distribution of the self-attention layer is not Gaussian (see Figure 2(left)), and it fails to capture large magnitude weights. Furthermore, DGMS does not allow customization of the sparsity level; thus, it presents an unreasonable performance drop.
| ResNet-18 | Bits | Non-zero rate (%) | Compression rate | Top-1 accuracy drop | |
|---|---|---|---|---|---|
| Gaussian prior | Spike-and-slab prior (Ours) | ||||
5.3 Ablation studies
Prior selection: Gaussian vs. spike-and-slab. We evaluate how the choice of prior affects compression performance, comparing a Gaussian prior (Appendix D.2) to the spike-and-slab prior in Equation (5). Specifically, we compress ResNet-18 at varying sparsity levels by representing each layer’s weights with components, and evaluate accuracy on CIFAR-100.
As shown in Table 4, the spike-and-slab prior consistently outperforms the Gaussian prior across all sparsity levels. The gap becomes particularly pronounced at high sparsity, where the Gaussian prior suffers substantial degradation, suggesting that it may be less effective at inducing posterior sparsity in DNN weights. We leave a deeper investigation of this behavior to future work.
Windowing strategy: equal-size vs. outlier-aware window. We use the first-layer attention weights of Llama3.2-1B as a case study. The full-precision weights exhibit a pronounced long-tail distribution, where a small fraction of entries have large magnitudes. Additional statistics on layer-wise weight distributions are reported in Appendix D.3. Motivated by this observation, our SQS uses an outlier-aware window strategy to better fit the full-precision weight distribution in Equation (12). Figure 2 visualizes the resulting quantized weights under the equal-window and outlier-aware window strategies.
| Qwen2.5 | Method | Windowing strategy | Bits | Non-zero rate (%) | Top-1 accuracy drop |
|---|---|---|---|---|---|
| SQS | Outlier-aware window | ||||
| Equal-size window |
As shown in Table 5, outlier-aware windowing reduces the Top-1 accuracy drop on Qwen2.5-0.5B from 5.40 to 2.46 percentage points at the same bit width and nonzero rate. The equal-size strategy spreads its windows uniformly over the weight range, so the few large-magnitude entries in the long tail fall into wide windows and are coarsely quantized, which is precisely the tail region that Figure 2 shows to be distorted. Preserving these outlier weights is therefore a primary driver of the accuracy that SQS retains on heavy-tailed LLM weights.
As shown in Figure 2 (left, middle), the quantized weights produced by SQS under the outlier-aware window strategy more closely match the full-precision distribution than those obtained with an equal-window strategy. Figure 2 (right) further highlights that the outlier-aware window improves fidelity in the tails, capturing extreme-magnitude weights more accurately than equal windowing.
Inference strategies: Bayesian averaging vs. greedy approach. We evaluate the effectiveness of two inference strategies within our SQS framework: (1) Bayesian averaging as defined in Equation (10), and (2) greedy approach as defined in Equation (11). To ensure a fair comparison, we assess the performance of compressing ResNet-18 and ResNet-50 models while varying the number of Gaussian components. The sparsity level is fixed to zero (i.e., no pruning), so that all performance degradation arises purely from quantization.
As shown in Figure 3, using fewer components results in a larger accuracy drop. Under the same number of components, SQS with Bayesian averaging consistently achieves a smaller accuracy drop compared to the greedy approach.
Effect of the number of Bayesian-averaging samples. As shown in Table 6, the number of averaged posterior samples directly affects the robustness of the compressed model: a single posterior sample () incurs a Top-1 accuracy drop, whereas averaging over samples reduces the drop to . Most of this benefit is realized with only a few samples, as already attains a drop and further increasing yields only marginal improvement. This indicates that a small number of posterior samples suffices in practice, and that Bayesian averaging mainly improves robustness to quantization noise rather than providing a large accuracy gain.
| Inference strategy | ||||
| Top-1 accuracy Drop (%) |
| Method | Compression | Weight bits | Non-zero | Effective | Compression | Top-1 accuracy |
|---|---|---|---|---|---|---|
| type | rate (%) | bits/weight | rate | drop | ||
| Bayesian Bits | P+Q | (mixed) | ||||
| SQS (Ours) | P+Q |
Comparison with Bayesian Bits. In Table 7, effective bits per weight is the average storage cost per weight of the original dense model. We define , where is the original dense model’s weight count and is the total bit cost counted for the compressed weight representation in this comparison. Relative to an FP32 baseline, the corresponding compression rate is .
We use the official Bayesian Bits implementation released by Qualcomm AI Research11 1 https://github.com/Qualcomm-AI-research/BayesianBits and follow the published CIFAR-10 recipe without modification. We use a batch size of and Adam to optimize the weights, quantization, and gate parameters. Bayesian Bits is trained from scratch for epochs for hours, whereas SQS is initialized from a pre-trained model and trained for epochs.
We compare Bayesian Bits with SQS on ResNet-56/CIFAR-10, for which the full-precision model achieves Top-1 accuracy. As shown in Table 7, SQS achieves higher Top-1 accuracy and greater compression with fewer effective bits per weight.
6 Conclusion
In this paper, we proposed a unified framework for compressing full-precision DNNs by combining pruning and quantization into one integrated optimization process through variational learning. Unlike conventional approaches that apply pruning and quantization sequentially—often resulting in suboptimal solutions—our method jointly explores a broader solution space, achieving significantly higher compression rates with comparable performance degradation. To address the intractability of the original objective, we introduce an efficient approximation that enables scalable optimization. We evaluate our method across a range of benchmarks, including ResNets, BERT-base, Llama3.2, and Qwen2.5. Experimental results demonstrate that our approach consistently outperforms existing baselines in compression rate while maintaining competitive accuracy, highlighting its potential for efficient deployment in resource-constrained environments.
Broader Impact
Model compression can reduce memory use and may lower the cost of deploying DNNs on resource-constrained hardware. These benefits should be weighed against training-to-deployment trade-offs. Our method requires additional optimization after pretraining, which can increase training time and energy use before deployment. Inference also involves a design choice: Bayesian averaging can improve robustness to quantization noise, but using multiple posterior samples may increase latency compared with a single compressed model. Practical deployments should therefore measure end-to-end latency, memory use, energy consumption, and accuracy under the target hardware and workload.
SQS does not change the task or data distribution of the original model, so its safety risks largely inherit those of the base model. Compression may make models easier to deploy more widely, including in settings where monitoring is limited. For applications involving sensitive, high-stakes, or user-facing decisions, compressed models should be evaluated for robustness, fairness, privacy leakage, and failure modes after compression, rather than assuming that performance on aggregate benchmarks is sufficient.
Limitations
SQS is a compression-aware training method that optimizes its objective using task-specific training data. In our LLM experiments, the base Llama3.2-1B and Qwen2.5-0.5B models are first fine-tuned on SST-2 before compression; omitting this task-adaptation step leads to substantial performance degradation. Consequently, the reported results characterize the compression of task-adapted models rather than the preservation of the models’ general-purpose capabilities. Evaluating whether SQS maintains performance on unrelated tasks or under distribution shifts is beyond the scope of this study. In addition, our theoretical analysis is currently restricted to regression problems with fully connected neural networks and does not directly cover transformer architectures, classification settings, or other model families.
AI Use Statement
The authors used AI to assist with proofreading, grammar and mathematical-notation consistency checks, and the identification of potential derivation, citation, and presentation issues. All suggestions were reviewed and verified by the authors, who take full responsibility for the accuracy, integrity, and content of the manuscript.
References
- Achterhold et al. (2018) J. Achterhold, J. M. Koehler, A. Schmeink, and T. Genewein Variational network quantization. In ICLR, Cited by: §4.
- Bai et al. (2020) J. Bai, Q. Song, and G. Cheng Efficient variational inference for sparse deep learning with theoretical guarantee. In NeurIPS, Vol. 33, pp. 466–476. Cited by: §B.2, §B.2, §B.2, Appendix B, Appendix B, §3.1.2, §3.1, §3.2, §4, Theorem 1.
- Bai et al. (2023) S. Bai, J. Chen, X. Shen, Y. Qian, and Y. Liu Unified data-free compression: pruning and quantization without fine-tuning. In ICCV, pp. 5876–5885. Cited by: §1, §4.
- Bai et al. (2022) Y. Bai, H. Wang, Z. Tao, K. Li, and Y. Fu Dual lottery ticket hypothesis. In ICLR, Cited by: §1.
- Banner et al. (2018) R. Banner, I. Hubara, E. Hoffer, and D. Soudry Scalable methods for 8-bit training of neural networks. In NeurIPS, Vol. 31, pp. 5151–5159. Cited by: §1.
- Beknazaryan (2022) A. Beknazaryan Function approximation by deep neural networks with parameters {0, ½, 1, 2}. Journal of Statistical Theory and Practice 16 (1), pp. 7. Cited by: §3.2.
- Blei et al. (2017) D. M. Blei, A. Kucukelbir, and J. D. McAuliffe Variational inference: a review for statisticians. Journal of the American statistical Association 112 (518), pp. 859–877. Cited by: §2.
- Blundell et al. (2015) C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra Weight uncertainty in neural network. In ICML, Vol. 37, pp. 1613–1622. Cited by: §4.
- Boucheron et al. (2013) S. Boucheron, G. Lugosi, and P. Massart Concentration inequalities: a nonasymptotic theory of independence. Oxford University press. Cited by: §B.2.
- Buciluǎ et al. (2006) C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil Model compression. In KDD, pp. 535–541. Cited by: §1.
- Chérief-Abdellatif and Alquier (2018) B. Chérief-Abdellatif and P. Alquier Consistency of variational bayes inference for estimation and model selection in mixtures. Electronic Journal of Statistics 12 (2), pp. 2995 – 3035. Cited by: §A.1, Lemma 3.
- Chérief-Abdellatif (2020) B. Chérief-Abdellatif Convergence rates of variational inference in sparse deep learning. In ICML, Vol. 119, pp. 1831–1842. Cited by: §B.1, Appendix B, §3.2.
- Choudhary et al. (2020) T. Choudhary, V. Mishra, A. Goswami, and J. Sarangapani A comprehensive survey on model compression and acceleration. Artificial Intelligence Review 53, pp. 5113–5155. Cited by: §1.
- Courbariaux et al. (2015) M. Courbariaux, Y. Bengio, and J. David Binaryconnect: training deep neural networks with binary weights during propagations. In NeurIPS, Vol. 28, pp. 3123–3131. Cited by: §1, §2, §4.
- Courbariaux and Bengio (2016) M. Courbariaux and Y. Bengio BinaryNet: training deep neural networks with weights and activations constrained to +1 or -1. CoRR abs/1602.02830. Cited by: §4.
- Csiszar (1975) I. Csiszar -Divergence Geometry of Probability Distributions and Minimization Problems. The Annals of Probability 3 (1), pp. 146 – 158. Cited by: §2.
- De Sa et al. (2018) C. De Sa, M. Leszczynski, J. Zhang, A. Marzoev, C. R. Aberger, K. Olukotun, and C. Ré High-accuracy low-precision training. arXiv preprint arXiv:1803.03383. Cited by: §4.
- Dekking et al. (2006) F.M. Dekking, C. Kraaikamp, H.P. Lopuhaä, and L.E. Meester A modern introduction to probability and statistics: understanding why and how. Springer Texts in Statistics, Springer London. External Links: LCCN 2004057700 Cited by: §3.1.2.
- Deng et al. (2019) W. Deng, X. Zhang, F. Liang, and G. Lin An adaptive empirical bayesian method for sparse deep learning. In NeurIPS, Vol. 32, pp. 5564–5574. Cited by: §4.
- Dettmers et al. (2023) T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer QLoRA: efficient finetuning of quantized llms. In NeurIPS, Vol. 36, pp. 10088–10115. Cited by: §1, §5.1.
- Dettmers et al. (2024) T. Dettmers, R. Svirschevski, V. Egiazarian, D. Kuznedelev, E. Frantar, S. Ashkboos, A. Borzunov, T. Hoefler, and D. Alistarh SpQR: A sparse-quantized representation for near-lossless LLM weight compression. In ICLR, Cited by: §5.1.
- Devlin et al. (2019) J. Devlin, M. Chang, K. Lee, and K. Toutanova BERT: pre-training of deep bidirectional transformers for language understanding. In NAACL, pp. 4171–4186. Cited by: §D.1, §5.2.
- Dong et al. (2022) R. Dong, Z. Tan, M. Wu, L. Zhang, and K. Ma Finding the task-optimal low-bit sub-distribution in deep neural networks. In ICML, Vol. 162, pp. 5343–5359. Cited by: 2nd item, §1, §2, §2, §2, §3.1.2, §4, §5.1, §5.2, Table 1, Table 1, Table 1, Table 3, Table 3.
- Dong et al. (2017) X. Dong, S. Chen, and S. Pan Learning to prune deep neural networks via layer-wise optimal brain surgeon. In NeurIPS, Vol. 30, pp. 4857–4867. Cited by: §1, §4, §5.1, Table 2.
- Frankle and Carbin (2019) J. Frankle and M. Carbin The lottery ticket hypothesis: finding sparse, trainable neural networks. In ICLR, Cited by: §4.
- Frantar et al. (2023) E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh OPTQ: accurate quantization for generative pre-trained transformers. In ICLR, Cited by: 6th item, §1, §4, Table 2.
- Frantar et al. (2022) E. Frantar, S. P. Singh, and D. Alistarh Optimal brain compression: a framework for accurate post-training quantization and pruning. In NeurIPS, Vol. 35, pp. 4475–4488. Cited by: 5th item, §1, §3.1.3, §4, §4, §5.1, Table 2, Table 2, Table 2.
- Gholami et al. (2022) A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, and K. Keutzer A survey of quantization methods for efficient neural network inference. In Low-power computer vision, pp. 291–326. Cited by: §2.
- Gou et al. (2021) J. Gou, B. Yu, S. J. Maybank, and D. Tao Knowledge distillation: a survey. International Journal of Computer Vision 129 (6), pp. 1789–1819. Cited by: §1.
- Guo et al. (2016) Y. Guo, A. Yao, and Y. Chen Dynamic network surgery for efficient dnns. In NeurIPS, Vol. 29, pp. 1379–1387. Cited by: §1.
- Gupta et al. (2015) S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan Deep learning with limited numerical precision. In ICML, Vol. 37, pp. 1737–1746. Cited by: §4.
- Han et al. (2016) S. Han, H. Mao, and W. J. Dally Deep compression: compressing deep neural network with pruning, trained quantization and huffman coding. In ICLR, Cited by: §1, §5.1.
- Hassibi et al. (1993) B. Hassibi, D. G. Stork, and G. J. Wolff Optimal brain surgeon and general network pruning. In IEEE International Conference on Neural Networks, pp. 293–299. Cited by: 1st item, §4.
- Hershey and Olsen (2007) J. R. Hershey and P. A. Olsen Approximating the kullback leibler divergence between gaussian mixture models. In IEEE International Conference on Acoustics, Speech, and Signal Processing, pp. 317–320. Cited by: §A.1.
- Hubara et al. (2018) I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio Quantized neural networks: training neural networks with low precision weights and activations. Journal of Machine Learning Research 18 (187), pp. 1–30. Cited by: §4.
- Hubara et al. (2021) I. Hubara, Y. Nahshan, Y. Hanani, R. Banner, and D. Soudry Accurate post training quantization with small calibration sets. In ICML, Vol. 139, pp. 4466–4475. Cited by: 3rd item, §3.1.3, §4.
- Ishwaran and Rao (2005) H. Ishwaran and J. S. Rao Spike and slab variable selection: frequentist and bayesian strategies. The Annals of Statistics 33, pp. 730–773. Cited by: §3.1.
- Jacob et al. (2018) B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko Quantization and training of neural networks for efficient integer-arithmetic-only inference. In CVPR, pp. 2704–2713. Cited by: §2.
- Jordan et al. (1999) M. I. Jordan, Z. Ghahramani, T. S. Jaakkola, and L. K. Saul An introduction to variational methods for graphical models. Machine learning 37 (2), pp. 183–233. Cited by: §2.
- Kumar et al. (2025) T. Kumar, Z. Ankner, B. F. Spector, B. Bordelon, N. Muennighoff, M. Paul, C. Pehlevan, C. Re, and A. Raghunathan Scaling laws for precision. In ICLR, Cited by: §1.
- LeCun et al. (1989) Y. LeCun, J. Denker, and S. Solla Optimal brain damage. In NeurIPS, Vol. 2, pp. 598–605. Cited by: §1, §4.
- Li et al. (2017) H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf Pruning filters for efficient convnets. In ICLR, Cited by: Table 1.
- Li et al. (2021) Y. Li, R. Gong, X. Tan, Y. Yang, P. Hu, Q. Zhang, F. Yu, W. Wang, and S. Gu BRECQ: pushing the limit of post-training quantization by block reconstruction. In ICLR, Cited by: 4th item, §4.
- Lin et al. (2024) J. Lin, J. Tang, H. Tang, S. Yang, W. Chen, W. Wang, G. Xiao, X. Dang, C. Gan, and S. Han AWQ: activation-aware weight quantization for on-device LLM compression and acceleration. In Annual Conference on Machine Learning and Systems, Cited by: 1st item, §1, §4, §5.2, Table 3, Table 3.
- Lin et al. (2025) M. Lin, S. Guan, W. Jing, G. Botterweck, and A. Patane Stochastic weight sharing for bayesian neural networks. In AISTATS, Proceedings of Machine Learning Research, Vol. 258, pp. 4519–4527. Cited by: §4.
- Liu et al. (2018) C. Liu, B. Zoph, M. Neumann, J. Shlens, W. Hua, L. Li, L. Fei-Fei, A. Yuille, J. Huang, and K. Murphy Progressive neural architecture search. In ECCV, pp. 19–34. Cited by: §1.
- Liu et al. (2025a) K. Liu, Q. Zheng, K. Tao, Z. Li, H. Qin, W. Li, Y. Guo, X. Liu, L. Kong, G. Chen, Y. Zhang, and X. Yang Low-bit model quantization for deep neural networks: a survey. External Links: 2505.05530 Cited by: §1.
- Liu et al. (2025b) Z. Liu, C. Zhao, I. Fedorov, B. Soran, D. Choudhary, R. Krishnamoorthi, V. Chandra, Y. Tian, and T. Blankevoort SpinQuant: LLM quantization with learned rotations. In ICLR, Cited by: §4.
- Louizos et al. (2019) C. Louizos, M. Reisser, T. Blankevoort, E. Gavves, and M. Welling Relaxed quantization for discretized neural networks. In ICLR, Cited by: §4.
- Louizos et al. (2017) C. Louizos, K. Ullrich, and M. Welling Bayesian compression for deep learning. In NeurIPS, Vol. 30, pp. 3288–3298. Cited by: §4, §4.
- Louizos et al. (2018) C. Louizos, M. Welling, and D. P. Kingma Learning sparse neural networks through regularization. In ICLR, Cited by: §4.
- Marchesi et al. (1993) M. Marchesi, G. Orlandi, F. Piazza, and A. Uncini Fast neural networks without multipliers. IEEE transactions on Neural Networks 4 (1), pp. 53–62. Cited by: §4.
- Nagel et al. (2020) M. Nagel, R. A. Amjad, M. Van Baalen, C. Louizos, and T. Blankevoort Up or down? adaptive rounding for post-training quantization. In ICML, Vol. 119, pp. 7197–7206. Cited by: §3.1.3.
- Nielsen et al. (2025) J. Nielsen, P. Schneider-Kamp, and L. Galke Continual quantization-aware pre-training: when to transition from 16-bit to 1.58-bit pre-training for bitnet language models?. In ACL (Findings), pp. 13483–13493. Cited by: §4.
- Park et al. (2019) W. Park, D. Kim, Y. Lu, and M. Cho Relational knowledge distillation. In CVPR, pp. 3967–3976. Cited by: §1.
- Radford et al. (2018) A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever Improving language understanding by generative pre-training. Technical Report OpenAI. Cited by: §1.
- Rajpurkar et al. (2016) P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang SQuAD: 100,000+ questions for machine comprehension of text. In EMNLP, pp. 2383–2392. Cited by: §D.1, §5.2.
- Rastegari et al. (2016) M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi Xnor-net: imagenet classification using binary convolutional neural networks. In ECCV, pp. 525–542. Cited by: §1, §4.
- Roth and Pernkopf (2018) W. Roth and F. Pernkopf Bayesian neural networks with weight sharing using dirichlet processes. IEEE Transactions on Pattern Analysis and Machine Intelligence 42 (1), pp. 246–252. Cited by: §4.
- Shayer et al. (2018) O. Shayer, D. Levi, and E. Fetaya Learning discrete weights using the local reparameterization trick. In ICLR, Cited by: §4.
- Singh and Alistarh (2020) S. P. Singh and D. Alistarh Woodfisher: efficient second-order approximation for neural network compression. In NeurIPS, Vol. 33, pp. 18098–18109. Cited by: §4.
- Sun et al. (2019) X. Sun, J. Choi, C. Chen, N. Wang, S. Venkataramani, V. V. Srinivasan, X. Cui, W. Zhang, and K. Gopalakrishnan Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks. In NeurIPS, Vol. 32, pp. 4901–4910. Cited by: §1.
- Sun et al. (2022) Y. Sun, Q. Song, and F. Liang Consistent sparse deep learning: theory and computation. Journal of the American Statistical Association 117 (540), pp. 1981–1995. Cited by: §3.1.2.
- Sze et al. (2017) V. Sze, Y. Chen, T. Yang, and J. S. Emer Efficient processing of deep neural networks: a tutorial and survey. Proceedings of the IEEE 105 (12), pp. 2295–2329. Cited by: §4.
- Touvron et al. (2023) H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. Llama: open and efficient foundation language models. arXiv preprint arXiv:2302.13971. Cited by: §1.
- Ullrich et al. (2017) K. Ullrich, E. Meeds, and M. Welling Soft weight-sharing for neural network compression. In ICLR, Cited by: §4, §4.
- Van Baalen et al. (2020) M. Van Baalen, C. Louizos, M. Nagel, R. A. Amjad, Y. Wang, T. Blankevoort, and M. Welling Bayesian bits: unifying quantization and pruning. In NeurIPS, Vol. 33, pp. 5741–5752. Cited by: §4.
- van den Oord et al. (2017) A. van den Oord, O. Vinyals, and K. Kavukcuoglu Neural discrete representation learning. In NeurIPS, Vol. 30. Cited by: §5.1.
- Wang et al. (2019) C. Wang, R. Grosse, S. Fidler, and G. Zhang Eigendamage: structured pruning in the kronecker-factored eigenbasis. In ICML, Vol. 97, pp. 6566–6575. Cited by: §4.
- Wang et al. (2020a) P. Wang, Q. Chen, X. He, and J. Cheng Towards accurate post-training network quantization via bit-split and stitching. In ICML, Vol. 119, pp. 9847–9856. Cited by: 2nd item, §4.
- Wang et al. (2025) R. Wang, Y. Gong, X. Liu, G. Zhao, Z. Yang, B. Guo, Z. Zha, and P. Cheng Optimizing large language model training using FP4 quantization. In ICML, Proceedings of Machine Learning Research, Vol. 267, pp. 62937–62957. Cited by: §1.
- Wang et al. (2020b) T. Wang, K. Wang, H. Cai, J. Lin, Z. Liu, H. Wang, Y. Lin, and S. Han Apq: joint search for network architecture, pruning and quantization policy. In CVPR, pp. 2078–2087. Cited by: §1, §1, §4.
- Wang et al. (2020c) Y. Wang, Y. Lu, and T. Blankevoort Differentiable joint pruning and quantization for hardware efficiency. In ECCV, pp. 259–277. Cited by: §1, §4.
- Wei et al. (2022) X. Wei, Y. Zhang, X. Zhang, R. Gong, S. Zhang, Q. Zhang, F. Yu, and X. Liu Outlier suppression: pushing the limit of low-bit transformer language models. In NeurIPS, Vol. 35, pp. 17402–17414. Cited by: §3.1.2.
- Wortsman et al. (2023) M. Wortsman, T. Dettmers, L. Zettlemoyer, A. Morcos, A. Farhadi, and L. Schmidt Stable and low-precision training for large-scale vision-language models. In NeurIPS, Vol. 36, pp. 10271–10298. Cited by: §1.
- Wu et al. (2018) S. Wu, G. Li, F. Chen, and L. Shi Training and inference with integers in deep neural networks. In ICLR, Cited by: §4.
- Xia et al. (2024) M. Xia, T. Gao, Z. Zeng, and D. Chen Sheared LLaMA: accelerating language model pre-training via structured pruning. In ICLR, Cited by: §4.
- Xu et al. (2020) C. Xu, W. Zhou, T. Ge, F. Wei, and M. Zhou BERT-of-theseus: compressing BERT by progressive module replacing. In EMNLP, pp. 7859–7869. Cited by: §1.
- Xu et al. (2025) Z. Xu, Y. Zhang, J. Li, J. Guo, Q. Zhu, and H. Huang Towards efficient vlms: information-theoretic driven compression via adaptive structural pruning. arXiv preprint arXiv:2511.19518. Cited by: §4.
- Yin et al. (2019) P. Yin, J. Lyu, S. Zhang, S. J. Osher, Y. Qi, and J. Xin Understanding straight-through estimator in training activation quantized neural nets. In ICLR, Cited by: §4, §4.
- You et al. (2019) Z. You, K. Yan, J. Ye, M. Ma, and P. Wang Gate decorator: global filter pruning method for accelerating deep convolutional neural networks. In NeurIPS, Vol. 32, pp. 2130–2141. Cited by: §1.
- Zhang et al. (2018) D. Zhang, J. Yang, D. Ye, and G. Hua Lq-nets: learned quantization for highly accurate and compact deep neural networks. In ECCV, pp. 365–382. Cited by: Table 1.
- Zhang et al. (2022) Q. Zhang, S. Zuo, C. Liang, A. Bukharin, P. He, W. Chen, and T. Zhao Platon: pruning large transformer models with upper confidence bound of weight importance. In ICML, Vol. 162, pp. 26809–26823. Cited by: §5.1, Table 2.
- Zhang and Shrivastava (2025) T. Zhang and A. Shrivastava LeanQuant: accurate and scalable large language model quantization with loss-error-aware grid. In ICLR, Cited by: §4.
- Zhao et al. (2025) S. Zhao, M. Pitts, and Z. Qin EfficientXpert: efficient domain adaptation for large language models via propagation-aware pruning. arXiv preprint arXiv:2511.19935. Cited by: §4.
- Zhu et al. (2017) C. Zhu, S. Han, H. Mao, and W. J. Dally Trained ternary quantization. In ICLR, Cited by: Table 1, Table 1.
- Zhu and Gupta (2017) M. Zhu and S. Gupta To prune, or not to prune: exploring the efficacy of pruning for model compression. arXiv preprint arXiv:1710.01878. Cited by: Table 2.
Appendix A Derivation of Approximate Objective
| Symbol | Meaning |
|---|---|
| Pre-trained full-precision weight that is fixed and provided as input to the compression procedure. | |
| Random sparse and quantized weight learned during compression. | |
| Marginal variational distribution of , defined by the spike-and-GMM model in Equation (6). | |
| Spike-and-slab prior defined in Equation (5). | |
| ; ; | Keep/prune indicator, prior retention probability, and variational retention probability, respectively. |
| Responsibility of GMM component for weight , evaluated as a function of the fixed weight . | |
| Learnable component mean and variance; the means define the quantization levels. | |
| Variance of the Gaussian slab in the spike-and-slab prior. | |
| Quantity estimated after optimization, such as , , or . | |
| , , | Number of GMM components, total number of weights, and number of posterior samples used for Bayesian averaging, respectively. |
A.1 An upper bound on the KL divergence between two mixtures
To simplify the ELBO and validate our approach, we reformulate a key lemma from previous work (Chérief-Abdellatif and Alquier, 2018, Lemma 6.1). This Lemma is a tool widely used in signal processing (Hershey and Olsen, 2007). We provide the proof for the sake of completeness.
Lemma 3(From Lemma 6.1 in (Chérief-Abdellatif and Alquier, 2018) ).
For any , the KL divergence between any two mixture densities and is upper bounded by
Proof.
We expand the KL divergence term by its definition and obtain:
where the first inequality is due to Jensen’s inequality and the convexity of the function . This completes the proof. ∎
A.2 Derivation of Approximate Objective
We aim to approximate the ELBO objective:
| (18) |
where is defined in Equation (5) and is defined in Equation (6):
It is important to note that the KL divergence between the variational distribution and the spike-and-slab prior distribution does not have a closed-form solution.
Step 1: Approximate the expected log-likelihood. The first term can be expensive to compute because it requires sampling from the spike-and-GMM distribution. We therefore use a plug-in approximation based on the posterior mean. For each coordinate ,
Collecting these coordinate-wise means gives the complete mean parameter vector
The plug-in approximation is therefore
Step 2: Upper Bound KL between spike-and-slab distributions. The KL divergence between the marginal variational posterior and the prior is intractable due to the presence of both the Dirac delta and the mixture components. To upper-bound the KL divergence between them, we apply Lemma 3 by matching component structure:
| with | |||||||||
Substituting into the bound, we obtain:
Combining the terms, we have:
Note that the first term on the right-hand side, which is the KL divergence between the GMM and the Gaussian distribution, does not have a closed form. But it can be further upper-bounded as:
where the inequality is obtained by Lemma 3. Empirically, we approximate the mixture KL by evaluating only the dominant component:
We approximate the inner sum over using the maximum-weight component, which is the -th component.
A small temperature is needed to avoid a flat posterior distribution, which could introduce large differences between the training phase and inference phase.
Finally, putting all approximations together, we obtain:
where is the mean parameter vector defined in Step 1, and . We thus obtain the result shown in Equation (8).
Appendix B Proof of Theorem 1
Consider an - hidden-layer fully connected neural network with the ReLU activation function defined as on some dimension and parameter . The number of neurons in each layer is defined as for . The weights and biases are denoted by and . Thus, given the parameters , let denote the vector obtained by stacking all entries of the weight matrices and bias vectors , then the fully connected network can be presented as:
The DNN also introduces a probability measure of the data, which we denote as , and is the corresponding density function; would be the likelihood of the data .
One can define the sparse parameter space with sparsity parameter as , where has only many non-zero entries. Then we can further introduce the sparse and quantized weights space as follows:
where is some constant that satisfies and is the indexing space; each element consists of many -dimensional one-hot rows, and the remaining rows are zero vectors indicating the corresponding weight is pruned. In such a way, any satisfies that and only have many distinct entry values then the DNN is sparse and quantized. The following conditions are assumed, similarly to (Bai et al., 2020):
Condition 1.
that can depend on n, and .
Condition 2.
is 1-Lipschitz continuous.
Condition 3.
The hyperparameter is set to be some constant, and satisfies
Condition 4.
and .
The “oracle” sparsity is defined in Equation (19).
Definition 1.
The true function defined in Equation (13) is -Hölder smoothness if
for some constant .
Following previous paper (Bai et al., 2020), we define:
| (19) |
where
Correspondingly, we define .
In this section, we reformulate the variational distribution by introducing a latent index variable. For any , it has the following equivalent form:
| (20) | ||||
In addition, for theoretical convenience, we further restrict the variational family to satisfy
Condition 5.
and .
Note that the requirement of is fairly reasonable, as most of the existing approximation results (Chérief-Abdellatif, 2020) only need bounded DNN weights.
We restate a formal version of our Theorem 1 as follows:
Theorem 2.
Under Conditions 1-2 and 4-5, Let be a constant and for any constant , Then with high probability:
| (21) |
where denotes the Hellinger distance, and and are some constants.
Proof.
The convergence in squared Hellinger distance follows directly from Lemmas 1 and Lemma 2, as the chosen value of meets the necessary assumptions. ∎
Remark. Compared to prior results, Lemma 1 demonstrates that a spike-and-slab prior combined with a Gaussian mixture model (GMM) with finitely many components can effectively approximate the true underlying function. In contrast, Lemma 2 establishes that the statistical estimation error of the spike-and-GMM variational distribution vanishes as the sample size .
B.1 Proof of Lemma 1
Proof.
Let . By definition, they are only many unique non-zero number in , denoted as , for . In other words, for any , must choose from the quantization set , and we denote the choice index of from as , (i.e., ). Now, given , we construct as follow:
where . 22 2 Notice that satisfies Condition 5 as with sufficiently large and , and . Thus, we can have the following marginal distribution:
Next we first need to bound . We can also write , then we define the following terms as:
Then, following the proof in (Chérief-Abdellatif, 2020, Proof of Theorem 7), we can have the following:
| (22) |
Then next we need to upper bound the term:
We first bound the following, for some ,
| (23) |
Notice that by definition , thus we can bound the first term as:
Next, we bound the second term of the Equation (23):
Notice that the last inequality is because of . And by choosing , the Equation (23), can be bounded by:
The second inequality is obtained by setting such that . Note that given a fixed DNN structure, such a always exists as the decrease monotonically to zero as decrease. Thus, we can have
Next we bound , following similar procedure, for some , we can have:
| (24) |
Note that with a slight abuse of notation, the in the above equation means the latent variable (introduced in (20)) corresponding to the weight . The first term of Equation (24) can be bounded for as follow:
And the second term of Equation (24) can be bounded as:
The last inequality is again because of the property that . Thus Equation (24) can be bounded by:
where the last inequality is because of choosing a small such , such always exist since decrease monotonically to as .
Thus we can bound by the following:
By choosing , we can have:
Similarly, we can have the following:
Combined with Equation (22), we can have:
With some algebra, we can have:
as ,
The last equality is due to the definition of , and the last inequality is due to the definition of .
In the next step, we aim to bound the integral . Note that by definition:
We can define the following:
Since , it follows that
Given , we have
whereby the Cauchy-Schwarz inequality,
Thus, , and with high probability, for some positive constant if , or for any diverging sequence if . Therefore,
| (25) |
In the next step, we try to bound the divergence between and ,
| (26) | ||||
| (27) | ||||
| (28) | ||||
where the inequality (26) and (27) are follows from Lemma 3 and the inequality (28) is because of the fact that and , thus combined with the result in equation (25), we finish the proof. ∎
B.2 Proof of Lemma 2
Proof.
Following previous work (Bai et al., 2020, proof of Lemma 4.2), we first define the space
By the above definitions, we now have:
| (29) |
Lemma 4 presents a variational characterization of the divergence, originally due to Donsker and Varadhan. The proof is available in (Boucheron et al., 2013).
Lemma 4.
Let be any probability measure and a measurable function with , then
We can define the truncation of distribution on the set denoted as , (i.e. ), similarly we can also define the . By adopting the arguments from (Bai et al., 2020), and following steps analogous to those leading to Equation (17) therein, we obtain:
| (30) |
for some constant , where .
Then, given the Lemma 1 and equation (30), we can show that the first term can be bounded w.h.p. as:
| (31) |
Additionally, since , the second and the third terms of Equation (29) are bounded by:
Substituting the bound on the first term from Equation (31), together with the two bounds above, into Equation (29), we obtain with high probability:
| (32) |
Following the procedure in (Bai et al., 2020, Lemma 4.2, equation 20), we can show with high probability that:
where is some constant. Next, we show that .
Lemma 5.
Given the Condition 5, holds.
Proof.
Let , then by definition of the variational distribution ,we can know that:
By the definition of variational distribution , we know:
And by Chernoff bound and the fact that , we can have:
Thus we show that . ∎
Then by Lemma 5, we complete the proof. ∎
Appendix C Implementation of SQS
Pretrained model setting. Taking the compression of the Llama3.2-1B model as an example, we first download the pre-trained model from Hugging Face33 3 https://huggingface.co/meta-llama/Llama-3.2-1B using the Python package “transformers”. We then fine-tune the model on the considered SST-2 task before applying our SQS for compression. We find that omitting the fine-tuning step significantly degrades the performance of SQS.
For ResNet models, we use publicly available pre-trained models obtained by the Python package “timm’’ on the CIFAR-10 and CIFAR-100 datasets and directly apply our SQS for compression. For the BERT-base model, we download the pre-trained model from Hugging Face44 4 https://huggingface.co/huggingface-course/bert-finetuned-squad using the Python package “transformers”.
Initialization. To initialize the learnable parameters of our SQS method, denoted as , we employ the K-means algorithm. Specifically, the DNN weights of a given layer are first clustered into groups. For each group , the mean and standard deviation are computed as the empirical statistics of the weights in that group, while the mixture coefficient is set to the proportion of weights in group relative to the total number of weights in the layer.
We assume that K-means yields disjoint groups of weights, denoted as , such that covers all weights in the selected layer. The initial parameters are then defined as:
Implementing marginal . To make the distribution differentiable, we reparameterize (as defined in Equation 6) using the following equation:
| (33) |
where the temperature is set to a fixed constant to stabilize the training process. To better exploit the learned pruning parameters in later stages of training, we halve after completing half of the total training steps to stabilize the training and sharpen the retention probabilities around their learned optima.
Hyperparameter Configuration. The number of Gaussian components is not fixed across experiments; it is chosen per model and can be read off the Bits column of each experiment table, since . Specifically, we use ( bits) for the ResNet models on CIFAR-10 in Table 1, ( bits) for BERT-base on SQuAD v1.1 in Table 2, and ( bits) for Llama3.2-1B and Qwen2.5-0.5B on SST-2 in Table 3. The ablation studies state their own in the corresponding table or figure caption. A larger reduces the performance drop at the cost of a lower compression rate; this trade-off is analyzed in the case study shown in Figure 3. For the experiments not reporting the Bayesian Averaging number , the default is set to 4. All models are trained on an NVIDIA H100 GPU with 80 GB of memory.
During training and testing, for ResNet-18, ResNet-20, ResNet-32, ResNet-50, ResNet-56, BERT-base, Llama3.2-1B, and Qwen2.5-0.5B models, the settings are:
Training Time: Approximately 30 minutes for ResNet models (ResNet-18 through ResNet-56); 4 hours for BERT-base; 24 hours for both Llama3.2-1B and Qwen2.5-0.5B.
Optimizer: AdamW is used consistently across all models.
Quantization Learning Rate: for ResNet-18; for all other models.
Pruning Learning Rate: Fixed at for all models.
Pruning Schedule: A polynomial schedule is used for all models.
Appendix D Experiment Settings
D.1 Experiment settings for benchmark with all baselines
Benchmark compression on ResNet models. We present experiments using ResNet architectures on the CIFAR-10 and CIFAR-100 datasets. When compressing ResNet models, our method requires fine-tuning over the training dataset, completing the compression process within epochs. To achieve high compression rates, we represent each layer’s weights with components (i.e., for each layer) and apply a sparsity level of . As shown in Table 1, our method achieves 32× compression on all three ResNet models, with reported accuracy drops below 1.5 percentage points. For example, compressing ResNet-32 by a factor of yields a minimal accuracy reduction of . Additionally, we compress ResNet-56 by a factor of , observing an accuracy drop of only . Compared to other methods, our approach achieves much higher compression rates with smaller decreases in accuracy.
Benchmark compression on BERT-base model. We further investigate our compression method on attention-based models. We apply our compression model on the BERT-base (Devlin et al., 2019) model and test it on the SQuAD V1.1 dataset (Rajpurkar et al., 2016). Similarly, we consider the F1 score drop and compression rate as the evaluation metrics. During the compression process, the BERT model is fine-tuned on the training dataset, with the entire procedure completed within epochs.
We compressed the BERT model using Gaussian components and pruned of its parameters, leading to a compression rate. We employed layer-wise quantization combined with unstructured pruning to attain these results.
Benchmark compression on Llama and Qwen models. Due to hardware limitations, we cannot run very large-scale LLMs, which are Llama3.1-8B and Qwen2.5-7B.
D.2 Experiment settings for ablation studies for SQS method
Impact of different priors. For comparison, we consider a zero-mean Gaussian distribution as the prior and replace the delta distribution with a Gaussian distribution in the variational family. That is, any has the form:
Based on this, we can get the modified marginal variational distribution as:
| (34) |
Thus, following the same reasoning and derivation, we get the equation (8), The Gaussian prior can be obtained by:
| (Gaussian prior) |
We compare the impact of the above Gaussian prior with the proposed Spike-and-GMM priors and summarize the result in Table 4.
Description of Baselines. For the following lines of baselines, we use the reported results in their papers:
Optimal Brain Surgeon (OBS) (Hassibi et al., 1993) selects weights for removal from a trained neural network using second-order information.
BitSplit (Wang et al., 2020a) incrementally constructs quantized values using a squared error metric based on residual errors.
AdaQuant (Hubara et al., 2021) utilizes STE for direct optimization.
BRECQ (Li et al., 2021) integrates Fisher information into the optimization process and focuses on the joint optimization of layers within individual residual blocks.
Exact Optimal Brain Quantization (OBQ) (Frantar et al., 2022) adapts second-order weight pruning methods to quantization tasks.
GPTQ (Frantar et al., 2023) employs second-order information for error compensation on calibration sets to speed up generative models.
We adopt their implemented code and use the same setting for training and testing:
AWQ (Lin et al., 2024) implements activation-aware quantization, selectively bypassing the quantization of key weights55 5 https://github.com/mit-han-lab/llm-awq. This method is training-free and does not need extra training on the selected dataset.
DGMS (Dong et al., 2022) is an automated quantization method that utilizes Mixtures of Gaussians to avoid the aforementioned problem66 6 https://github.com/RunpeiDong/DGMS. We use their codebase and configure it with the same hyperparameters. Their algorithm is trained on the same dataset for fairness of comparison.
Definition of Evaluation Metrics. Let denote the number of shared weight vectors. Then, the metric Bits is defined as .
D.3 Long-tailed Full-precision Weight Distribution of Llama3.2 and Qwen2.5 Models
We present the visualization of the long-tailed weight distributions for the Llama3.2 model in Figures 6 and 7, and for the Qwen2.5 model in Figures 4 and 5.