arXiv · 2609.38106
Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs
Abstract
Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we systematically study the effect of audio encoder pruning on SLAM-ASR for different demographic groups. Using the Fair-Speech and Common Voice datasets, we found that the pruning does not affect all demographic groups equally; the gap between best- and worst-performing groups increases in fold. These disparities appear across all three encoder scales, but only the largest model initially hides them behind aggregate WER. LoRA adaptation improves WER for every group, but benefits groups already performing well more strongly and widens for certain groups. On Common Voice English, Danish, and Dutch, accent gaps persist but do not clearly widen, showing that the fairness effects of pruning vary across datasets and must be measured directly. Our findings suggest that for pruned models, deployment decisions should include per-group WER, with the worst-performing group's error rate as an explicit criterion.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi Shekhar. 2026-09-29. Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs. https://arxiv.org/abs/2609.38106
Cite the original work for its findings. Save a collection to share your selection of sources.