Search arXiv⌕ Search

arXiv · 2609.26823

Text Scores Can Miss Waveform Use: A Qwen2-Audio Quantization Case Study

Abstract

Post-training quantization of speech language models is often summarized with text-output scores and nominal bit widths. Those numbers alone do not establish behavior that depends on information missing from a transcript, or efficiency for a particular runtime. We introduce an evaluation protocol that separately tests lexical output, a transcript-insufficient endpoint, and a measured packed implementation. In a Qwen2-Audio case study, a translation-selected 6-bit allocation improves chrF by 2.36 on a frozen English-to-German replay, with paired 95% bootstrap interval [1.04, 3.62], but loses 3.91 percentage points on speaker-disjoint emotion recognition. At the same 6-bit budget, the uniform structural control reaches higher emotion accuracy than the selected allocation, and the front-layer control is also higher by point estimate on the same frozen set. At 7 bits, chrF improves by 3.28 with interval [2.08, 4.59], the emotion interval against FP16 includes zero, and a same-budget front-layer control still exceeds the selected allocation. A separate matched-budget 4.08-bit study finds roughly 10-point emotion deficits for every tested low-bit allocation and no selected-allocation advantage over frozen controls. Finally, a dequantized average-6-bit simulation retains the FP16 peak memory. This case study identifies a precision-dependent mismatch between lexical output, waveform-dependent behavior, and nominal precision. It does not establish a general failure of low-bit speech models or a deployment benefit for the selected allocation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mengzhe Geng, Jinxi Jin, Junhao Xu. 2026-09-20. Text Scores Can Miss Waveform Use: A Qwen2-Audio Quantization Case Study. https://arxiv.org/abs/2609.26823

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Building on the Joint-Embedding Predictive Architecture (JEPA) paradigm, a recent self-supervised learning framework that predicts latent representations of masked regions in high-level feature spaces, we propose Audio-JEPA (Audio Joint-Embedding Predictive Architecture), tailored specifically for audio data. Audio-JEPA uses a simple Vision Transformer backbone to predict latent representations of masked spectrogram patches rather than reconstructing raw audio. We pre-train on unlabeled AudioSet clips (10s, 32kHz) with random patch masking on mel-spectrograms. We evaluate on the X-ARES suite covering speech, music, and environmental sound tasks. Although our implementation is a straightforward translation of the original model to audio, the results still show comparable performance to wav2vec 2.0 and data2vec while using less than one-fifth of their training data and with no hyper-parameter tuning. All code and pretrained checkpoints are available on GitHub.

cs.SD↗

Channel-Preserving Representation Alignment for EEG-to-Music Reconstruction

Reconstructing naturalistic music from EEG requires extracting music-specific information from weak signals distributed across electrodes. Existing spectrogram regression and diffusion approaches provide audio synthesis mechanisms, but learning the EEG-to-music mapping faces an information--estimation tradeoff. Combining electrode measurements into fewer features can ease estimation from limited recordings while obscuring distinctions between music segments. We propose channel-preserving representation alignment, separating electrode information retention from prediction regularization. Per-electrode tokenization preserves spatial identity for learned aggregation, while multi-view self-distillation and channel dropout encourage consistency across electrode subsets during pretraining and music alignment. The aligned representation conditions a pretrained audio generator. Our theory explains these complementary choices. A finite-sample regression analysis identifies when retained task information outweighs estimation cost. For nonlinear alignment, an exact decomposition reveals a prediction-consistency penalty in the masked contrastive objective, with margin-dependent bounds connecting prediction variability to ranking errors across electrode subsets. On NMED-T/H, our method improves within-subject 50-way identification from 0.388 to 0.485 and generated-audio CLAP similarity from 0.635 to 0.683 compared with previous work.

cs.SD↗

Patient-level validation of a foundation-model pipeline for pediatric lung sounds: physician-annotated adventitious events are recognized, disease groups are not reliably predicted

Background: Lung-sound classifiers are usually evaluated on recording- or event-level splits, although each child contributes many recordings. We evaluated a foundation-model pipeline (PulmoVec) with the patient as the unit of partitioning and asked whether disease group can be predicted beyond age and sex. Methods: We analyzed 19693 physician-annotated respiratory events from 736 children in the public SPRSound database. A frozen Health Acoustic Representations (HeAR) encoder with low-rank adapters was trained for screening (normal versus adventitious), sound pattern (normal, crackles, wheeze/rhonchi) and disease group (pneumonia, bronchial disease, normal/other); event probabilities were stacked with age, sex and auscultation site. The primary analysis was nested patient-grouped cross-validation; comparators were the majority class, annotated event duration and demographics. Results: The area under the receiver operating characteristic curve (AUC) was 0.95 (95% CI 0.94 to 0.96) for both screening and sound pattern, against 0.78 and 0.75 for event duration alone; the positive predictive value for adventitious events was 0.72. Disease-group prediction reached an AUC of 0.58 (95% CI 0.54 to 0.62), and its accuracy was below the majority-class rate at event level (0.44 versus 0.62) and at patient level (0.46 versus 0.60). Conclusions: Recognition of physician-annotated adventitious events holds under patient-level validation and is not explained by event duration or demographics, although its positive predictive value would limit unaided use. Prediction of disease group, a clinical diagnosis, from lung-sound events alone was not demonstrated. Lung-sound studies should partition data by patient and report non-acoustic comparators.

cs.SD↗