Search arXiv⌕ Search

arXiv subjects

Yinghong Lan

Publications and source records attributed to Yinghong Lan.

2 recordsLinked to original sources

Louder, Longer, Livelier: Acoustic Shortcuts and Underspecified Rationales in Speech LLM Judges

LLM-as-a-judge is widely used for evaluating text, but extending this paradigm to speech requires models to interpret acoustic as well as linguistic evidence. This introduces a modality-specific risk: a speech judge may treat a perceptually salient cue as evidence of quality even when that cue is irrelevant to the target criterion or receives more weight than human listeners give it. We call this behavior an acoustic shortcut. To study it, we audit six speech LLM judges using controlled manipulations of intensity, content richness, and emotional delivery. We evaluate both pointwise scoring and pairwise comparison, using human preference calibration to interpret the results. The judges consistently reward louder audio, prefer content-rich speech more strongly than human listeners do, and map emotional delivery into quality preferences. These effects are most visible in pairwise comparison, while pointwise scores often obscure them. More concerningly, the accompanying rationales rarely identify the acoustic cue that changes a judgment and instead repeatedly rely on a limited vocabulary, leaving them acoustically underspecified. Together, these findings show that reliable speech judges must both resist acoustic shortcuts and ground their rationales in the acoustic evidence behind their decisions. To support reproducibility and future audits, we also release SpeechJudgeAudit, the controlled stimuli and evaluation tools used in this study.

eess.AS↗

SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences

Human speech conveys rich perceptual information, such as emotion and speaker identity, yet most automatic speech quality judges reduce it to a single naturalness score. We study diagnostic speech judges: given two candidates, a diagnostic judge decides which is better, along which perceptual dimensions (e.g., timbre, emotion, timing) they differ, and which audible cues support its decision. Learning such judges is challenging: expert annotation is costly, and simply prompting a frontier audio-language model to produce labels is unreliable: our probing reveals substantial errors and unstable instruction following. We introduce SpeechCritic, which learns a diagnostic judge in a reference-conditioned cross-lingual setting from only about 300 human-labeled comparisons. Rather than replacing the frontier model, SpeechCritic calibrates it with these labels: for each dimension, it selects the acoustic measurements that agree with human judgments, maps them to A/Tie/B probabilities, and passes these to the model as non-binding hints alongside the audio. Compared with the same model labeling without hints, this raises dimension-level agreement with humans by 6.3 points and cuts the mismatch with human Tie rates by 10.4 points. We then train a 7B judge on this supervision and find that different training signals shape different judge behaviors: SFT establishes the task, OPD transfers the teacher's dimension-level strengths and weaknesses, and RL helps most on clear-cut comparisons where human raters agree. Notably, human listeners also find that RL makes rationales cite more specific, localized acoustic cues, although it never directly rewards rationale text. Finally, we show that the pipeline is language-pair agnostic by instantiating it on both English-Japanese and English-Spanish. Together, these results demonstrate a path from limited human preferences to a diagnostic speech judge.

cs.AI↗