arXiv · 2610.11131
SEER: Source-Conditioned Emotion Enhancement via Retrieval for Cochlear-Implant Speech
Abstract
Cochlear implants (CIs) restore speech access but weaken cues needed for vocal emotion recognition. Prior CI-oriented enhancement requires parallel normal/strong recordings and intensity labels. We propose SEER, a retrieval-based framework that learns which same-emotion reference helps each source remain recognizable after CI processing. A source-conditioned retriever learns CI-aware utility from sampled emotional voice conversion outcomes, while uncertainty-guided exploration avoids exhaustive pair evaluation; neither parallel recordings nor intensity labels are required. SEER improves Source macro-F1 at N8 by 7.30 points on RAVDESS and 11.66 points on ESD, with significant ESD gains across N4/N8/N16. Sixteen-listener RAVDESS gains are significant across all conditions. Exhaustive analysis finds an aggregate benefit from stronger references but little effect from matching gender or content.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hsing-Hang Chou, Yun-Shao Lin, Ching-Chin Sung, Chi-Chun Lee. 2026-10-08. SEER: Source-Conditioned Emotion Enhancement via Retrieval for Cochlear-Implant Speech. https://arxiv.org/abs/2610.11131
Cite the original work for its findings. Save a collection to share your selection of sources.