Search arXiv⌕ Search

arXiv subjects

Yunda Chen

Publications and source records attributed to Yunda Chen.

3 recordsLinked to original sources

Adaptive Depth and Expert Refinement for Efficient Speech Enhancement

Most neural speech enhancement systems use a fixed processing depth for all inputs, which can introduce unnecessary computation when fewer refinement steps are sufficient. We propose Adaptive Depth and Expert Refinement (ADER), a parameter-shared progressive enhancement framework with input-dependent computation. ADER combines an Adaptive Depth Controller (ADC) for hard early termination with a Conditional Expert Router (CER) that selects one lightweight residual adapter at each executed refinement iteration. We further introduce Exit-aware Intermediate Supervision (EIS) to directly optimize candidate intermediate outputs for early exit. On VCTK-DEMAND, ADER reduces the parameter count and average computation of MP-SENet by 70.4% and 51.3%, respectively, while achieving a WB-PESQ of 3.37. Overall, ADER enables input-dependent refinement and reduces redundant computation during inference.

eess.AS↗

Local chord corruption is not recognizer replay: structure-matched calibration

Synthetic chord substitutions offer controlled tests of music generation, but their effects can differ from those of a complete recognized chord sequence. We propose structure-matched calibration, which constructs synthetic chord sequences that match the changed positions and harmonic-relation composition of recognizer replay. Paired generation compares both target-response magnitude and output-chord agreement with replay. On 29 of 30 MUSDB18-HQ songs, central four-second tritone corruption produces a larger target response than complete replay in MIDI-SAG. On 24 held-out MoisesDB songs, structure matching reduces target-response distance to replay by 81% for MIDI-SAG and 77% for MusicGen-Chord. Distance decreases on every song in both models with CNN--CRF. Joint matching also reduces output-chord mismatch with replay by 8--17 percentage points relative to temporal or relational matching alone.

cs.SD↗

Deep Learning for Personalized Binaural Audio Reproduction

Personalized binaural audio reproduction is the basis of realistic spatial localization, sound externalization, and immersive listening, directly shaping user experience and listening effort. This survey reviews recent advances in deep learning for this task and organizes them by generation mechanism into two paradigms: explicit personalized filtering and end-to-end rendering. Explicit methods predict personalized head-related transfer functions (HRTFs) from sparse measurements, morphological features, or environmental cues, and then use them in the conventional rendering pipeline. End-to-end methods map source signals directly to binaural signals, aided by other inputs such as visual, textual, or parametric guidance, and they learn personalization within the model. We also summarize the field's main datasets and evaluation metrics to support fair and repeatable comparison. Finally, we conclude with a discussion of key applications enabled by these technologies, current technical limitations, and potential research directions for deep learning-based spatial audio systems.

eess.AS↗