Search arXiv⌕ Search

arXiv · 2610.02752

GAANet: Global-guided Asymmetric Attention Network for Audio-Visual Speech Separation

Abstract

Multi-scale design is crucial for efficient audio-visual speech separation, yet effectively modeling multi-scale information for audio-visual feature fusion remains challenging. We argue that the limited capacity of existing approaches primarily arises from: 1) treating features from different modalities in the same manner, and 2) overlooking the role of global features. To address these issues, we propose a Global-guided Asymmetric Attention Network (GAANet). Our model introduces two core innovations: first, an asymmetric multi-scale fusion framework that allows audio and visual streams to extract and interact with features at their respective optimal temporal resolutions, removing the need for symmetric temporal downsampling; second, a global-guided attention mechanism that compresses each modality into a compact global token with a temporal dimension of one, which then provides high-level semantic cues to guide both intra- and inter-modal fusion across scales. Experiments on LRS2 and VoxCeleb2 demonstrate that GAANet achieves state-of-the-art performance, reaching 16.5 dB SI-SNRi on LRS2 and 14.0 dB on VoxCeleb2, while maintaining a lightweight computational profile with only 3.3M parameters and 19.8G MACs. These results highlight the strong potential of asymmetric temporal modeling and global guidance for efficient and robust multimodal fusion. The source code is publicly accessible at https://github.com/redizzy/GAANet

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhiyuan Zhang, Jingyuan Xu, Yiming Tang, Liu Liu, Dan Guo. 2026-10-02. GAANet: Global-guided Asymmetric Attention Network for Audio-Visual Speech Separation. https://arxiv.org/abs/2610.02752

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Towards Unified Music Emotion Recognition across Dimensional and Categorical Models

One of the most significant challenges in Music Emotion Recognition (MER) comes from the fact that emotion labels can be heterogeneous across datasets with regard to the emotion representation, including categorical (e.g., happy, sad) versus dimensional labels (e.g., valence-arousal). In this paper, we present a unified multitask learning framework that combines these two types of labels and is thus able to be trained on multiple datasets. This framework uses an effective input representation that combines musical features (i.e., key and chords) and MERT embeddings. Moreover, knowledge distillation is employed to transfer the knowledge of teacher models trained on individual datasets to a student model, enhancing its ability to generalize across multiple tasks. To validate our proposed framework, we conducted extensive experiments on a variety of datasets, including MTG-Jamendo, DEAM, PMEmo, and EmoMusic. According to our experimental results, the inclusion of musical features, multitask learning, and knowledge distillation significantly enhances performance. In particular, our model outperforms the state-of-the-art models, including the best-performing model from the MediaEval 2021 competition on the MTG-Jamendo dataset. Our work makes a significant contribution to MER by allowing the combination of categorical and dimensional emotion labels in one unified framework, thus enabling training across datasets.

cs.SD↗

Controllable Embedding Transformation for Mood-Guided Music Retrieval

Music representations are the backbone of modern recommendation systems, powering playlist generation, similarity search, and personalized discovery. Yet most embeddings offer little control for adjusting a single musical attribute, e.g., changing only the mood of a track while preserving its genre or instrumentation. In this work, we address the problem of controllable music retrieval through embedding-based transformation, where the objective is to retrieve songs that remain similar to a seed track but are modified along one chosen dimension. We propose a novel framework for mood-guided music embedding transformation, which learns a mapping from a seed audio embedding to a target embedding guided by mood labels, while preserving other musical attributes. Because mood cannot be directly altered in the seed audio, we introduce a sampling mechanism that retrieves proxy targets to balance diversity with similarity to the seed. We train a lightweight translation model using this sampling strategy and introduce a novel joint objective that encourages transformation and information preservation. Extensive experiments on two datasets show strong mood transformation performance while retaining genre and instrumentation far better than training-free baselines, establishing controllable embedding transformation as a promising paradigm for personalized music retrieval.

cs.SD↗

Hearing is Believing? Evaluating and Analyzing Audio Language Model Sycophancy with SYAUDIO

Audio Language Models (ALMs) have recently shown strong capabilities in unified reasoning over speech, sound, and natural language; yet we find that they can inherit sycophancy, the tendency to agree with user assertions even when they contradict objective evidence. This failure mode is especially concerning for audio-conditioned reasoning, where a model must preserve evidence from acoustic events, speaker characteristics, and speech rate while responding to potentially misleading user feedback. However, unlike text and vision-language sycophancy, ALM sycophancy has not been systematically studied. We therefore introduce SYAUDIO, the first benchmark dedicated to evaluating sycophancy in ALMs, consisting of 4,319 audio questions spanning Audio Perception, Audio Reasoning, Audio Math, and Audio Ethics. Built upon established audio benchmarks and augmented with TTS-generated arithmetic and moral reasoning tasks, SYAUDIO enables systematic evaluation across multiple domains and sycophancy types with a human-speaker validation. Using this benchmark, we identify substantial and audio-specific sycophancy patterns under realistic conditions involving noise and speech rate, and further show that supervised fine-tuning reduces misleading susceptibility while decode-time steering reveals controllable hidden-state directions.

cs.SD↗