Search arXiv⌕ Search

arXiv · 2610.03016

From Expression to Reaction: Role-aware Visual Transfer and Stimulus-guided Reasoning for Interlocutor Emotion Recognition

Abstract

In this paper, we propose a Role-aware Stimulus-guided (RASG) framework for interlocutor emotion recognition, which predicts listener emotions from listener-only videos and speaker-only audios. RASG consists of Role-aware Visual Transfer (RVT) and Stimulus-guided Boundary Reasoning (SBR) modules, which address supervision mismatch due to the lack of labeled listener data and ambiguity among visually similar listener reactions whose interpretation depends on speaker context, respectively. More specifically, RVT selects speaker samples whose facial expressions support their emotion labels. It then filters listener tracks and uses reliable pseudo-labels to train a listener-centric visual expert. SBR uses a two-class language reasoner only when the visual model is uncertain. It treats speaker audio and text as context rather than direct emotion evidence to distinguish similar listener reactions. Experiments conducted on MER-Cross dataset shows that RASG achieves 76.25\% on MER-Cross and improves the performance of the baseline over 17\%. Our team ranks second in Track 1 (MER-Cross) of the MER Grand Challenge at ACM MM 2026.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wei Wang, Zhaowu Li, Jianjie Luo, Fu Lee Wang, Lap-Kei Lee, Zhenguo Yang. 2026-10-02. From Expression to Reaction: Role-aware Visual Transfer and Stimulus-guided Reasoning for Interlocutor Emotion Recognition. https://arxiv.org/abs/2610.03016

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Supporting Perspective Acquisition and Opinion Formation on Societal Issues Through AI-Generated Japanese Rap Battle Debates

Acquiring diverse perspectives and forming informed opinions on societal issues are essential for critical thinking and informed decision-making. Although observing debates between opposing viewpoints can promote perspective acquisition, conventional debates require substantial time and human resources, limiting their accessibility. This study investigates whether brief AI-generated rap battle-style debates can support perspective acquisition and opinion formation on societal issues. Rap battles are debate-style performances in which opposing viewpoints are expressed through rhymed verses, making them engaging while preserving the structure of argumentative dialogue. To conduct this investigation, we automatically generated rap battle-style debates using GPT-5.4-mini and presented them to participants with synthesized audio through a web-based viewing interface. We evaluated the educational effectiveness of this approach through a user study involving six participants who viewed 40 debates generated from topics on societal issues used in the All Japan Junior and Senior High School Debate Championships. The results showed that watching approximately 1.5 minute rap battle-style debates nearly doubled the number of opinions and discussion points identified by participants. Furthermore, participants reconsidered or changed their stances in 52.5% of the evaluated cases. These findings demonstrate that even brief AI-generated rap battle-style debates can effectively support perspective acquisition and opinion formation on societal issues.

cs.MM↗

Enhancing Relation Modeling with Social Attributes for Social Media Popularity Prediction

Recent studies highlight the critical role of retrieval-augmented mechanisms in social media popularity prediction (SMPP). Although such frameworks have improved SMPP performance by leveraging historical posts, existing methods still suffer from the low retrieval accuracy due to the oversight of relative relationships among UGC instances. To address this limitation, we propose a novel Relation-Enhanced Retrieval-Augmented framework (RE-Rag) that models UGC similarity as a continuous relation jointly driven by semantic content and social attributes. Specifically, RE-Rag employs a Semantic-Attribute Retriever (SAR) to obtain instances aligned in both semantic and social-attribute distributions. Subsequently, we design a Relation-Guided Predictor (RGP): first, cross-attention encodes multimodal features of retrieved instances; then, a relative relation graph is introduced to guide attention weight allocation, forming a Relation-Guided Transformer (RGTs) that dynamically modulate attention weights based on relative attribute relations to capture the interplay between semantics and various social attributes. The refined features are fused with the target instance for popularity prediction. Experiments on three public benchmarks show that RE-Rag consistently outperforms state-of-the-art methods in both prediction accuracy and retrieval efficiency.

cs.MM↗

PrecipJEPA: JEPA-Regularized Future-State Prediction with Motion-Source Rendering for Precipitation Nowcasting

Long-term precipitation nowcasting requires modeling radar-echo evolution while preserving localized high-intensity structures. Recent radar-specific studies motivate location-aware prediction and separating echo displacement from intensity change. However existing encoders learn historical representations mainly from final forecast errors. We propose PrecipJEPA, which couples a structured forecasting path with an auxiliary path that enriches its encoder from observed radar history. In the forecasting path, an online encoder first converts the observations into spatiotemporal tokens. The Task-Driven Future-State Predictor (TFP) combines these tokens with a recent-dynamics summary and spatiotemporal queries to construct future radar states. The Parallel Motion-Source Renderer (PMSR) decodes these states into motion and source-sink fields that transform the latest observation into future frames. During joint training, the History-Masked JEPA (H-JEPA) operates on the auxiliary path to predict masked historical features from visible context, directly supervising the same online encoder from the observed sequence. Experiments on SEVIR and MeteoNet show that PrecipJEPA improves highest-threshold CSI by 118.6% and 35.1%, respectively, over the strongest baselines, while maintaining the highest mean CSI throughout the 3-hour forecast.

cs.MM↗