Search arXiv⌕ Search

arXiv · 2009.04107

Multi-modal Attention for Speech Emotion Recognition

Abstract

Emotion represents an essential aspect of human speech that is manifested in speech prosody. Speech, visual, and textual cues are complementary in human communication. In this paper, we study a hybrid fusion method, referred to as multi-modal attention network (MMAN) to make use of visual and textual cues in speech emotion recognition. We propose a novel multi-modal attention mechanism, cLSTM-MMA, which facilitates the attention across three modalities and selectively fuse the information. cLSTM-MMA is fused with other uni-modal sub-networks in the late fusion. The experiments show that speech emotion recognition benefits significantly from visual and textual cues, and the proposed cLSTM-MMA alone is as competitive as other fusion methods in terms of accuracy, but with a much more compact network structure. The proposed hybrid network MMAN achieves state-of-the-art performance on IEMOCAP database for emotion recognition.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zexu Pan, Zhaojie Luo, Jichen Yang, Haizhou Li. 2020-09-09. Multi-modal Attention for Speech Emotion Recognition. https://arxiv.org/abs/2009.04107

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

HuPER: A Human-Inspired Framework for Phonetic Perception

We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic-phonetics evidence and linguistic knowledge. With only 100 hours of training data, HuPER achieves state-of-the-art phonetic error rates on five English benchmarks and strong zero-shot transfer to 95 unseen languages. HuPER is also the first framework to enable adaptive, multi-path phonetic perception under diverse acoustic conditions. All training data, models, and code are open-sourced. Code and demo avaliable at https://github.com/Berkeley-Speech-Group/HuPER.

eess.AS↗

SHroom: A Python Framework for Ambisonics Room Acoustics Simulation and Binaural Rendering

Spatial audio research for virtual and augmented reality, teleconferencing and hearing devices often represents sound fields in the Spherical Harmonics (SH) domain, known as Ambisonics. A typical study simulates a room, renders what a listener or a microphone array would capture in it, and processes those signals in the SH domain. We present SHroom (Spherical Harmonics ROOM), an open-source Python library that performs this whole workflow in one package, from room simulation to binaural rendering, head rotation, microphone-array simulation and Ambisonics encoding. Existing tools cover either the room simulation or the downstream processing, so researchers bridge them with ad-hoc code that is hard to reproduce and compare. In SHroom, every step operates on one shared signal type through one processing interface, so a simulated room flows through the complete chain without format conversions. Built on the image-source engine of `pyroomacoustics`, SHroom reproduces the Ambisonic Room Impulse Response (ARIR) of its spherical-harmonic receivers while computing it about 3x faster for SH orders 4 to 12. SHroom is available at 'https://github.com/Yhonatangayer/shroom' and installable via `pip install pyshroom`.

eess.AS↗

Exploring Second-Order Pattern Recognition in Speaker Recognition

In traditional pattern recognition tasks, neural networks are trained to recognise human-defined patterns (e.g. audio categories) in model inputs (e.g. audio). Meanwhile, some Explainable AI (XAI) methods explain latent patterns characterising the network's recognition of inputs as human-defined patterns; this work calls these latent patterns second-order patterns and proposes to discover them. Accordingly, we apply a hierarchical clustering algorithm to analyse whether our speaker recognition network's representations learned from known utterances naturally form hierarchical clusters. Each resulting cluster is a second-order pattern that characterises a context in our network's recognition of the known utterances as speaker identities. All discovered second-order patterns are then interpreted using the Hierarchical Cluster-Class Matching (HCCM) method. Moreover, we propose a new task, second-order pattern recognition, to identify which of the discovered second-order patterns characterising the recognition of known utterances also apply to unseen utterances, thereby characterising the recognition of unseen utterances. Accordingly, we design the Hierarchical Cluster Navigation and Assignment (HCNA) method. HCNA recognises a second-order pattern as applying to an unseen utterance when the utterance's network representation lies within the extrapolation space of the cluster regarded as that second-order pattern. Experimental results demonstrate that the extrapolation space introduced in HCNA substantially improves task performance.

eess.AS↗