Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations

Chen, Jinming; Fang, Jingyi; Zheng, Yuanzhong; Wang, Yaoxuan; Fei, Haojun

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2503.22687 (eess)

[Submitted on 5 Mar 2025]

Title:Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations

Authors:Jinming Chen, Jingyi Fang, Yuanzhong Zheng, Yaoxuan Wang, Haojun Fei

View PDF HTML (experimental)

Abstract:Emotion recognition plays a pivotal role in intelligent human-machine interaction systems. Multimodal approaches benefit from the fusion of diverse modalities, thereby improving the recognition accuracy. However, the lack of high-quality multimodal data and the challenge of achieving optimal alignment between different modalities significantly limit the potential for improvement in multimodal approaches. In this paper, the proposed Qieemo framework effectively utilizes the pretrained automatic speech recognition (ASR) model backbone which contains naturally frame aligned textual and emotional features, to achieve precise emotion classification solely based on the audio modality. Furthermore, we design the multimodal fusion (MMF) module and cross-modal attention (CMA) module in order to fuse the phonetic posteriorgram (PPG) and emotional features extracted by the ASR encoder for improving recognition accuracy. The experimental results on the IEMOCAP dataset demonstrate that Qieemo outperforms the benchmark unimodal, multimodal, and self-supervised models with absolute improvements of 3.0%, 1.2%, and 1.9% respectively.

Subjects:	Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2503.22687 [eess.AS]
	(or arXiv:2503.22687v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2503.22687

Submission history

From: Jinming Chen [view email]
[v1] Wed, 5 Mar 2025 07:02:30 UTC (2,016 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators