Multi-level Fusion of Wav2vec 2.0 and BERT for Multimodal Emotion Recognition

Zhao, Zihan; Wang, Yanfeng; Wang, Yu

Computer Science > Computation and Language

arXiv:2207.04697 (cs)

[Submitted on 11 Jul 2022 (v1), last revised 12 Jul 2022 (this version, v2)]

Title:Multi-level Fusion of Wav2vec 2.0 and BERT for Multimodal Emotion Recognition

Authors:Zihan Zhao, Yanfeng Wang, Yu Wang

View PDF

Abstract:The research and applications of multimodal emotion recognition have become increasingly popular recently. However, multimodal emotion recognition faces the challenge of lack of data. To solve this problem, we propose to use transfer learning which leverages state-of-the-art pre-trained models including wav2vec 2.0 and BERT for this task. Multi-level fusion approaches including coattention-based early fusion and late fusion with the models trained on both embeddings are explored. Also, a multi-granularity framework which extracts not only frame-level speech embeddings but also segment-level embeddings including phone, syllable and word-level speech embeddings is proposed to further boost the performance. By combining our coattention-based early fusion model and late fusion model with the multi-granularity feature extraction framework, we obtain result that outperforms best baseline approaches by 1.3% unweighted accuracy (UA) on the IEMOCAP dataset.

Comments:	Accepted to INTERSPEECH 2022
Subjects:	Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2207.04697 [cs.CL]
	(or arXiv:2207.04697v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2207.04697

Submission history

From: Zihan Zhao [view email]
[v1] Mon, 11 Jul 2022 08:20:53 UTC (1,315 KB)
[v2] Tue, 12 Jul 2022 04:21:25 UTC (857 KB)

Computer Science > Computation and Language

Title:Multi-level Fusion of Wav2vec 2.0 and BERT for Multimodal Emotion Recognition

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Multi-level Fusion of Wav2vec 2.0 and BERT for Multimodal Emotion Recognition

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators