End-to-End Label Uncertainty Modeling in Speech Emotion Recognition using Bayesian Neural Networks and Label Distribution Learning

Prabhu, Navin Raj; Lehmann-Willenbrock, Nale; Gerkman, Timo

doi:10.1109/TAFFC.2023.3283595

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2209.15449 (eess)

[Submitted on 30 Sep 2022 (v1), last revised 13 Jun 2023 (this version, v2)]

Title:End-to-End Label Uncertainty Modeling in Speech Emotion Recognition using Bayesian Neural Networks and Label Distribution Learning

Authors:Navin Raj Prabhu, Nale Lehmann-Willenbrock, Timo Gerkman

View PDF

Abstract:To train machine learning algorithms to predict emotional expressions in terms of arousal and valence, annotated datasets are needed. However, as different people perceive others' emotional expressions differently, their annotations are subjective. To account for this, annotations are typically collected from multiple annotators and averaged to obtain ground-truth labels. However, when exclusively trained on this averaged ground-truth, the model is agnostic to the inherent subjectivity in emotional expressions. In this work, we therefore propose an end-to-end Bayesian neural network capable of being trained on a distribution of annotations to also capture the subjectivity-based label uncertainty. Instead of a Gaussian, we model the annotation distribution using Student's t-distribution, which also accounts for the number of annotations available. We derive the corresponding Kullback-Leibler divergence loss and use it to train an estimator for the annotation distribution, from which the mean and uncertainty can be inferred. We validate the proposed method using two in-the-wild datasets. We show that the proposed t-distribution based approach achieves state-of-the-art uncertainty modeling results in speech emotion recognition, and also consistent results in cross-corpora evaluations. Furthermore, analyses reveal that the advantage of a t-distribution over a Gaussian grows with increasing inter-annotator correlation and a decreasing number of annotations available.

Comments:	Accepted Paper at IEEE Transactions on Affective Computing, June 2023. Contains main paper with supplementary material. arXiv admin note: text overlap with arXiv:2207.12135
Subjects:	Audio and Speech Processing (eess.AS); Machine Learning (cs.LG)
Cite as:	arXiv:2209.15449 [eess.AS]
	(or arXiv:2209.15449v2 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2209.15449
Journal reference:	IEEE Transactions on Affective Computing, June 2023
Related DOI:	https://doi.org/10.1109/TAFFC.2023.3283595

Submission history

From: Navin Raj Prabhu [view email]
[v1] Fri, 30 Sep 2022 12:55:43 UTC (11,144 KB)
[v2] Tue, 13 Jun 2023 08:55:11 UTC (12,243 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:End-to-End Label Uncertainty Modeling in Speech Emotion Recognition using Bayesian Neural Networks and Label Distribution Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:End-to-End Label Uncertainty Modeling in Speech Emotion Recognition using Bayesian Neural Networks and Label Distribution Learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators