MASIVE: Open-Ended Affective State Identification in English and Spanish

Deas, Nicholas; Turcan, Elsbeth; Mejía, Iván Pérez; McKeown, Kathleen

Computer Science > Computation and Language

arXiv:2407.12196 (cs)

[Submitted on 16 Jul 2024]

Title:MASIVE: Open-Ended Affective State Identification in English and Spanish

Authors:Nicholas Deas, Elsbeth Turcan, Iván Pérez Mejía, Kathleen McKeown

View PDF HTML (experimental)

Abstract:In the field of emotion analysis, much NLP research focuses on identifying a limited number of discrete emotion categories, often applied across languages. These basic sets, however, are rarely designed with textual data in mind, and culture, language, and dialect can influence how particular emotions are interpreted. In this work, we broaden our scope to a practically unbounded set of \textit{affective states}, which includes any terms that humans use to describe their experiences of feeling. We collect and publish MASIVE, a dataset of Reddit posts in English and Spanish containing over 1,000 unique affective states each. We then define the new problem of \textit{affective state identification} for language generation models framed as a masked span prediction task. On this task, we find that smaller finetuned multilingual models outperform much larger LLMs, even on region-specific Spanish affective states. Additionally, we show that pretraining on MASIVE improves model performance on existing emotion benchmarks. Finally, through machine translation experiments, we find that native speaker-written data is vital to good performance on this task.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2407.12196 [cs.CL]
	(or arXiv:2407.12196v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2407.12196

Submission history

From: Nicholas Deas [view email]
[v1] Tue, 16 Jul 2024 21:43:47 UTC (497 KB)

Computer Science > Computation and Language

Title:MASIVE: Open-Ended Affective State Identification in English and Spanish

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:MASIVE: Open-Ended Affective State Identification in English and Spanish

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators