Toward Corpus Size Requirements for Training and Evaluating Depression Risk Models Using Spoken Language

Rutowski, Tomek; Harati, Amir; Shriberg, Elizabeth; Lu, Yang; Chlebek, Piotr; Oliveira, Ricardo

Computer Science > Computation and Language

arXiv:2501.00617 (cs)

[Submitted on 31 Dec 2024]

Title:Toward Corpus Size Requirements for Training and Evaluating Depression Risk Models Using Spoken Language

Authors:Tomek Rutowski, Amir Harati, Elizabeth Shriberg, Yang Lu, Piotr Chlebek, Ricardo Oliveira

View PDF

Abstract:Mental health risk prediction is a growing field in the speech community, but many studies are based on small corpora. This study illustrates how variations in test and train set sizes impact performance in a controlled study. Using a corpus of over 65K labeled data points, results from a fully crossed design of different train/test size combinations are provided. Two model types are included: one based on language and the other on speech acoustics. Both use methods current in this domain. An age-mismatched test set was also included. Results show that (1) test sizes below 1K samples gave noisy results, even for larger training set sizes; (2) training set sizes of at least 2K were needed for stable results; (3) NLP and acoustic models behaved similarly with train/test size variations, and (4) the mismatched test set showed the same patterns as the matched test set. Additional factors are discussed, including label priors, model strength and pre-training, unique speakers, and data lengths. While no single study can specify exact size requirements, results demonstrate the need for appropriately sized train and test sets for future studies of mental health risk prediction from speech and language.

Subjects:	Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2501.00617 [cs.CL]
	(or arXiv:2501.00617v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2501.00617
Journal reference:	Proceedings Interspeech, 2022

Submission history

From: Elizabeth Shriberg [view email]
[v1] Tue, 31 Dec 2024 19:32:25 UTC (308 KB)

Computer Science > Computation and Language

Title:Toward Corpus Size Requirements for Training and Evaluating Depression Risk Models Using Spoken Language

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Toward Corpus Size Requirements for Training and Evaluating Depression Risk Models Using Spoken Language

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators