Using a Large Language Model to Control Speaking Style for Expressive TTS

Sigurgeirsson, Atli Thor; King, Simon

Computer Science > Computation and Language

arXiv:2305.10321v1 (cs)

[Submitted on 17 May 2023 (this version), latest version 19 Sep 2023 (v2)]

Title:Using a Large Language Model to Control Speaking Style for Expressive TTS

Authors:Atli Thor Sigurgeirsson, Simon King

View PDF

Abstract:Appropriate prosody is critical for successful spoken communication. Contextual word embeddings are proven to be helpful in predicting prosody but do not allow for choosing between plausible prosodic renditions. Reference-based TTS models attempt to address this by conditioning speech generation on a reference speech sample. These models can generate expressive speech but this requires finding an appropriate reference.
Sufficiently large generative language models have been used to solve various language-related tasks. We explore whether such models can be used to suggest appropriate prosody for expressive TTS. We train a TTS model on a non-expressive corpus and then prompt the language model to suggest changes to pitch, energy and duration. The prompt can be designed for any task and we prompt the model to make suggestions based on target speaking style and dialogue context. The proposed method is rated most appropriate in 49.9\% of cases compared to 31.0\% for a baseline model.

Comments:	Submitted to Interspeech 2023
Subjects:	Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2305.10321 [cs.CL]
	(or arXiv:2305.10321v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2305.10321

Submission history

From: Atli Sigurgeirsson [view email]
[v1] Wed, 17 May 2023 16:01:50 UTC (215 KB)
[v2] Tue, 19 Sep 2023 16:35:57 UTC (237 KB)

Computer Science > Computation and Language

Title:Using a Large Language Model to Control Speaking Style for Expressive TTS

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Using a Large Language Model to Control Speaking Style for Expressive TTS

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators