Incremental Text to Speech for Neural Sequence-to-Sequence Models using Reinforcement Learning

Mohan, Devang S Ram; Lenain, Raphael; Foglianti, Lorenzo; Teh, Tian Huey; Staib, Marlene; Torresquintero, Alexandra; Gao, Jiameng

doi:10.21437/Interspeech.2020-1822

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2008.03096 (eess)

[Submitted on 7 Aug 2020]

Title:Incremental Text to Speech for Neural Sequence-to-Sequence Models using Reinforcement Learning

Authors:Devang S Ram Mohan, Raphael Lenain, Lorenzo Foglianti, Tian Huey Teh, Marlene Staib, Alexandra Torresquintero, Jiameng Gao

View PDF

Abstract:Modern approaches to text to speech require the entire input character sequence to be processed before any audio is synthesised. This latency limits the suitability of such models for time-sensitive tasks like simultaneous interpretation. Interleaving the action of reading a character with that of synthesising audio reduces this latency. However, the order of this sequence of interleaved actions varies across sentences, which raises the question of how the actions should be chosen. We propose a reinforcement learning based framework to train an agent to make this decision. We compare our performance against that of deterministic, rule-based systems. Our results demonstrate that our agent successfully balances the trade-off between the latency of audio generation and the quality of synthesised audio. More broadly, we show that neural sequence-to-sequence models can be adapted to run in an incremental manner.

Comments:	To be published in Interspeech 2020. 5 pages, 4 figures
Subjects:	Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD); Machine Learning (stat.ML)
Cite as:	arXiv:2008.03096 [eess.AS]
	(or arXiv:2008.03096v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2008.03096
Related DOI:	https://doi.org/10.21437/Interspeech.2020-1822

Submission history

From: Devang S Ram Mohan [view email]
[v1] Fri, 7 Aug 2020 11:48:05 UTC (257 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Incremental Text to Speech for Neural Sequence-to-Sequence Models using Reinforcement Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Incremental Text to Speech for Neural Sequence-to-Sequence Models using Reinforcement Learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators