Efficient Neural Audio Synthesis

Kalchbrenner, Nal; Elsen, Erich; Simonyan, Karen; Noury, Seb; Casagrande, Norman; Lockhart, Edward; Stimberg, Florian; Oord, Aaron van den; Dieleman, Sander; Kavukcuoglu, Koray

Computer Science > Sound

arXiv:1802.08435 (cs)

[Submitted on 23 Feb 2018 (v1), last revised 25 Jun 2018 (this version, v2)]

Title:Efficient Neural Audio Synthesis

Authors:Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron van den Oord, Sander Dieleman, Koray Kavukcuoglu

View PDF

Abstract:Sequential models achieve state-of-the-art results in audio, visual and textual domains with respect to both estimating the data distribution and generating high-quality samples. Efficient sampling for this class of models has however remained an elusive problem. With a focus on text-to-speech synthesis, we describe a set of general techniques for reducing sampling time while maintaining high output quality. We first describe a single-layer recurrent neural network, the WaveRNN, with a dual softmax layer that matches the quality of the state-of-the-art WaveNet model. The compact form of the network makes it possible to generate 24kHz 16-bit audio 4x faster than real time on a GPU. Second, we apply a weight pruning technique to reduce the number of weights in the WaveRNN. We find that, for a constant number of parameters, large sparse networks perform better than small dense networks and this relationship holds for sparsity levels beyond 96%. The small number of weights in a Sparse WaveRNN makes it possible to sample high-fidelity audio on a mobile CPU in real time. Finally, we propose a new generation scheme based on subscaling that folds a long sequence into a batch of shorter sequences and allows one to generate multiple samples at once. The Subscale WaveRNN produces 16 samples per step without loss of quality and offers an orthogonal method for increasing sampling efficiency.

Comments:	10 pages
Subjects:	Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:1802.08435 [cs.SD]
	(or arXiv:1802.08435v2 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.1802.08435

Submission history

From: Nal Kalchbrenner [view email]
[v1] Fri, 23 Feb 2018 08:20:23 UTC (804 KB)
[v2] Mon, 25 Jun 2018 19:45:25 UTC (784 KB)

Computer Science > Sound

Title:Efficient Neural Audio Synthesis

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Efficient Neural Audio Synthesis

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators