VaSAB: The variable size adaptive information bottleneck for disentanglement on speech and singing voice

Bous, Frederik; Roebel, Axel

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2310.03444 (eess)

[Submitted on 5 Oct 2023]

Title:VaSAB: The variable size adaptive information bottleneck for disentanglement on speech and singing voice

Authors:Frederik Bous, Axel Roebel

View PDF

Abstract:The information bottleneck auto-encoder is a tool for disentanglement commonly used for voice transformation. The successful disentanglement relies on the right choice of bottleneck size. Previous bottleneck auto-encoders created the bottleneck by the dimension of the latent space or through vector quantization and had no means to change the bottleneck size of a specific model. As the bottleneck removes information from the disentangled representation, the choice of bottleneck size is a trade-off between disentanglement and synthesis quality. We propose to build the information bottleneck using dropout which allows us to change the bottleneck through the dropout rate and investigate adapting the bottleneck size depending on the context. We experimentally explore into using the adaptive bottleneck for pitch transformation and demonstrate that the adaptive bottleneck leads to improved disentanglement of the F0 parameter for both, speech and singing voice leading to improved synthesis quality. Using the variable bottleneck size, we were able to achieve disentanglement for singing voice including extremely high pitches and create a universal voice model, that works on both speech and singing voice with improved synthesis quality.

Comments:	Submitted to ICASSP 2024
Subjects:	Audio and Speech Processing (eess.AS); Sound (cs.SD)
Cite as:	arXiv:2310.03444 [eess.AS]
	(or arXiv:2310.03444v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2310.03444

Submission history

From: Frederik Bous [view email]
[v1] Thu, 5 Oct 2023 10:30:33 UTC (158 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:VaSAB: The variable size adaptive information bottleneck for disentanglement on speech and singing voice

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:VaSAB: The variable size adaptive information bottleneck for disentanglement on speech and singing voice

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators