Self-Supervised Representations for Singing Voice Conversion

Jayashankar, Tejas; Wu, Jilong; Sari, Leda; Kant, David; Manohar, Vimal; He, Qing

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2303.12197 (eess)

[Submitted on 21 Mar 2023]

Title:Self-Supervised Representations for Singing Voice Conversion

Authors:Tejas Jayashankar, Jilong Wu, Leda Sari, David Kant, Vimal Manohar, Qing He

View PDF

Abstract:A singing voice conversion model converts a song in the voice of an arbitrary source singer to the voice of a target singer. Recently, methods that leverage self-supervised audio representations such as HuBERT and Wav2Vec 2.0 have helped further the state-of-the-art. Though these methods produce more natural and melodic singing outputs, they often rely on confusion and disentanglement losses to render the self-supervised representations speaker and pitch-invariant. In this paper, we circumvent disentanglement training and propose a new model that leverages ASR fine-tuned self-supervised representations as inputs to a HiFi-GAN neural vocoder for singing voice conversion. We experiment with different f0 encoding schemes and show that an f0 harmonic generation module that uses a parallel bank of transposed convolutions (PBTC) alongside ASR fine-tuned Wav2Vec 2.0 features results in the best singing voice conversion quality. Additionally, the model is capable of making a spoken voice sing. We also show that a simple f0 shifting scheme during inference helps retain singer identity and bolsters the performance of our singing voice conversion model. Our results are backed up by extensive MOS studies that compare different ablations and baselines.

Subjects:	Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2303.12197 [eess.AS]
	(or arXiv:2303.12197v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2303.12197

Submission history

From: Vimal Manohar [view email]
[v1] Tue, 21 Mar 2023 21:04:03 UTC (520 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Self-Supervised Representations for Singing Voice Conversion

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Self-Supervised Representations for Singing Voice Conversion

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators