MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization

Zhu, Haina; Zhou, Yizhi; Chen, Hangting; Yu, Jianwei; Ma, Ziyang; Gu, Rongzhi; Luo, Yi; Tan, Wei; Chen, Xie

Computer Science > Sound

arXiv:2501.01108 (cs)

[Submitted on 2 Jan 2025 (v1), last revised 3 Jan 2025 (this version, v2)]

Title:MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization

Authors:Haina Zhu, Yizhi Zhou, Hangting Chen, Jianwei Yu, Ziyang Ma, Rongzhi Gu, Yi Luo, Wei Tan, Xie Chen

View PDF HTML (experimental)

Abstract:Recent years have witnessed the success of foundation models pre-trained with self-supervised learning (SSL) in various music informatics understanding tasks, including music tagging, instrument classification, key detection, and more. In this paper, we propose a self-supervised music representation learning model for music understanding. Distinguished from previous studies adopting random projection or existing neural codec, the proposed model, named MuQ, is trained to predict tokens generated by Mel Residual Vector Quantization (Mel-RVQ). Our Mel-RVQ utilizes residual linear projection structure for Mel spectrum quantization to enhance the stability and efficiency of target extraction and lead to better performance. Experiments in a large variety of downstream tasks demonstrate that MuQ outperforms previous self-supervised music representation models with only 0.9K hours of open-source pre-training data. Scaling up the data to over 160K hours and adopting iterative training consistently improve the model performance. To further validate the strength of our model, we present MuQ-MuLan, a joint music-text embedding model based on contrastive learning, which achieves state-of-the-art performance in the zero-shot music tagging task on the MagnaTagATune dataset. Code and checkpoints are open source in this https URL.

Subjects:	Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2501.01108 [cs.SD]
	(or arXiv:2501.01108v2 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2501.01108

Submission history

From: Haina Zhu [view email]
[v1] Thu, 2 Jan 2025 07:08:29 UTC (1,574 KB)
[v2] Fri, 3 Jan 2025 08:35:34 UTC (1,574 KB)

Computer Science > Sound

Title:MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators