Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification

Li, Jingyu; Tian, Yusheng; Lee, Tan

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2210.17310 (eess)

[Submitted on 31 Oct 2022]

Title:Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification

Authors:Jingyu Li, Yusheng Tian, Tan Lee

View PDF

Abstract:Deep convolutional neural networks (CNNs) have been applied to extracting speaker embeddings with significant success in speaker verification. Incorporating the attention mechanism has shown to be effective in improving the model performance. This paper presents an efficient two-dimensional convolution-based attention module, namely C2D-Att. The interaction between the convolution channel and frequency is involved in the attention calculation by lightweight convolution layers. This requires only a small number of parameters. Fine-grained attention weights are produced to represent channel and frequency-specific information. The weights are imposed on the input features to improve the representation ability for speaker modeling. The C2D-Att is integrated into a modified version of ResNet for speaker embedding extraction. Experiments are conducted on VoxCeleb datasets. The results show that C2DAtt is effective in generating discriminative attention maps and outperforms other attention methods. The proposed model shows robust performance with different scales of model size and achieves state-of-the-art results.

Subjects:	Audio and Speech Processing (eess.AS); Sound (cs.SD)
Cite as:	arXiv:2210.17310 [eess.AS]
	(or arXiv:2210.17310v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2210.17310

Submission history

From: Jingyu Li [view email]
[v1] Mon, 31 Oct 2022 13:34:22 UTC (463 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators