Learning Audio-Text Agreement for Open-vocabulary Keyword Spotting

Shin, Hyeon-Kyeong; Han, Hyewon; Kim, Doyeon; Chung, Soo-Whan; Kang, Hong-Goo

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2206.15400 (eess)

[Submitted on 30 Jun 2022 (v1), last revised 1 Jul 2022 (this version, v2)]

Title:Learning Audio-Text Agreement for Open-vocabulary Keyword Spotting

Authors:Hyeon-Kyeong Shin, Hyewon Han, Doyeon Kim, Soo-Whan Chung, Hong-Goo Kang

View PDF

Abstract:In this paper, we propose a novel end-to-end user-defined keyword spotting method that utilizes linguistically corresponding patterns between speech and text sequences. Unlike previous approaches requiring speech keyword enrollment, our method compares input queries with an enrolled text keyword sequence. To place the audio and text representations within a common latent space, we adopt an attention-based cross-modal matching approach that is trained in an end-to-end manner with monotonic matching loss and keyword classification loss. We also utilize a de-noising loss for the acoustic embedding network to improve robustness in noisy environments. Additionally, we introduce the LibriPhrase dataset, a new short-phrase dataset based on LibriSpeech for efficiently training keyword spotting models. Our proposed method achieves competitive results on various evaluation sets compared to other single-modal and cross-modal baselines.

Comments:	Accepted to Interspeech 2022
Subjects:	Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2206.15400 [eess.AS]
	(or arXiv:2206.15400v2 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2206.15400

Submission history

From: Hyeon-Kyeong Shin [view email]
[v1] Thu, 30 Jun 2022 16:40:31 UTC (985 KB)
[v2] Fri, 1 Jul 2022 06:42:55 UTC (914 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Learning Audio-Text Agreement for Open-vocabulary Keyword Spotting

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Learning Audio-Text Agreement for Open-vocabulary Keyword Spotting

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators