Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario

Medennikov, Ivan; Korenevsky, Maxim; Prisyach, Tatiana; Khokhlov, Yuri; Korenevskaya, Mariya; Sorokin, Ivan; Timofeeva, Tatiana; Mitrofanov, Anton; Andrusenko, Andrei; Podluzhny, Ivan; Laptev, Aleksandr; Romanenko, Aleksei

doi:10.21437/Interspeech.2020-1602

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2005.07272 (eess)

[Submitted on 14 May 2020 (v1), last revised 27 Jul 2020 (this version, v2)]

Title:Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario

Authors:Ivan Medennikov, Maxim Korenevsky, Tatiana Prisyach, Yuri Khokhlov, Mariya Korenevskaya, Ivan Sorokin, Tatiana Timofeeva, Anton Mitrofanov, Andrei Andrusenko, Ivan Podluzhny, Aleksandr Laptev, Aleksei Romanenko

View PDF

Abstract:Speaker diarization for real-life scenarios is an extremely challenging problem. Widely used clustering-based diarization approaches perform rather poorly in such conditions, mainly due to the limited ability to handle overlapping speech. We propose a novel Target-Speaker Voice Activity Detection (TS-VAD) approach, which directly predicts an activity of each speaker on each time frame. TS-VAD model takes conventional speech features (e.g., MFCC) along with i-vectors for each speaker as inputs. A set of binary classification output layers produces activities of each speaker. I-vectors can be estimated iteratively, starting with a strong clustering-based diarization. We also extend the TS-VAD approach to the multi-microphone case using a simple attention mechanism on top of hidden representations extracted from the single-channel TS-VAD model. Moreover, post-processing strategies for the predicted speaker activity probabilities are investigated. Experiments on the CHiME-6 unsegmented data show that TS-VAD achieves state-of-the-art results outperforming the baseline x-vector-based system by more than 30% Diarization Error Rate (DER) abs.

Comments:	Accepted to Interspeech 2020
Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
Cite as:	arXiv:2005.07272 [eess.AS]
	(or arXiv:2005.07272v2 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2005.07272
Related DOI:	https://doi.org/10.21437/Interspeech.2020-1602

Submission history

From: Ivan Medennikov [view email]
[v1] Thu, 14 May 2020 21:24:56 UTC (142 KB)
[v2] Mon, 27 Jul 2020 13:28:20 UTC (141 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators