A language model based approach towards large scale and lightweight language identification systems

Srivastava, Brij Mohan Lal; Vydana, Hari Krishna; Vuppala, Anil Kumar; Shrivastava, Manish

Computer Science > Sound

arXiv:1510.03602 (cs)

[Submitted on 13 Oct 2015]

Title:A language model based approach towards large scale and lightweight language identification systems

Authors:Brij Mohan Lal Srivastava, Hari Krishna Vydana, Anil Kumar Vuppala, Manish Shrivastava

View PDF

Abstract:Multilingual spoken dialogue systems have gained prominence in the recent past necessitating the requirement for a front-end Language Identification (LID) system. Most of the existing LID systems rely on modeling the language discriminative information from low-level acoustic features. Due to the variabilities of speech (speaker and emotional variabilities, etc.), large-scale LID systems developed using low-level acoustic features suffer from a degradation in the performance. In this approach, we have attempted to model the higher level language discriminative phonotactic information for developing an LID system. In this paper, the input speech signal is tokenized to phone sequences by using a language independent phone recognizer. The language discriminative phonotactic information in the obtained phone sequences are modeled using statistical and recurrent neural network based language modeling approaches. As this approach, relies on higher level phonotactical information it is more robust to variabilities of speech. Proposed approach is computationally light weight, highly scalable and it can be used in complement with the existing LID systems.

Comments:	Under review at ICASSP 2016
Subjects:	Sound (cs.SD); Computation and Language (cs.CL)
Cite as:	arXiv:1510.03602 [cs.SD]
	(or arXiv:1510.03602v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.1510.03602

Submission history

From: Brij Mohan Lal Srivastava [view email]
[v1] Tue, 13 Oct 2015 09:51:23 UTC (52 KB)

Computer Science > Sound

Title:A language model based approach towards large scale and lightweight language identification systems

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:A language model based approach towards large scale and lightweight language identification systems

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators