Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator

Sun, Guangzhi; Zhang, Chao; Woodland, Philip C

Computer Science > Computation and Language

arXiv:2205.09058 (cs)

[Submitted on 18 May 2022 (v1), last revised 23 May 2022 (this version, v2)]

Title:Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator

Authors:Guangzhi Sun, Chao Zhang, Philip C Woodland

View PDF

Abstract:Contextual knowledge is essential for reducing speech recognition errors on high-valued long-tail words. This paper proposes a novel tree-constrained pointer generator (TCPGen) component that enables end-to-end ASR models to bias towards a list of long-tail words obtained using external contextual information. With only a small overhead in memory use and computation cost, TCPGen can structure thousands of biasing words efficiently into a symbolic prefix-tree and creates a neural shortcut between the tree and the final ASR output to facilitate the recognition of the biasing words. To enhance TCPGen, we further propose a novel minimum biasing word error (MBWE) loss that directly optimises biasing word errors during training, along with a biasing-word-driven language model discounting (BLMD) method during the test. All contextual ASR systems were evaluated on the public Librispeech audiobook corpus and the data from the dialogue state tracking challenges (DSTC) with the biasing lists extracted from the dialogue-system ontology. Consistent word error rate (WER) reductions were achieved with TCPGen, which were particularly significant on the biasing words with around 40\% relative reductions in the recognition error rates. MBWE and BLMD further improved the effectiveness of TCPGen and achieved more significant WER reductions on the biasing words. TCPGen also achieved zero-shot learning of words not in the audio training set with large WER reductions on the out-of-vocabulary words in the biasing list.

Comments:	This work has been submitted to the IEEE Transactions on Audio, Speech, and Language Processing for possible publication
Subjects:	Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2205.09058 [cs.CL]
	(or arXiv:2205.09058v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2205.09058

Submission history

From: Guangzhi Sun [view email]
[v1] Wed, 18 May 2022 16:40:50 UTC (2,462 KB)
[v2] Mon, 23 May 2022 19:48:26 UTC (2,462 KB)

Computer Science > Computation and Language

Title:Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators