Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR

Lu, Xugang; Shen, Peng; Tsao, Yu; Kawai, Hisashi

Computer Science > Sound

arXiv:2409.02239v1 (cs)

[Submitted on 3 Sep 2024 (this version), latest version 5 Sep 2024 (v2)]

Title:Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR

Authors:Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai

View PDF HTML (experimental)

Abstract:Transferring linguistic knowledge from a pretrained language model (PLM) to an acoustic model has been shown to greatly improve the performance of automatic speech recognition (ASR). However, due to the heterogeneous feature distributions in cross-modalities, designing an effective model for feature alignment and knowledge transfer between linguistic and acoustic sequences remains a challenging task. Optimal transport (OT), which efficiently measures probability distribution discrepancies, holds great potential for aligning and transferring knowledge between acoustic and linguistic modalities. Nonetheless, the original OT treats acoustic and linguistic feature sequences as two unordered sets in alignment and neglects temporal order information during OT coupling estimation. Consequently, a time-consuming pretraining stage is required to learn a good alignment between the acoustic and linguistic representations. In this paper, we propose a Temporal Order Preserved OT (TOT)-based Cross-modal Alignment and Knowledge Transfer (CAKT) (TOT-CAKT) for ASR. In the TOT-CAKT, local neighboring frames of acoustic sequences are smoothly mapped to neighboring regions of linguistic sequences, preserving their temporal order relationship in feature alignment and matching. With the TOT-CAKT model framework, we conduct Mandarin ASR experiments with a pretrained Chinese PLM for linguistic knowledge transfer. Our results demonstrate that the proposed TOT-CAKT significantly improves ASR performance compared to several state-of-the-art models employing linguistic knowledge transfer, and addresses the weaknesses of the original OT-based method in sequential feature alignment for ASR.

Comments:	Accepted to IEEE SLT 2024
Subjects:	Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2409.02239 [cs.SD]
	(or arXiv:2409.02239v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2409.02239

Submission history

From: Yu Tsao [view email]
[v1] Tue, 3 Sep 2024 19:11:15 UTC (245 KB)
[v2] Thu, 5 Sep 2024 11:34:00 UTC (268 KB)

Computer Science > Sound

Title:Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators