Linguistically-driven Multi-task Pre-training for Low-resource Neural Machine Translation

Mao, Zhuoyuan; Chu, Chenhui; Kurohashi, Sadao

doi:10.1145/3491065

Computer Science > Computation and Language

arXiv:2201.08070 (cs)

[Submitted on 20 Jan 2022]

Title:Linguistically-driven Multi-task Pre-training for Low-resource Neural Machine Translation

Authors:Zhuoyuan Mao, Chenhui Chu, Sadao Kurohashi

View PDF

Abstract:In the present study, we propose novel sequence-to-sequence pre-training objectives for low-resource machine translation (NMT): Japanese-specific sequence to sequence (JASS) for language pairs involving Japanese as the source or target language, and English-specific sequence to sequence (ENSS) for language pairs involving English. JASS focuses on masking and reordering Japanese linguistic units known as bunsetsu, whereas ENSS is proposed based on phrase structure masking and reordering tasks. Experiments on ASPEC Japanese--English & Japanese--Chinese, Wikipedia Japanese--Chinese, News English--Korean corpora demonstrate that JASS and ENSS outperform MASS and other existing language-agnostic pre-training methods by up to +2.9 BLEU points for the Japanese--English tasks, up to +7.0 BLEU points for the Japanese--Chinese tasks and up to +1.3 BLEU points for English--Korean tasks. Empirical analysis, which focuses on the relationship between individual parts in JASS and ENSS, reveals the complementary nature of the subtasks of JASS and ENSS. Adequacy evaluation using LASER, human evaluation, and case studies reveals that our proposed methods significantly outperform pre-training methods without injected linguistic knowledge and they have a larger positive impact on the adequacy as compared to the fluency. We release codes here: this https URL.

Comments:	An extension of work arXiv:2005.03361
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2201.08070 [cs.CL]
	(or arXiv:2201.08070v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2201.08070
Journal reference:	TALLIP Volume 21, Issue 4, July 2022
Related DOI:	https://doi.org/10.1145/3491065

Submission history

From: Zhuoyuan Mao [view email]
[v1] Thu, 20 Jan 2022 09:10:08 UTC (1,480 KB)

Computer Science > Computation and Language

Title:Linguistically-driven Multi-task Pre-training for Low-resource Neural Machine Translation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Linguistically-driven Multi-task Pre-training for Low-resource Neural Machine Translation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators