Understanding Pure Character-Based Neural Machine Translation: The Case of Translating Finnish into English

Tang, Gongbo; Sennrich, Rico; Nivre, Joakim

Computer Science > Computation and Language

arXiv:2011.03469 (cs)

[Submitted on 6 Nov 2020]

Title:Understanding Pure Character-Based Neural Machine Translation: The Case of Translating Finnish into English

Authors:Gongbo Tang, Rico Sennrich, Joakim Nivre

View PDF

Abstract:Recent work has shown that deeper character-based neural machine translation (NMT) models can outperform subword-based models. However, it is still unclear what makes deeper character-based models successful. In this paper, we conduct an investigation into pure character-based models in the case of translating Finnish into English, including exploring the ability to learn word senses and morphological inflections and the attention mechanism. We demonstrate that word-level information is distributed over the entire character sequence rather than over a single character, and characters at different positions play different roles in learning linguistic knowledge. In addition, character-based models need more layers to encode word senses which explains why only deeper models outperform subword-based models. The attention distribution pattern shows that separators attract a lot of attention and we explore a sparse word-level attention to enforce character hidden states to capture the full word-level information. Experimental results show that the word-level attention with a single head results in 1.2 BLEU points drop.

Comments:	accepted by COLING 2020, camera-ready version
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2011.03469 [cs.CL]
	(or arXiv:2011.03469v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2011.03469

Submission history

From: Gongbo Tang [view email]
[v1] Fri, 6 Nov 2020 16:47:43 UTC (270 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2020-11

Change to browse by:

References & Citations

DBLP - CS Bibliography

listing | bibtex

Gongbo Tang
Rico Sennrich
Joakim Nivre

export BibTeX citation

Computer Science > Computation and Language

Title:Understanding Pure Character-Based Neural Machine Translation: The Case of Translating Finnish into English

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Understanding Pure Character-Based Neural Machine Translation: The Case of Translating Finnish into English

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators