BilBOWA: Fast Bilingual Distributed Representations without Word Alignments

Gouws, Stephan; Bengio, Yoshua; Corrado, Greg

Statistics > Machine Learning

arXiv:1410.2455 (stat)

[Submitted on 9 Oct 2014 (v1), last revised 4 Feb 2016 (this version, v3)]

Title:BilBOWA: Fast Bilingual Distributed Representations without Word Alignments

Authors:Stephan Gouws, Yoshua Bengio, Greg Corrado

View PDF

Abstract:We introduce BilBOWA (Bilingual Bag-of-Words without Alignments), a simple and computationally-efficient model for learning bilingual distributed representations of words which can scale to large monolingual datasets and does not require word-aligned parallel training data. Instead it trains directly on monolingual data and extracts a bilingual signal from a smaller set of raw-text sentence-aligned data. This is achieved using a novel sampled bag-of-words cross-lingual objective, which is used to regularize two noise-contrastive language models for efficient cross-lingual feature learning. We show that bilingual embeddings learned using the proposed model outperform state-of-the-art methods on a cross-lingual document classification task as well as a lexical translation task on WMT11 data.

Subjects:	Machine Learning (stat.ML); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:1410.2455 [stat.ML]
	(or arXiv:1410.2455v3 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.1410.2455

Submission history

From: Stephan Gouws [view email]
[v1] Thu, 9 Oct 2014 13:41:18 UTC (307 KB)
[v2] Thu, 4 Dec 2014 20:52:32 UTC (242 KB)
[v3] Thu, 4 Feb 2016 05:51:59 UTC (627 KB)

Statistics > Machine Learning

Title:BilBOWA: Fast Bilingual Distributed Representations without Word Alignments

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:BilBOWA: Fast Bilingual Distributed Representations without Word Alignments

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators