Untangling the Influence of Typology, Data and Model Architecture on Ranking Transfer Languages for Cross-Lingual POS Tagging

Rice, Enora; Marashian, Ali; Haynie, Hannah; von der Wense, Katharina; Palmer, Alexis

Computer Science > Computation and Language

arXiv:2503.19979 (cs)

[Submitted on 25 Mar 2025]

Title:Untangling the Influence of Typology, Data and Model Architecture on Ranking Transfer Languages for Cross-Lingual POS Tagging

Authors:Enora Rice, Ali Marashian, Hannah Haynie, Katharina von der Wense, Alexis Palmer

View PDF HTML (experimental)

Abstract:Cross-lingual transfer learning is an invaluable tool for overcoming data scarcity, yet selecting a suitable transfer language remains a challenge. The precise roles of linguistic typology, training data, and model architecture in transfer language choice are not fully understood. We take a holistic approach, examining how both dataset-specific and fine-grained typological features influence transfer language selection for part-of-speech tagging, considering two different sources for morphosyntactic features. While previous work examines these dynamics in the context of bilingual biLSTMS, we extend our analysis to a more modern transfer learning pipeline: zero-shot prediction with pretrained multilingual models. We train a series of transfer language ranking systems and examine how different feature inputs influence ranker performance across architectures. Word overlap, type-token ratio, and genealogical distance emerge as top features across all architectures. Our findings reveal that a combination of typological and dataset-dependent features leads to the best rankings, and that good performance can be obtained with either feature group on its own.

Comments:	Accepted to NAACL 2025 Workshop Language Models for Underserved Communities
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2503.19979 [cs.CL]
	(or arXiv:2503.19979v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2503.19979

Submission history

From: Enora Rice [view email]
[v1] Tue, 25 Mar 2025 18:05:40 UTC (55 KB)

Computer Science > Computation and Language

Title:Untangling the Influence of Typology, Data and Model Architecture on Ranking Transfer Languages for Cross-Lingual POS Tagging

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Untangling the Influence of Typology, Data and Model Architecture on Ranking Transfer Languages for Cross-Lingual POS Tagging

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators