CoAM: Corpus of All-Type Multiword Expressions

Ide, Yusuke; Tanner, Joshua; Nohejl, Adam; Hoffman, Jacob; Vasselli, Justin; Kamigaito, Hidetaka; Watanabe, Taro

Computer Science > Computation and Language

arXiv:2412.18151 (cs)

[Submitted on 24 Dec 2024]

Title:CoAM: Corpus of All-Type Multiword Expressions

Authors:Yusuke Ide, Joshua Tanner, Adam Nohejl, Jacob Hoffman, Justin Vasselli, Hidetaka Kamigaito, Taro Watanabe

View PDF HTML (experimental)

Abstract:Multiword expressions (MWEs) refer to idiomatic sequences of multiple words. MWE identification, i.e., detecting MWEs in text, can play a key role in downstream tasks such as machine translation. Existing datasets for MWE identification are inconsistently annotated, limited to a single type of MWE, or limited in size. To enable reliable and comprehensive evaluation, we created CoAM: Corpus of All-Type Multiword Expressions, a dataset of 1.3K sentences constructed through a multi-step process to enhance data quality consisting of human annotation, human review, and automated consistency checking. MWEs in CoAM are tagged with MWE types, such as Noun and Verb, to enable fine-grained error analysis. Annotations for CoAM were collected using a new interface created with our interface generator, which allows easy and flexible annotation of MWEs in any form, including discontinuous ones. Through experiments using CoAM, we find that a fine-tuned large language model outperforms the current state-of-the-art approach for MWE identification. Furthermore, analysis using our MWE type tagged data reveals that Verb MWEs are easier than Noun MWEs to identify across approaches.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2412.18151 [cs.CL]
	(or arXiv:2412.18151v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2412.18151

Submission history

From: Yusuke Ide [view email]
[v1] Tue, 24 Dec 2024 04:09:33 UTC (9,123 KB)

Computer Science > Computation and Language

Title:CoAM: Corpus of All-Type Multiword Expressions

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:CoAM: Corpus of All-Type Multiword Expressions

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators