TNT-KID: Transformer-based Neural Tagger for Keyword Identification

Martinc, Matej; Škrlj, Blaž; Pollak, Senja

doi:10.1017/S1351324921000127

Computer Science > Computation and Language

arXiv:2003.09166 (cs)

[Submitted on 20 Mar 2020 (v1), last revised 30 Nov 2021 (this version, v3)]

Title:TNT-KID: Transformer-based Neural Tagger for Keyword Identification

Authors:Matej Martinc, Blaž Škrlj, Senja Pollak

View PDF

Abstract:With growing amounts of available textual data, development of algorithms capable of automatic analysis, categorization and summarization of these data has become a necessity. In this research we present a novel algorithm for keyword identification, i.e., an extraction of one or multi-word phrases representing key aspects of a given document, called Transformer-based Neural Tagger for Keyword IDentification (TNT-KID). By adapting the transformer architecture for a specific task at hand and leveraging language model pretraining on a domain specific corpus, the model is capable of overcoming deficiencies of both supervised and unsupervised state-of-the-art approaches to keyword extraction by offering competitive and robust performance on a variety of different datasets while requiring only a fraction of manually labeled data required by the best performing systems. This study also offers thorough error analysis with valuable insights into the inner workings of the model and an ablation study measuring the influence of specific components of the keyword identification workflow on the overall performance.

Comments:	Accepted to Natural Language Engineering journal
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2003.09166 [cs.CL]
	(or arXiv:2003.09166v3 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2003.09166
Journal reference:	Martinc, M., Škrlj, B., & Pollak, S. (2021). TNT-KID: Transformer-based neural tagger for keyword identification. Natural Language Engineering, 1-40. doi:10.1017/S1351324921000127
Related DOI:	https://doi.org/10.1017/S1351324921000127

Submission history

From: Matej Martinc [view email]
[v1] Fri, 20 Mar 2020 09:55:10 UTC (1,100 KB)
[v2] Tue, 8 Dec 2020 11:45:00 UTC (1,430 KB)
[v3] Tue, 30 Nov 2021 14:56:28 UTC (2,345 KB)

Computer Science > Computation and Language

Title:TNT-KID: Transformer-based Neural Tagger for Keyword Identification

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:TNT-KID: Transformer-based Neural Tagger for Keyword Identification

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators