RMDM: A Multilabel Fakenews Dataset for Vietnamese Evidence Verification

Nguyen, Hai-Long; Pham, Thi-Kieu-Trang; Le, Thai-Son; Nguyen, Tan-Minh; Vuong, Thi-Hai-Yen; Nguyen, Ha-Thanh

Computer Science > Computation and Language

arXiv:2309.09071 (cs)

[Submitted on 16 Sep 2023]

Title:RMDM: A Multilabel Fakenews Dataset for Vietnamese Evidence Verification

Authors:Hai-Long Nguyen, Thi-Kieu-Trang Pham, Thai-Son Le, Tan-Minh Nguyen, Thi-Hai-Yen Vuong, Ha-Thanh Nguyen

View PDF

Abstract:In this study, we present a novel and challenging multilabel Vietnamese dataset (RMDM) designed to assess the performance of large language models (LLMs), in verifying electronic information related to legal contexts, focusing on fake news as potential input for electronic evidence. The RMDM dataset comprises four labels: real, mis, dis, and mal, representing real information, misinformation, disinformation, and mal-information, respectively. By including these diverse labels, RMDM captures the complexities of differing fake news categories and offers insights into the abilities of different language models to handle various types of information that could be part of electronic evidence. The dataset consists of a total of 1,556 samples, with 389 samples for each label. Preliminary tests on the dataset using GPT-based and BERT-based models reveal variations in the models' performance across different labels, indicating that the dataset effectively challenges the ability of various language models to verify the authenticity of such information. Our findings suggest that verifying electronic information related to legal contexts, including fake news, remains a difficult problem for language models, warranting further attention from the research community to advance toward more reliable AI models for potential legal applications.

Comments:	ISAILD@KSE 2023
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2309.09071 [cs.CL]
	(or arXiv:2309.09071v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2309.09071

Submission history

From: Ha Thanh Nguyen [view email]
[v1] Sat, 16 Sep 2023 18:35:08 UTC (724 KB)

Computer Science > Computation and Language

Title:RMDM: A Multilabel Fakenews Dataset for Vietnamese Evidence Verification

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:RMDM: A Multilabel Fakenews Dataset for Vietnamese Evidence Verification

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators