Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities

Dutta, Senjuti; Mittal, Sid; Chen, Sherol; Ramachandran, Deepak; Rajakumar, Ravi; Kivlichan, Ian; Mak, Sunny; Butryna, Alena; Paritosh, Praveen

Computer Science > Artificial Intelligence

arXiv:2311.00203 (cs)

[Submitted on 1 Nov 2023]

Title:Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities

Authors:Senjuti Dutta (1), Sid Mittal (2), Sherol Chen (2), Deepak Ramachandran (2), Ravi Rajakumar (2), Ian Kivlichan (2), Sunny Mak (2), Alena Butryna (2), Praveen Paritosh (2) ((1) University of Tennessee, Knoxville, (2) Google LLC)

View PDF

Abstract:The prevalence and impact of toxic discussions online have made content moderation crucial.Automated systems can play a vital role in identifying toxicity, and reducing the reliance on human moderation.Nevertheless, identifying toxic comments for diverse communities continues to present challenges that are addressed in this paper.The two-part goal of this study is to(1)identify intuitive variances from annotator disagreement using quantitative analysis and (2)model the subjectivity of these this http URL achieve our goal, we published a new dataset\footnote{\url{this https URL}} with expert annotators' annotations and used two other public datasets to identify the subjectivity of toxicity.Then leveraging the Large Language Model(LLM),we evaluate the model's ability to mimic diverse viewpoints on toxicity by varying size of the training data and utilizing same set of annotators as the test set used during model training and a separate set of annotators as the test set.We conclude that subjectivity is evident across all annotator groups, demonstrating the shortcomings of majority-rule voting. Moving forward, subjective annotations should serve as ground truth labels for training models for domains like toxicity in diverse communities.

Subjects:	Artificial Intelligence (cs.AI)
Cite as:	arXiv:2311.00203 [cs.AI]
	(or arXiv:2311.00203v1 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2311.00203

Submission history

From: Senjuti Dutta [view email]
[v1] Wed, 1 Nov 2023 00:17:11 UTC (1,987 KB)

Computer Science > Artificial Intelligence

Title:Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators