Expose Backdoors on the Way: A Feature-Based Efficient Defense against Textual Backdoor Attacks

Chen, Sishuo; Yang, Wenkai; Zhang, Zhiyuan; Bi, Xiaohan; Sun, Xu

Computer Science > Computation and Language

arXiv:2210.07907 (cs)

[Submitted on 14 Oct 2022]

Title:Expose Backdoors on the Way: A Feature-Based Efficient Defense against Textual Backdoor Attacks

Authors:Sishuo Chen, Wenkai Yang, Zhiyuan Zhang, Xiaohan Bi, Xu Sun

View PDF

Abstract:Natural language processing (NLP) models are known to be vulnerable to backdoor attacks, which poses a newly arisen threat to NLP models. Prior online backdoor defense methods for NLP models only focus on the anomalies at either the input or output level, still suffering from fragility to adaptive attacks and high computational cost. In this work, we take the first step to investigate the unconcealment of textual poisoned samples at the intermediate-feature level and propose a feature-based efficient online defense method. Through extensive experiments on existing attacking methods, we find that the poisoned samples are far away from clean samples in the intermediate feature space of a poisoned NLP model. Motivated by this observation, we devise a distance-based anomaly score (DAN) to distinguish poisoned samples from clean samples at the feature level. Experiments on sentiment analysis and offense detection tasks demonstrate the superiority of DAN, as it substantially surpasses existing online defense methods in terms of defending performance and enjoys lower inference costs. Moreover, we show that DAN is also resistant to adaptive attacks based on feature-level regularization. Our code is available at this https URL.

Comments:	Findings of EMNLP 2022
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2210.07907 [cs.CL]
	(or arXiv:2210.07907v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2210.07907

Submission history

From: Sishuo Chen [view email]
[v1] Fri, 14 Oct 2022 15:44:28 UTC (719 KB)

Computer Science > Computation and Language

Title:Expose Backdoors on the Way: A Feature-Based Efficient Defense against Textual Backdoor Attacks

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Expose Backdoors on the Way: A Feature-Based Efficient Defense against Textual Backdoor Attacks

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators