Robustifying Models Against Adversarial Attacks by Langevin Dynamics

Srinivasan, Vignesh; Marban, Arturo; Müller, Klaus-Robert; Samek, Wojciech; Nakajima, Shinichi

Computer Science > Machine Learning

arXiv:1805.12017 (cs)

[Submitted on 30 May 2018 (v1), last revised 6 Jun 2019 (this version, v2)]

Title:Robustifying Models Against Adversarial Attacks by Langevin Dynamics

Authors:Vignesh Srinivasan, Arturo Marban, Klaus-Robert Müller, Wojciech Samek, Shinichi Nakajima

View PDF

Abstract:Adversarial attacks on deep learning models have compromised their performance considerably. As remedies, a lot of defense methods were proposed, which however, have been circumvented by newer attacking strategies. In the midst of this ensuing arms race, the problem of robustness against adversarial attacks still remains unsolved. This paper proposes a novel, simple yet effective defense strategy where adversarial samples are relaxed onto the underlying manifold of the (unknown) target class distribution. Specifically, our algorithm drives off-manifold adversarial samples towards high density regions of the data generating distribution of the target class by the Metroplis-adjusted Langevin algorithm (MALA) with perceptual boundary taken into account. Although the motivation is similar to projection methods, e.g., Defense-GAN, our algorithm, called MALA for DEfense (MALADE), is equipped with significant dispersion - projection is distributed broadly, and therefore any whitebox attack cannot accurately align the input so that the MALADE moves it to a targeted untrained spot where the model predicts a wrong label. In our experiments, MALADE exhibited state-of-the-art performance against various elaborate attacking strategies.

Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1805.12017 [cs.LG]
	(or arXiv:1805.12017v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1805.12017

Submission history

From: Vignesh Srinivasan [view email]
[v1] Wed, 30 May 2018 15:01:38 UTC (895 KB)
[v2] Thu, 6 Jun 2019 15:25:03 UTC (8,459 KB)

Computer Science > Machine Learning

Title:Robustifying Models Against Adversarial Attacks by Langevin Dynamics

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Robustifying Models Against Adversarial Attacks by Langevin Dynamics

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators