If You've Seen One, You've Seen Them All: Leveraging AST Clustering Using MCL to Mimic Expertise to Detect Software Supply Chain Attacks

Ohm, Marc; Kempf, Lukas; Boes, Felix; Meier, Michael

Computer Science > Cryptography and Security

arXiv:2011.02235v1 (cs)

[Submitted on 4 Nov 2020 (this version), latest version 19 Mar 2021 (v2)]

Title:If You've Seen One, You've Seen Them All: Leveraging AST Clustering Using MCL to Mimic Expertise to Detect Software Supply Chain Attacks

Authors:Marc Ohm, Lukas Kempf, Felix Boes, Michael Meier

View PDF

Abstract:Trojanized software packages used in software supply chain attacks constitute an merging threat. Unfortunately, there is still a lack of scalable approaches that allow automated and timely detection of malicious software packages. However, it has been observed that most attack campaigns comprise multiple packages that share the same or similar malicious code. We leverage that fact to automatically reproduce manually identified clusters of known malicious packages that have been used in real world attacks, thus, reducing the need for expert knowledge and manual inspection. Our approach, AST Clustering using MCL to mimic Expertise (ACME), yields promising results with a $F_1$ score of 0.99. Signatures are automatically generated based on representative code fragments from clusters and are subsequently used to scan the whole npm registry for unreported malicious packages. We are able to identify and report six malicious packages that have been removed from npm consequentially. Therefore, our approach is able to reproduce clustering based on expert knowledge and hence may be employed by maintainers of package repositories like npm to timely detect possible maliciousness of newly uploaded or updated packages.

Comments:	submitted to IEEE EuroS&P 2021
Subjects:	Cryptography and Security (cs.CR)
Cite as:	arXiv:2011.02235 [cs.CR]
	(or arXiv:2011.02235v1 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.2011.02235

Submission history

From: Marc Ohm [view email]
[v1] Wed, 4 Nov 2020 11:26:07 UTC (356 KB)
[v2] Fri, 19 Mar 2021 10:30:24 UTC (329 KB)

Computer Science > Cryptography and Security

Title:If You've Seen One, You've Seen Them All: Leveraging AST Clustering Using MCL to Mimic Expertise to Detect Software Supply Chain Attacks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:If You've Seen One, You've Seen Them All: Leveraging AST Clustering Using MCL to Mimic Expertise to Detect Software Supply Chain Attacks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators