On the privacy-utility trade-off in differentially private hierarchical text classification

Wunderlich, Dominik; Bernau, Daniel; Aldà, Francesco; Parra-Arnau, Javier; Strufe, Thorsten

Computer Science > Cryptography and Security

arXiv:2103.02895 (cs)

[Submitted on 4 Mar 2021 (v1), last revised 9 Dec 2021 (this version, v2)]

Title:On the privacy-utility trade-off in differentially private hierarchical text classification

Authors:Dominik Wunderlich, Daniel Bernau, Francesco Aldà, Javier Parra-Arnau, Thorsten Strufe

View PDF

Abstract:Hierarchical text classification consists in classifying text documents into a hierarchy of classes and sub-classes. Although artificial neural networks have proved useful to perform this task, unfortunately they can leak training data information to adversaries due to training data memorization. Using differential privacy during model training can mitigate leakage attacks against trained models, enabling the models to be shared safely at the cost of reduced model accuracy. This work investigates the privacy-utility trade-off in hierarchical text classification with differential privacy guarantees, and identifies neural network architectures that offer superior trade-offs. To this end, we use a white-box membership inference attack to empirically assess the information leakage of three widely used neural network architectures. We show that large differential privacy parameters already suffice to completely mitigate membership inference attacks, thus resulting only in a moderate decrease in model utility. More specifically, for large datasets with long texts we observed Transformer-based models to achieve an overall favorable privacy-utility trade-off, while for smaller datasets with shorter texts convolutional neural networks are preferable.

Subjects:	Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2103.02895 [cs.CR]
	(or arXiv:2103.02895v2 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.2103.02895

Submission history

From: Daniel Bernau [view email]
[v1] Thu, 4 Mar 2021 08:51:00 UTC (397 KB)
[v2] Thu, 9 Dec 2021 10:11:11 UTC (533 KB)

Computer Science > Cryptography and Security

Title:On the privacy-utility trade-off in differentially private hierarchical text classification

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:On the privacy-utility trade-off in differentially private hierarchical text classification

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators