Disclosure risk assessment with Bayesian non-parametric hierarchical modelling

Battiston, Marco; Rimella, Lorenzo

Statistics > Applications

arXiv:2408.12521 (stat)

[Submitted on 22 Aug 2024 (v1), last revised 23 Aug 2024 (this version, v2)]

Title:Disclosure risk assessment with Bayesian non-parametric hierarchical modelling

Authors:Marco Battiston, Lorenzo Rimella

View PDF HTML (experimental)

Abstract:Micro and survey datasets often contain private information about individuals, like their health status, income or political preferences. Previous studies have shown that, even after data anonymization, a malicious intruder could still be able to identify individuals in the dataset by matching their variables to external information. Disclosure risk measures are statistical measures meant to quantify how big such a risk is for a specific dataset. One of the most common measures is the number of sample unique values that are also population-unique. \cite{Man12} have shown how mixed membership models can provide very accurate estimates of this measure. A limitation of that approach is that the number of extreme profiles has to be chosen by the modeller. In this article, we propose a non-parametric version of the model, based on the Hierarchical Dirichlet Process (HDP). The proposed approach does not require any tuning parameter or model selection step and provides accurate estimates of the disclosure risk measure, even with samples as small as 1$\%$ of the population size. Moreover, a data augmentation scheme to address the presence of structural zeros is presented. The proposed methodology is tested on a real dataset from the New York census.

Subjects:	Applications (stat.AP); Computation (stat.CO)
Cite as:	arXiv:2408.12521 [stat.AP]
	(or arXiv:2408.12521v2 [stat.AP] for this version)
	https://doi.org/10.48550/arXiv.2408.12521

Submission history

From: Lorenzo Rimella [view email]
[v1] Thu, 22 Aug 2024 16:23:09 UTC (790 KB)
[v2] Fri, 23 Aug 2024 07:21:18 UTC (790 KB)

Statistics > Applications

Title:Disclosure risk assessment with Bayesian non-parametric hierarchical modelling

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Applications

Title:Disclosure risk assessment with Bayesian non-parametric hierarchical modelling

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators