Subspace Clustering with Missing and Corrupted Data

Charles, Zachary; Jalali, Amin; Willett, Rebecca

Statistics > Machine Learning

arXiv:1707.02461 (stat)

[Submitted on 8 Jul 2017 (v1), last revised 15 Jan 2018 (this version, v2)]

Title:Subspace Clustering with Missing and Corrupted Data

Authors:Zachary Charles, Amin Jalali, Rebecca Willett

View PDF

Abstract:Given full or partial information about a collection of points that lie close to a union of several subspaces, subspace clustering refers to the process of clustering the points according to their subspace and identifying the subspaces. One popular approach, sparse subspace clustering (SSC), represents each sample as a weighted combination of the other samples, with weights of minimal $\ell_1$ norm, and then uses those learned weights to cluster the samples. SSC is stable in settings where each sample is contaminated by a relatively small amount of noise. However, when there is a significant amount of additive noise, or a considerable number of entries are missing, theoretical guarantees are scarce. In this paper, we study a robust variant of SSC and establish clustering guarantees in the presence of corrupted or missing data. We give explicit bounds on amount of noise and missing data that the algorithm can tolerate, both in deterministic settings and in a random generative model. Notably, our approach provides guarantees for higher tolerance to noise and missing data than existing analyses for this method. By design, the results hold even when we do not know the locations of the missing data; e.g., as in presence-only data.

Comments:	31 pages, 2 figures
Subjects:	Machine Learning (stat.ML)
Cite as:	arXiv:1707.02461 [stat.ML]
	(or arXiv:1707.02461v2 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.1707.02461

Submission history

From: Zachary Charles [view email]
[v1] Sat, 8 Jul 2017 16:24:11 UTC (283 KB)
[v2] Mon, 15 Jan 2018 18:26:24 UTC (90 KB)

Statistics > Machine Learning

Title:Subspace Clustering with Missing and Corrupted Data

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:Subspace Clustering with Missing and Corrupted Data

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators