Sketching Datasets for Large-Scale Learning (long version)

Gribonval, Rémi; Chatalic, Antoine; Keriven, Nicolas; Schellekens, Vincent; Jacques, Laurent; Schniter, Philip

Statistics > Machine Learning

arXiv:2008.01839 (stat)

[Submitted on 4 Aug 2020 (v1), last revised 24 Jun 2021 (this version, v3)]

Title:Sketching Datasets for Large-Scale Learning (long version)

Authors:Rémi Gribonval, Antoine Chatalic, Nicolas Keriven, Vincent Schellekens, Laurent Jacques, Philip Schniter

View PDF

Abstract:This article considers "compressive learning," an approach to large-scale machine learning where datasets are massively compressed before learning (e.g., clustering, classification, or regression) is performed. In particular, a "sketch" is first constructed by computing carefully chosen nonlinear random features (e.g., random Fourier features) and averaging them over the whole dataset. Parameters are then learned from the sketch, without access to the original dataset. This article surveys the current state-of-the-art in compressive learning, including the main concepts and algorithms, their connections with established signal-processing methods, existing theoretical guarantees -- on both information preservation and privacy preservation, and important open problems.

Subjects:	Machine Learning (stat.ML); Information Theory (cs.IT); Machine Learning (cs.LG)
Cite as:	arXiv:2008.01839 [stat.ML]
	(or arXiv:2008.01839v3 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.2008.01839

Submission history

From: Philip Schniter [view email]
[v1] Tue, 4 Aug 2020 21:29:05 UTC (17,703 KB)
[v2] Tue, 19 Jan 2021 20:41:41 UTC (20,095 KB)
[v3] Thu, 24 Jun 2021 21:36:36 UTC (19,141 KB)

Full-text links:

Access Paper:

view license

Current browse context:

stat.ML

< prev | next >

new | recent | 2020-08

Change to browse by:

cs
cs.IT
cs.LG
math
math.IT
stat

References & Citations

export BibTeX citation

Statistics > Machine Learning

Title:Sketching Datasets for Large-Scale Learning (long version)

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:Sketching Datasets for Large-Scale Learning (long version)

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators