Statistical learning on measures: an application to persistence diagrams

Hacquard, Olympio; Blanchard, Gilles; Levrard, Clément

Computer Science > Computational Geometry

arXiv:2303.08456 (cs)

[Submitted on 15 Mar 2023 (v1), last revised 31 May 2023 (this version, v2)]

Title:Statistical learning on measures: an application to persistence diagrams

Authors:Olympio Hacquard (LMO, DATASHAPE), Gilles Blanchard (LMO, DATASHAPE), Clément Levrard (LPSM)

View PDF

Abstract:We consider a binary supervised learning classification problem where instead of having data in a finite-dimensional Euclidean space, we observe measures on a compact space $\mathcal{X}$. Formally, we observe data $D_N = (\mu_1, Y_1), \ldots, (\mu_N, Y_N)$ where $\mu_i$ is a measure on $\mathcal{X}$ and $Y_i$ is a label in $\{0, 1\}$. Given a set $\mathcal{F}$ of base-classifiers on $\mathcal{X}$, we build corresponding classifiers in the space of measures. We provide upper and lower bounds on the Rademacher complexity of this new class of classifiers that can be expressed simply in terms of corresponding quantities for the class $\mathcal{F}$. If the measures $\mu_i$ are uniform over a finite set, this classification task boils down to a multi-instance learning problem. However, our approach allows more flexibility and diversity in the input data we can deal with. While such a framework has many possible applications, this work strongly emphasizes on classifying data via topological descriptors called persistence diagrams. These objects are discrete measures on $\mathbb{R}^2$, where the coordinates of each point correspond to the range of scales at which a topological feature exists. We will present several classifiers on measures and show how they can heuristically and theoretically enable a good classification performance in various settings in the case of persistence diagrams.

Subjects:	Computational Geometry (cs.CG); Statistics Theory (math.ST); Machine Learning (stat.ML)
Cite as:	arXiv:2303.08456 [cs.CG]
	(or arXiv:2303.08456v2 [cs.CG] for this version)
	https://doi.org/10.48550/arXiv.2303.08456

Submission history

From: Olympio Hacquard [view email] [via CCSD proxy]
[v1] Wed, 15 Mar 2023 09:01:37 UTC (7,831 KB)
[v2] Wed, 31 May 2023 08:09:26 UTC (1,134 KB)

Computer Science > Computational Geometry

Title:Statistical learning on measures: an application to persistence diagrams

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computational Geometry

Title:Statistical learning on measures: an application to persistence diagrams

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators