Cross-validation on Extreme Regions

Aghbalou, Anass; Bertail, Patrice; Portier, François; Sabourin, Anne

Mathematics > Statistics Theory

arXiv:2202.00488 (math)

[Submitted on 1 Feb 2022 (v1), last revised 10 Sep 2024 (this version, v3)]

Title:Cross-validation on Extreme Regions

Authors:Anass Aghbalou, Patrice Bertail, François Portier, Anne Sabourin

View PDF

Abstract:We conduct a non asymptotic study of the Cross Validation (CV) estimate of the generalization risk for learning algorithms dedicated to extreme regions of the covariates space. In this Extreme Value Analysis context, the risk function measures the algorithm's error given that the norm of the input exceeds a high quantile. The main challenge within this framework is the negligible size of the extreme training sample with respect to the full sample size and the necessity to re-scale the risk function by a probability tending to zero. We open the road to a finite sample understanding of CV for extreme values by establishing two new results: an exponential probability bound on the \Kfold CV error and a polynomial probability bound on the leave-\textrm{p}-out CV. Our bounds are sharp in the sense that they match state-of-the-art guarantees for standard CV estimates while extending them to encompass a conditioning event of small probability. We illustrate the significance of our results regarding high dimensional classification in extreme regions via a Lasso-type logistic regression algorithm. The tightness of our bounds is investigated in numerical experiments.

Comments:	53 pages, 2 figures. To appear in Extremes
Subjects:	Statistics Theory (math.ST)
Cite as:	arXiv:2202.00488 [math.ST]
	(or arXiv:2202.00488v3 [math.ST] for this version)
	https://doi.org/10.48550/arXiv.2202.00488

Submission history

From: Anne Sabourin [view email]
[v1] Tue, 1 Feb 2022 15:40:53 UTC (313 KB)
[v2] Sun, 11 Jun 2023 15:11:31 UTC (1,319 KB)
[v3] Tue, 10 Sep 2024 20:41:57 UTC (485 KB)

Mathematics > Statistics Theory

Title:Cross-validation on Extreme Regions

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Mathematics > Statistics Theory

Title:Cross-validation on Extreme Regions

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators