Dynamics of SGD with Stochastic Polyak Stepsizes: Truly Adaptive Variants and Convergence to Exact Solution

Orvieto, Antonio; Lacoste-Julien, Simon; Loizou, Nicolas

Mathematics > Optimization and Control

arXiv:2205.04583 (math)

[Submitted on 9 May 2022 (v1), last revised 16 Feb 2024 (this version, v5)]

Title:Dynamics of SGD with Stochastic Polyak Stepsizes: Truly Adaptive Variants and Convergence to Exact Solution

Authors:Antonio Orvieto, Simon Lacoste-Julien, Nicolas Loizou

View PDF HTML (experimental)

Abstract:Recently, Loizou et al. (2021), proposed and analyzed stochastic gradient descent (SGD) with stochastic Polyak stepsize (SPS). The proposed SPS comes with strong convergence guarantees and competitive performance; however, it has two main drawbacks when it is used in non-over-parameterized regimes: (i) It requires a priori knowledge of the optimal mini-batch losses, which are not available when the interpolation condition is not satisfied (e.g., regularized objectives), and (ii) it guarantees convergence only to a neighborhood of the solution. In this work, we study the dynamics and the convergence properties of SGD equipped with new variants of the stochastic Polyak stepsize and provide solutions to both drawbacks of the original SPS. We first show that a simple modification of the original SPS that uses lower bounds instead of the optimal function values can directly solve issue (i). On the other hand, solving issue (ii) turns out to be more challenging and leads us to valuable insights into the method's behavior. We show that if interpolation is not satisfied, the correlation between SPS and stochastic gradients introduces a bias, which effectively distorts the expectation of the gradient signal near minimizers, leading to non-convergence - even if the stepsize is scaled down during training. To fix this issue, we propose DecSPS, a novel modification of SPS, which guarantees convergence to the exact minimizer - without a priori knowledge of the problem parameters. For strongly-convex optimization problems, DecSPS is the first stochastic adaptive optimization method that converges to the exact solution without restrictive assumptions like bounded iterates/gradients.

Comments:	Accepted at NeurIPS 2022 v4: tiny mistake in the main proof (result unchanged) is now fixed, v5: confusing typo fixed
Subjects:	Optimization and Control (math.OC)
Cite as:	arXiv:2205.04583 [math.OC]
	(or arXiv:2205.04583v5 [math.OC] for this version)
	https://doi.org/10.48550/arXiv.2205.04583

Submission history

From: Antonio Orvieto [view email]
[v1] Mon, 9 May 2022 22:15:52 UTC (4,165 KB)
[v2] Sun, 3 Jul 2022 22:07:01 UTC (4,715 KB)
[v3] Wed, 28 Dec 2022 11:15:35 UTC (5,021 KB)
[v4] Wed, 14 Feb 2024 11:53:32 UTC (5,021 KB)
[v5] Fri, 16 Feb 2024 23:16:13 UTC (5,021 KB)

Mathematics > Optimization and Control

Title:Dynamics of SGD with Stochastic Polyak Stepsizes: Truly Adaptive Variants and Convergence to Exact Solution

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Mathematics > Optimization and Control

Title:Dynamics of SGD with Stochastic Polyak Stepsizes: Truly Adaptive Variants and Convergence to Exact Solution

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators