(S)GD over Diagonal Linear Networks: Implicit Regularisation, Large Stepsizes and Edge of Stability

Even, Mathieu; Pesme, Scott; Gunasekar, Suriya; Flammarion, Nicolas

Computer Science > Machine Learning

arXiv:2302.08982 (cs)

[Submitted on 17 Feb 2023 (v1), last revised 25 Oct 2023 (this version, v2)]

Title:(S)GD over Diagonal Linear Networks: Implicit Regularisation, Large Stepsizes and Edge of Stability

Authors:Mathieu Even, Scott Pesme, Suriya Gunasekar, Nicolas Flammarion

View PDF

Abstract:In this paper, we investigate the impact of stochasticity and large stepsizes on the implicit regularisation of gradient descent (GD) and stochastic gradient descent (SGD) over diagonal linear networks. We prove the convergence of GD and SGD with macroscopic stepsizes in an overparametrised regression setting and characterise their solutions through an implicit regularisation problem. Our crisp characterisation leads to qualitative insights about the impact of stochasticity and stepsizes on the recovered solution. Specifically, we show that large stepsizes consistently benefit SGD for sparse regression problems, while they can hinder the recovery of sparse solutions for GD. These effects are magnified for stepsizes in a tight window just below the divergence threshold, in the "edge of stability" regime. Our findings are supported by experimental results.

Subjects:	Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
Cite as:	arXiv:2302.08982 [cs.LG]
	(or arXiv:2302.08982v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2302.08982

Submission history

From: Scott Pesme [view email]
[v1] Fri, 17 Feb 2023 16:37:08 UTC (5,244 KB)
[v2] Wed, 25 Oct 2023 16:09:19 UTC (8,954 KB)

Computer Science > Machine Learning

Title:(S)GD over Diagonal Linear Networks: Implicit Regularisation, Large Stepsizes and Edge of Stability

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:(S)GD over Diagonal Linear Networks: Implicit Regularisation, Large Stepsizes and Edge of Stability

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators