Training Diagonal Linear Networks with Stochastic Sharpness-Aware Minimization

Clara, Gabriel; Langer, Sophie; Schmidt-Hieber, Johannes

Computer Science > Machine Learning

arXiv:2503.11891 (cs)

[Submitted on 14 Mar 2025]

Title:Training Diagonal Linear Networks with Stochastic Sharpness-Aware Minimization

Authors:Gabriel Clara, Sophie Langer, Johannes Schmidt-Hieber

View PDF HTML (experimental)

Abstract:We analyze the landscape and training dynamics of diagonal linear networks in a linear regression task, with the network parameters being perturbed by small isotropic normal noise. The addition of such noise may be interpreted as a stochastic form of sharpness-aware minimization (SAM) and we prove several results that relate its action on the underlying landscape and training dynamics to the sharpness of the loss. In particular, the noise changes the expected gradient to force balancing of the weight matrices at a fast rate along the descent trajectory. In the diagonal linear model, we show that this equates to minimizing the average sharpness, as well as the trace of the Hessian matrix, among all possible factorizations of the same matrix. Further, the noise forces the gradient descent iterates towards a shrinkage-thresholding of the underlying true parameter, with the noise level explicitly regulating both the shrinkage factor and the threshold.

Comments:	54 pages, 3 figures
Subjects:	Machine Learning (cs.LG); Statistics Theory (math.ST); Machine Learning (stat.ML)
Cite as:	arXiv:2503.11891 [cs.LG]
	(or arXiv:2503.11891v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2503.11891

Submission history

From: Gabriel Clara [view email]
[v1] Fri, 14 Mar 2025 21:45:12 UTC (1,319 KB)

Full-text links:

Access Paper:

view license

Current browse context:

stat.TH

< prev | next >

new | recent | 2025-03

Change to browse by:

cs
cs.LG
math
math.ST
stat
stat.ML

References & Citations

export BibTeX citation

Computer Science > Machine Learning

Title:Training Diagonal Linear Networks with Stochastic Sharpness-Aware Minimization

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Training Diagonal Linear Networks with Stochastic Sharpness-Aware Minimization

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators