A Framework for Overparameterized Learning

Terjék, Dávid; González-Sánchez, Diego

Computer Science > Machine Learning

arXiv:2205.13507 (cs)

[Submitted on 26 May 2022 (v1), last revised 13 Feb 2023 (this version, v2)]

Title:A Framework for Overparameterized Learning

Authors:Dávid Terjék, Diego González-Sánchez

View PDF

Abstract:A candidate explanation of the good empirical performance of deep neural networks is the implicit regularization effect of first order optimization methods. Inspired by this, we prove a convergence theorem for nonconvex composite optimization, and apply it to a general learning problem covering many machine learning applications, including supervised learning. We then present a deep multilayer perceptron model and prove that, when sufficiently wide, it $(i)$ leads to the convergence of gradient descent to a global optimum with a linear rate, $(ii)$ benefits from the implicit regularization effect of gradient descent, $(iii)$ is subject to novel bounds on the generalization error, $(iv)$ exhibits the lazy training phenomenon and $(v)$ enjoys learning rate transfer across different widths. The corresponding coefficients, such as the convergence rate, improve as width is further increased, and depend on the even order moments of the data generating distribution up to an order depending on the number of layers. The only non-mild assumption we make is the concentration of the smallest eigenvalue of the neural tangent kernel at initialization away from zero, which has been shown to hold for a number of less general models in contemporary works. We present empirical evidence supporting this assumption as well as our theoretical claims.

Comments:	31 pages, 5 figures
Subjects:	Machine Learning (cs.LG); Functional Analysis (math.FA); Optimization and Control (math.OC); Machine Learning (stat.ML)
MSC classes:	68T07 (Primary) 46N10, 90C26 (Secondary)
Cite as:	arXiv:2205.13507 [cs.LG]
	(or arXiv:2205.13507v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2205.13507

Submission history

From: Dávid Terjék [view email]
[v1] Thu, 26 May 2022 17:17:46 UTC (42 KB)
[v2] Mon, 13 Feb 2023 16:32:28 UTC (349 KB)

Computer Science > Machine Learning

Title:A Framework for Overparameterized Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:A Framework for Overparameterized Learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators