Loss landscapes and optimization in over-parameterized non-linear systems and neural networks

Liu, Chaoyue; Zhu, Libin; Belkin, Mikhail

Computer Science > Machine Learning

arXiv:2003.00307 (cs)

[Submitted on 29 Feb 2020 (v1), last revised 26 May 2021 (this version, v2)]

Title:Loss landscapes and optimization in over-parameterized non-linear systems and neural networks

Authors:Chaoyue Liu, Libin Zhu, Mikhail Belkin

View PDF

Abstract:The success of deep learning is due, to a large extent, to the remarkable effectiveness of gradient-based optimization methods applied to large neural networks. The purpose of this work is to propose a modern view and a general mathematical framework for loss landscapes and efficient optimization in over-parameterized machine learning models and systems of non-linear equations, a setting that includes over-parameterized deep neural networks. Our starting observation is that optimization problems corresponding to such systems are generally not convex, even locally. We argue that instead they satisfy PL$^*$, a variant of the Polyak-Lojasiewicz condition on most (but not all) of the parameter space, which guarantees both the existence of solutions and efficient optimization by (stochastic) gradient descent (SGD/GD). The PL$^*$ condition of these systems is closely related to the condition number of the tangent kernel associated to a non-linear system showing how a PL$^*$-based non-linear theory parallels classical analyses of over-parameterized linear equations. We show that wide neural networks satisfy the PL$^*$ condition, which explains the (S)GD convergence to a global minimum. Finally we propose a relaxation of the PL$^*$ condition applicable to "almost" over-parameterized systems.

Comments:	The discussion on transition to linearity in Version 1 has been moved to arXiv:2010.01092 (appeared in NeurIPS 2020)
Subjects:	Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
Cite as:	arXiv:2003.00307 [cs.LG]
	(or arXiv:2003.00307v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2003.00307

Submission history

From: Chaoyue Liu [view email]
[v1] Sat, 29 Feb 2020 17:18:28 UTC (505 KB)
[v2] Wed, 26 May 2021 19:22:33 UTC (1,662 KB)

Computer Science > Machine Learning

Title:Loss landscapes and optimization in over-parameterized non-linear systems and neural networks

Submission history

Access Paper:

References & Citations

1 blog link

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Loss landscapes and optimization in over-parameterized non-linear systems and neural networks

Submission history

Access Paper:

References & Citations

1 blog link

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators