How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study

Erben, Alexander; Mayer, Ruben; Jacobsen, Hans-Arno

Computer Science > Machine Learning

arXiv:2306.03163 (cs)

[Submitted on 5 Jun 2023 (v1), last revised 2 Jun 2024 (this version, v4)]

Title:How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study

Authors:Alexander Erben, Ruben Mayer, Hans-Arno Jacobsen

View PDF HTML (experimental)

Abstract:This paper aims to answer the question: Can deep learning models be cost-efficiently trained on a global market of spot VMs spanning different data centers and cloud providers? To provide guidance, we extensively evaluate the cost and throughput implications of training in different zones, continents, and clouds for representative CV, NLP, and ASR models. To expand the current training options further, we compare the scalability potential for hybrid-cloud scenarios by adding cloud resources to on-premise hardware to improve training throughput. Finally, we show how leveraging spot instance pricing enables a new cost-efficient way to train models with multiple cheap VMs, trumping both more centralized and powerful hardware and even on-demand cloud offerings at competitive prices.

Comments:	Published at VLDB 2024. Artifacts and Code: this https URL
Subjects:	Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC); Networking and Internet Architecture (cs.NI); Performance (cs.PF)
ACM classes:	I.2.11; C.2.4; C.4; D.2.8
Cite as:	arXiv:2306.03163 [cs.LG]
	(or arXiv:2306.03163v4 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2306.03163

Submission history

From: Alexander Erben [view email]
[v1] Mon, 5 Jun 2023 18:17:37 UTC (909 KB)
[v2] Tue, 21 Nov 2023 14:10:03 UTC (924 KB)
[v3] Sat, 3 Feb 2024 23:19:08 UTC (925 KB)
[v4] Sun, 2 Jun 2024 09:53:59 UTC (1,705 KB)

Computer Science > Machine Learning

Title:How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators