QuaRL: Quantization for Fast and Environmentally Sustainable Reinforcement Learning

Krishnan, Srivatsan; Lam, Maximilian; Chitlangia, Sharad; Wan, Zishen; Barth-Maron, Gabriel; Faust, Aleksandra; Reddi, Vijay Janapa

Computer Science > Machine Learning

arXiv:1910.01055 (cs)

[Submitted on 2 Oct 2019 (v1), last revised 14 Nov 2022 (this version, v6)]

Title:QuaRL: Quantization for Fast and Environmentally Sustainable Reinforcement Learning

Authors:Srivatsan Krishnan, Maximilian Lam, Sharad Chitlangia, Zishen Wan, Gabriel Barth-Maron, Aleksandra Faust, Vijay Janapa Reddi

View PDF

Abstract:Deep reinforcement learning continues to show tremendous potential in achieving task-level autonomy, however, its computational and energy demands remain prohibitively high. In this paper, we tackle this problem by applying quantization to reinforcement learning. To that end, we introduce a novel Reinforcement Learning (RL) training paradigm, \textit{ActorQ}, to speed up actor-learner distributed RL training. \textit{ActorQ} leverages 8-bit quantized actors to speed up data collection without affecting learning convergence. Our quantized distributed RL training system, \textit{ActorQ}, demonstrates end-to-end speedups \blue{between 1.5 $\times$ and 5.41$\times$}, and faster convergence over full precision training on a range of tasks (Deepmind Control Suite) and different RL algorithms (D4PG, DQN). Furthermore, we compare the carbon emissions (Kgs of CO2) of \textit{ActorQ} versus standard reinforcement learning \blue{algorithms} on various tasks. Across various settings, we show that \textit{ActorQ} enables more environmentally friendly reinforcement learning by achieving \blue{carbon emission improvements between 1.9$\times$ and 3.76$\times$} compared to training RL-agents in full-precision. We believe that this is the first of many future works on enabling computationally energy-efficient and sustainable reinforcement learning. The source code is available here for the public to use: \url{this https URL}.

Comments:	Equal contribution from first three authors. Updating with QuaRL for sustainable (carbon emissions) RL results
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
Cite as:	arXiv:1910.01055 [cs.LG]
	(or arXiv:1910.01055v6 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1910.01055
Journal reference:	Published in Transactions on Machine Learning Research (07/2022)

Submission history

From: Srivatsan Krishnan [view email]
[v1] Wed, 2 Oct 2019 16:22:29 UTC (1,893 KB)
[v2] Fri, 4 Oct 2019 21:44:52 UTC (1,893 KB)
[v3] Sat, 7 Dec 2019 00:57:49 UTC (1,952 KB)
[v4] Mon, 18 Jan 2021 20:05:26 UTC (13,577 KB)
[v5] Sun, 28 Nov 2021 03:15:24 UTC (684 KB)
[v6] Mon, 14 Nov 2022 01:42:10 UTC (3,029 KB)

Computer Science > Machine Learning

Title:QuaRL: Quantization for Fast and Environmentally Sustainable Reinforcement Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:QuaRL: Quantization for Fast and Environmentally Sustainable Reinforcement Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators