Observe and Look Further: Achieving Consistent Performance on Atari

Pohlen, Tobias; Piot, Bilal; Hester, Todd; Azar, Mohammad Gheshlaghi; Horgan, Dan; Budden, David; Barth-Maron, Gabriel; van Hasselt, Hado; Quan, John; Večerík, Mel; Hessel, Matteo; Munos, Rémi; Pietquin, Olivier

Computer Science > Machine Learning

arXiv:1805.11593 (cs)

[Submitted on 29 May 2018]

Title:Observe and Look Further: Achieving Consistent Performance on Atari

Authors:Tobias Pohlen, Bilal Piot, Todd Hester, Mohammad Gheshlaghi Azar, Dan Horgan, David Budden, Gabriel Barth-Maron, Hado van Hasselt, John Quan, Mel Večerík, Matteo Hessel, Rémi Munos, Olivier Pietquin

View PDF

Abstract:Despite significant advances in the field of deep Reinforcement Learning (RL), today's algorithms still fail to learn human-level policies consistently over a set of diverse tasks such as Atari 2600 games. We identify three key challenges that any algorithm needs to master in order to perform well on all games: processing diverse reward distributions, reasoning over long time horizons, and exploring efficiently. In this paper, we propose an algorithm that addresses each of these challenges and is able to learn human-level policies on nearly all Atari games. A new transformed Bellman operator allows our algorithm to process rewards of varying densities and scales; an auxiliary temporal consistency loss allows us to train stably using a discount factor of $\gamma = 0.999$ (instead of $\gamma = 0.99$) extending the effective planning horizon by an order of magnitude; and we ease the exploration problem by using human demonstrations that guide the agent towards rewarding states. When tested on a set of 42 Atari games, our algorithm exceeds the performance of an average human on 40 games using a common set of hyper parameters. Furthermore, it is the first deep RL algorithm to solve the first level of Montezuma's Revenge.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as:	arXiv:1805.11593 [cs.LG]
	(or arXiv:1805.11593v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1805.11593

Submission history

From: Tobias Pohlen [view email]
[v1] Tue, 29 May 2018 17:19:59 UTC (3,227 KB)

Computer Science > Machine Learning

Title:Observe and Look Further: Achieving Consistent Performance on Atari

Submission history

Access Paper:

References & Citations

1 blog link

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Observe and Look Further: Achieving Consistent Performance on Atari

Submission history

Access Paper:

References & Citations

1 blog link

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators