An Information-Theoretic Optimality Principle for Deep Reinforcement Learning

Leibfried, Felix; Grau-Moya, Jordi; Bou-Ammar, Haitham

Computer Science > Artificial Intelligence

arXiv:1708.01867 (cs)

[Submitted on 6 Aug 2017 (v1), last revised 20 Nov 2018 (this version, v5)]

Title:An Information-Theoretic Optimality Principle for Deep Reinforcement Learning

Authors:Felix Leibfried, Jordi Grau-Moya, Haitham Bou-Ammar

View PDF

Abstract:We methodologically address the problem of Q-value overestimation in deep reinforcement learning to handle high-dimensional state spaces efficiently. By adapting concepts from information theory, we introduce an intrinsic penalty signal encouraging reduced Q-value estimates. The resultant algorithm encompasses a wide range of learning outcomes containing deep Q-networks as a special case. Different learning outcomes can be demonstrated by tuning a Lagrange multiplier accordingly. We furthermore propose a novel scheduling scheme for this Lagrange multiplier to ensure efficient and robust learning. In experiments on Atari, our algorithm outperforms other algorithms (e.g. deep and double deep Q-networks) in terms of both game-play performance and sample complexity. These results remain valid under the recently proposed dueling architecture.

Comments:	Presented at the NIPS Deep Reinforcement Learning Workshop, Montreal, Canada, 2018
Subjects:	Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1708.01867 [cs.AI]
	(or arXiv:1708.01867v5 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.1708.01867

Submission history

From: Felix Leibfried [view email]
[v1] Sun, 6 Aug 2017 09:23:22 UTC (527 KB)
[v2] Wed, 8 Nov 2017 16:27:50 UTC (989 KB)
[v3] Thu, 8 Feb 2018 14:07:53 UTC (1,166 KB)
[v4] Thu, 6 Sep 2018 09:27:21 UTC (692 KB)
[v5] Tue, 20 Nov 2018 11:55:21 UTC (557 KB)

Full-text links:

Access Paper:

view license

Current browse context:

stat

< prev | next >

new | recent | 2017-08

Change to browse by:

cs
cs.AI
cs.LG
stat.ML

References & Citations

DBLP - CS Bibliography

listing | bibtex

Felix Leibfried
Jordi Grau-Moya
Haitham Bou-Ammar

export BibTeX citation

Computer Science > Artificial Intelligence

Title:An Information-Theoretic Optimality Principle for Deep Reinforcement Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:An Information-Theoretic Optimality Principle for Deep Reinforcement Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators