Periodic Intra-Ensemble Knowledge Distillation for Reinforcement Learning

Hong, Zhang-Wei; Nagarajan, Prabhat; Maeda, Guilherme

Computer Science > Machine Learning

arXiv:2002.00149 (cs)

[Submitted on 1 Feb 2020]

Title:Periodic Intra-Ensemble Knowledge Distillation for Reinforcement Learning

Authors:Zhang-Wei Hong, Prabhat Nagarajan, Guilherme Maeda

View PDF

Abstract:Off-policy ensemble reinforcement learning (RL) methods have demonstrated impressive results across a range of RL benchmark tasks. Recent works suggest that directly imitating experts' policies in a supervised manner before or during the course of training enables faster policy improvement for an RL agent. Motivated by these recent insights, we propose Periodic Intra-Ensemble Knowledge Distillation (PIEKD). PIEKD is a learning framework that uses an ensemble of policies to act in the environment while periodically sharing knowledge amongst policies in the ensemble through knowledge distillation. Our experiments demonstrate that PIEKD improves upon a state-of-the-art RL method in sample efficiency on several challenging MuJoCo benchmark tasks. Additionally, we perform ablation studies to better understand PIEKD.

Comments:	8 pages
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2002.00149 [cs.LG]
	(or arXiv:2002.00149v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2002.00149

Submission history

From: Zhang-Wei Hong [view email]
[v1] Sat, 1 Feb 2020 06:00:12 UTC (908 KB)

Full-text links:

Access Paper:

view license

Current browse context:

< prev | next >

new | recent | 2020-02

Change to browse by:

cs.AI
cs.LG

References & Citations

DBLP - CS Bibliography

listing | bibtex

Zhang-Wei Hong
Prabhat Nagarajan
Guilherme Maeda

export BibTeX citation

Computer Science > Machine Learning

Title:Periodic Intra-Ensemble Knowledge Distillation for Reinforcement Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Periodic Intra-Ensemble Knowledge Distillation for Reinforcement Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators