Learn to Exceed: Stereo Inverse Reinforcement Learning with Concurrent Policy Optimization

Tao, Feng; Cao, Yongcan

Computer Science > Machine Learning

arXiv:2009.09577 (cs)

[Submitted on 21 Sep 2020 (v1), last revised 22 Sep 2020 (this version, v2)]

Title:Learn to Exceed: Stereo Inverse Reinforcement Learning with Concurrent Policy Optimization

Authors:Feng Tao, Yongcan Cao

View PDF

Abstract:In this paper, we study the problem of obtaining a control policy that can mimic and then outperform expert demonstrations in Markov decision processes where the reward function is unknown to the learning agent. One main relevant approach is the inverse reinforcement learning (IRL), which mainly focuses on inferring a reward function from expert demonstrations. The obtained control policy by IRL and the associated algorithms, however, can hardly outperform expert demonstrations. To overcome this limitation, we propose a novel method that enables the learning agent to outperform the demonstrator via a new concurrent reward and action policy learning approach. In particular, we first propose a new stereo utility definition that aims to address the bias in the interpretation of expert demonstrations. We then propose a loss function for the learning agent to learn reward and action policies concurrently such that the learning agent can outperform expert demonstrations. The performance of the proposed method is first demonstrated in OpenAI environments. Further efforts are conducted to experimentally validate the proposed method via an indoor drone flight scenario.

Comments:	12 pages, 5 figures
Subjects:	Machine Learning (cs.LG); Robotics (cs.RO); Machine Learning (stat.ML)
Cite as:	arXiv:2009.09577 [cs.LG]
	(or arXiv:2009.09577v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2009.09577

Submission history

From: Yongcan Cao [view email]
[v1] Mon, 21 Sep 2020 02:16:21 UTC (502 KB)
[v2] Tue, 22 Sep 2020 23:04:09 UTC (503 KB)

Computer Science > Machine Learning

Title:Learn to Exceed: Stereo Inverse Reinforcement Learning with Concurrent Policy Optimization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Learn to Exceed: Stereo Inverse Reinforcement Learning with Concurrent Policy Optimization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators