Context-Based Soft Actor Critic for Environments with Non-stationary Dynamics

Pu, Yuan; Wang, Shaochen; Yao, Xin; Li, Bin

Computer Science > Machine Learning

arXiv:2105.03310 (cs)

[Submitted on 7 May 2021 (v1), last revised 10 May 2021 (this version, v2)]

Title:Context-Based Soft Actor Critic for Environments with Non-stationary Dynamics

Authors:Yuan Pu, Shaochen Wang, Xin Yao, Bin Li

View PDF

Abstract:The performance of deep reinforcement learning methods prone to degenerate when applied to environments with non-stationary dynamics. In this paper, we utilize the latent context recurrent encoders motivated by recent Meta-RL materials, and propose the Latent Context-based Soft Actor Critic (LC-SAC) method to address aforementioned issues. By minimizing the contrastive prediction loss function, the learned context variables capture the information of the environment dynamics and the recent behavior of the agent. Then combined with the soft policy iteration paradigm, the LC-SAC method alternates between soft policy evaluation and soft policy improvement until it converges to the optimal policy. Experimental results show that the performance of LC-SAC is significantly better than the SAC algorithm on the MetaWorld ML1 tasks whose dynamics changes drasticly among different episodes, and is comparable to SAC on the continuous control benchmark task MuJoCo whose dynamics changes slowly or doesn't change between different episodes. In addition, we also conduct relevant experiments to determine the impact of different hyperparameter settings on the performance of the LC-SAC algorithm and give the reasonable suggestions of hyperparameter setting.

Comments:	12 pages, 11 figures
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as:	arXiv:2105.03310 [cs.LG]
	(or arXiv:2105.03310v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2105.03310

Submission history

From: Yuan Pu [view email]
[v1] Fri, 7 May 2021 15:00:59 UTC (7,167 KB)
[v2] Mon, 10 May 2021 09:25:53 UTC (7,167 KB)

Computer Science > Machine Learning

Title:Context-Based Soft Actor Critic for Environments with Non-stationary Dynamics

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Context-Based Soft Actor Critic for Environments with Non-stationary Dynamics

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators