Value Memory Graph: A Graph-Structured World Model for Offline Reinforcement Learning

Zhu, Deyao; Li, Li Erran; Elhoseiny, Mohamed

Computer Science > Machine Learning

arXiv:2206.04384 (cs)

[Submitted on 9 Jun 2022 (v1), last revised 2 May 2023 (this version, v3)]

Title:Value Memory Graph: A Graph-Structured World Model for Offline Reinforcement Learning

Authors:Deyao Zhu, Li Erran Li, Mohamed Elhoseiny

View PDF

Abstract:Reinforcement Learning (RL) methods are typically applied directly in environments to learn policies. In some complex environments with continuous state-action spaces, sparse rewards, and/or long temporal horizons, learning a good policy in the original environments can be difficult. Focusing on the offline RL setting, we aim to build a simple and discrete world model that abstracts the original environment. RL methods are applied to our world model instead of the environment data for simplified policy learning. Our world model, dubbed Value Memory Graph (VMG), is designed as a directed-graph-based Markov decision process (MDP) of which vertices and directed edges represent graph states and graph actions, separately. As state-action spaces of VMG are finite and relatively small compared to the original environment, we can directly apply the value iteration algorithm on VMG to estimate graph state values and figure out the best graph actions. VMG is trained from and built on the offline RL dataset. Together with an action translator that converts the abstract graph actions in VMG to real actions in the original environment, VMG controls agents to maximize episode returns. Our experiments on the D4RL benchmark show that VMG can outperform state-of-the-art offline RL methods in several goal-oriented tasks, especially when environments have sparse rewards and long temporal horizons. Code is available at this https URL

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2206.04384 [cs.LG]
	(or arXiv:2206.04384v3 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2206.04384

Submission history

From: Deyao Zhu [view email]
[v1] Thu, 9 Jun 2022 09:51:42 UTC (7,065 KB)
[v2] Mon, 3 Oct 2022 19:30:04 UTC (9,449 KB)
[v3] Tue, 2 May 2023 14:15:02 UTC (9,654 KB)

Computer Science > Machine Learning

Title:Value Memory Graph: A Graph-Structured World Model for Offline Reinforcement Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Value Memory Graph: A Graph-Structured World Model for Offline Reinforcement Learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators