Stabilizing Policy Gradients for Stochastic Differential Equations via Consistency with Perturbation Process

Zhou, Xiangxin; Wang, Liang; Zhou, Yichi

Computer Science > Machine Learning

arXiv:2403.04154 (cs)

[Submitted on 7 Mar 2024 (v1), last revised 26 Jun 2024 (this version, v2)]

Title:Stabilizing Policy Gradients for Stochastic Differential Equations via Consistency with Perturbation Process

Authors:Xiangxin Zhou, Liang Wang, Yichi Zhou

View PDF HTML (experimental)

Abstract:Considering generating samples with high rewards, we focus on optimizing deep neural networks parameterized stochastic differential equations (SDEs), the advanced generative models with high expressiveness, with policy gradient, the leading algorithm in reinforcement learning. Nevertheless, when applying policy gradients to SDEs, since the policy gradient is estimated on a finite set of trajectories, it can be ill-defined, and the policy behavior in data-scarce regions may be uncontrolled. This challenge compromises the stability of policy gradients and negatively impacts sample complexity. To address these issues, we propose constraining the SDE to be consistent with its associated perturbation process. Since the perturbation process covers the entire space and is easy to sample, we can mitigate the aforementioned problems. Our framework offers a general approach allowing for a versatile selection of policy gradient methods to effectively and efficiently train SDEs. We evaluate our algorithm on the task of structure-based drug design and optimize the binding affinity of generated ligand molecules. Our method achieves the best Vina score -9.07 on the CrossDocked2020 dataset.

Comments:	Accepted to ICML 2024
Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2403.04154 [cs.LG]
	(or arXiv:2403.04154v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2403.04154

Submission history

From: Xiangxin Zhou [view email]
[v1] Thu, 7 Mar 2024 02:24:45 UTC (15,416 KB)
[v2] Wed, 26 Jun 2024 02:28:07 UTC (15,417 KB)

Computer Science > Machine Learning

Title:Stabilizing Policy Gradients for Stochastic Differential Equations via Consistency with Perturbation Process

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Stabilizing Policy Gradients for Stochastic Differential Equations via Consistency with Perturbation Process

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators