On the Stochastic (Variance-Reduced) Proximal Gradient Method for Regularized Expected Reward Optimization

Liang, Ling; Yang, Haizhao

Computer Science > Machine Learning

arXiv:2401.12508 (cs)

[Submitted on 23 Jan 2024 (v1), last revised 19 Aug 2024 (this version, v2)]

Title:On the Stochastic (Variance-Reduced) Proximal Gradient Method for Regularized Expected Reward Optimization

Authors:Ling Liang, Haizhao Yang

View PDF HTML (experimental)

Abstract:We consider a regularized expected reward optimization problem in the non-oblivious setting that covers many existing problems in reinforcement learning (RL). In order to solve such an optimization problem, we apply and analyze the classical stochastic proximal gradient method. In particular, the method has shown to admit an $O(\epsilon^{-4})$ sample complexity to an $\epsilon$-stationary point, under standard conditions. Since the variance of the classical stochastic gradient estimator is typically large, which slows down the convergence, we also apply an efficient stochastic variance-reduce proximal gradient method with an importance sampling based ProbAbilistic Gradient Estimator (PAGE). Our analysis shows that the sample complexity can be improved from $O(\epsilon^{-4})$ to $O(\epsilon^{-3})$ under additional conditions. Our results on the stochastic (variance-reduced) proximal gradient method match the sample complexity of their most competitive counterparts for discounted Markov decision processes under similar settings. To the best of our knowledge, the proposed methods represent a novel approach in addressing the general regularized reward optimization problem.

Comments:	23 pages
Subjects:	Machine Learning (cs.LG); Optimization and Control (math.OC)
Cite as:	arXiv:2401.12508 [cs.LG]
	(or arXiv:2401.12508v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2401.12508

Submission history

From: Ling Liang [view email]
[v1] Tue, 23 Jan 2024 06:01:29 UTC (53 KB)
[v2] Mon, 19 Aug 2024 19:18:48 UTC (46 KB)

Computer Science > Machine Learning

Title:On the Stochastic (Variance-Reduced) Proximal Gradient Method for Regularized Expected Reward Optimization

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:On the Stochastic (Variance-Reduced) Proximal Gradient Method for Regularized Expected Reward Optimization

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators