Orthogonalized Estimation of Difference of $Q$-functions

Cao, Defu; Zhou, Angela

Statistics > Machine Learning

arXiv:2406.08697 (stat)

[Submitted on 12 Jun 2024 (v1), last revised 16 Oct 2024 (this version, v2)]

Title:Orthogonalized Estimation of Difference of $Q$-functions

Authors:Defu Cao, Angela Zhou

View PDF HTML (experimental)

Abstract:Offline reinforcement learning is important in many settings with available observational data but the inability to deploy new policies online due to safety, cost, and other concerns. Many recent advances in causal inference and machine learning target estimation of causal contrast functions such as CATE, which is sufficient for optimizing decisions and can adapt to potentially smoother structure. We develop a dynamic generalization of the R-learner (Nie and Wager 2021, Lewis and Syrgkanis 2021) for estimating and optimizing the difference of $Q^\pi$-functions, $Q^\pi(s,1)-Q^\pi(s,0)$ (which can be used to optimize multiple-valued actions). We leverage orthogonal estimation to improve convergence rates in the presence of slower nuisance estimation rates and prove consistency of policy optimization under a margin condition. The method can leverage black-box nuisance estimators of the $Q$-function and behavior policy to target estimation of a more structured $Q$-function contrast.

Subjects:	Machine Learning (stat.ML); Machine Learning (cs.LG); Optimization and Control (math.OC); Methodology (stat.ME)
Cite as:	arXiv:2406.08697 [stat.ML]
	(or arXiv:2406.08697v2 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.2406.08697

Submission history

From: Angela Zhou [view email]
[v1] Wed, 12 Jun 2024 23:41:43 UTC (69 KB)
[v2] Wed, 16 Oct 2024 23:41:36 UTC (123 KB)

Statistics > Machine Learning

Title:Orthogonalized Estimation of Difference of $Q$-functions

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:Orthogonalized Estimation of Difference of $Q$-functions

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators