Second Order Bounds for Contextual Bandits with Function Approximation

Pacchiano, Aldo

Computer Science > Machine Learning

arXiv:2409.16197 (cs)

[Submitted on 24 Sep 2024]

Title:Second Order Bounds for Contextual Bandits with Function Approximation

Authors:Aldo Pacchiano

View PDF HTML (experimental)

Abstract:Many works have developed algorithms no-regret algorithms for contextual bandits with function approximation, where the mean rewards over context-action pairs belongs to a function class. Although there are many approaches to this problem, one that has gained in importance is the use of algorithms based on the optimism principle such as optimistic least squares. It can be shown the regret of this algorithm scales as square root of the product of the eluder dimension (a statistical measure of the complexity of the function class), the logarithm of the function class size and the time horizon. Unfortunately, even if the variance of the measurement noise of the rewards at each time is changing and is very small, the regret of the optimistic least squares algorithm scales with square root of the time horizon. In this work we are the first to develop algorithms that satisfy regret bounds of scaling not with the square root of the time horizon, but the square root of the sum of the measurement variances in the setting of contextual bandits with function approximation when the variances are unknown. These bounds generalize existing techniques for deriving second order bounds in contextual linear problems.

Comments:	12 pages main, 33 pages total
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as:	arXiv:2409.16197 [cs.LG]
	(or arXiv:2409.16197v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2409.16197

Submission history

From: Aldo Pacchiano [view email]
[v1] Tue, 24 Sep 2024 15:42:04 UTC (44 KB)

Computer Science > Machine Learning

Title:Second Order Bounds for Contextual Bandits with Function Approximation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Second Order Bounds for Contextual Bandits with Function Approximation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators