Fixed-Confidence Best Arm Identification with Decreasing Variance

Roychowdhury, Tamojeet; Reddy, Kota Srinivas; Jagannathan, Krishna P; Moharir, Sharayu

Computer Science > Machine Learning

arXiv:2502.07199 (cs)

[Submitted on 11 Feb 2025]

Title:Fixed-Confidence Best Arm Identification with Decreasing Variance

Authors:Tamojeet Roychowdhury, Kota Srinivas Reddy, Krishna P Jagannathan, Sharayu Moharir

View PDF HTML (experimental)

Abstract:We focus on the problem of best-arm identification in a stochastic multi-arm bandit with temporally decreasing variances for the arms' rewards. We model arm rewards as Gaussian random variables with fixed means and variances that decrease with time. The cost incurred by the learner is modeled as a weighted sum of the time needed by the learner to identify the best arm, and the number of samples of arms collected by the learner before termination. Under this cost function, there is an incentive for the learner to not sample arms in all rounds, especially in the initial rounds. On the other hand, not sampling increases the termination time of the learner, which also increases cost. This trade-off necessitates new sampling strategies. We propose two policies. The first policy has an initial wait period with no sampling followed by continuous sampling. The second policy samples periodically and uses a weighted average of the rewards observed to identify the best arm. We provide analytical guarantees on the performance of both policies and supplement our theoretical results with simulations which show that our polices outperform the state-of-the-art policies for the classical best arm identification problem.

Comments:	6 pages, 2 figures, accepted in the National Conference on Communications 2025
Subjects:	Machine Learning (cs.LG); Information Theory (cs.IT); Statistics Theory (math.ST); Machine Learning (stat.ML)
Cite as:	arXiv:2502.07199 [cs.LG]
	(or arXiv:2502.07199v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2502.07199

Submission history

From: Tamojeet Roychowdhury [view email]
[v1] Tue, 11 Feb 2025 02:47:20 UTC (119 KB)

Computer Science > Machine Learning

Title:Fixed-Confidence Best Arm Identification with Decreasing Variance

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Fixed-Confidence Best Arm Identification with Decreasing Variance

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators