Robust Thompson Sampling Algorithms Against Reward Poisoning Attacks

Xu, Yinglun; Wang, Zhiwei; Singh, Gagandeep

Computer Science > Machine Learning

arXiv:2410.19705 (cs)

[Submitted on 25 Oct 2024]

Title:Robust Thompson Sampling Algorithms Against Reward Poisoning Attacks

Authors:Yinglun Xu, Zhiwei Wang, Gagandeep Singh

View PDF HTML (experimental)

Abstract:Thompson sampling is one of the most popular learning algorithms for online sequential decision-making problems and has rich real-world applications. However, current Thompson sampling algorithms are limited by the assumption that the rewards received are uncorrupted, which may not be true in real-world applications where adversarial reward poisoning exists. To make Thompson sampling more reliable, we want to make it robust against adversarial reward poisoning. The main challenge is that one can no longer compute the actual posteriors for the true reward, as the agent can only observe the rewards after corruption. In this work, we solve this problem by computing pseudo-posteriors that are less likely to be manipulated by the attack. We propose robust algorithms based on Thompson sampling for the popular stochastic and contextual linear bandit settings in both cases where the agent is aware or unaware of the budget of the attacker. We theoretically show that our algorithms guarantee near-optimal regret under any attack strategy.

Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2410.19705 [cs.LG]
	(or arXiv:2410.19705v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2410.19705

Submission history

From: Yinglun Xu [view email]
[v1] Fri, 25 Oct 2024 17:27:58 UTC (1,407 KB)

Computer Science > Machine Learning

Title:Robust Thompson Sampling Algorithms Against Reward Poisoning Attacks

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Robust Thompson Sampling Algorithms Against Reward Poisoning Attacks

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators