Estimating Quality in User-Guided Multi-Objective Bandits Optimization

Durand, Audrey; Gagné, Christian

Computer Science > Machine Learning

arXiv:1701.01095v1 (cs)

[Submitted on 4 Jan 2017 (this version), latest version 20 Apr 2017 (v3)]

Title:Estimating Quality in User-Guided Multi-Objective Bandits Optimization

Authors:Audrey Durand, Christian Gagné

View PDF

Abstract:Many real-world applications are characterized by a number of conflicting performance measures. As optimizing in a multi-objective setting leads to a set of non-dominated solutions, a preference function is required for selecting the solution with the appropriate trade-off between the objectives. This preference function is often unknown, especially when it comes from an expert human user. However, if we could provide the expert user with a proper estimation for each action, she would be able to pick her best choice. The question is: how good do these estimations have to be in order for her choice to remain the same as if she had access to the exact values? In this paper, we introduce the concept of preference radius to characterize the robustness of the preference function and provide guidelines for controlling the quality of estimations in the multi-objective setting. More specifically, we provide a general formulation of multi-objective optimization under the bandits setting and the pure exploration setting with user feedback for articulating the preferences. We show how the preference radius relates to the optimal gap and how it can be used to analyze algorithms in the bandits and pure exploration settings. We finally present experiments in the bandits setting, where we evaluate the impact of noise and delayed expert user feedback, and in the pure exploration setting, where we compare multi-objective Thompson sampling with uniform sampling.

Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1701.01095 [cs.LG]
	(or arXiv:1701.01095v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1701.01095

Submission history

From: Audrey Durand [view email]
[v1] Wed, 4 Jan 2017 18:20:47 UTC (5,548 KB)
[v2] Sat, 1 Apr 2017 21:23:44 UTC (2,117 KB)
[v3] Thu, 20 Apr 2017 20:37:39 UTC (1,938 KB)

Computer Science > Machine Learning

Title:Estimating Quality in User-Guided Multi-Objective Bandits Optimization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Estimating Quality in User-Guided Multi-Objective Bandits Optimization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators