Proximal Ranking Policy Optimization for Practical Safety in Counterfactual Learning to Rank

Gupta, Shashank; Oosterhuis, Harrie; de Rijke, Maarten

Computer Science > Machine Learning

arXiv:2409.09881 (cs)

[Submitted on 15 Sep 2024]

Title:Proximal Ranking Policy Optimization for Practical Safety in Counterfactual Learning to Rank

Authors:Shashank Gupta, Harrie Oosterhuis, Maarten de Rijke

View PDF HTML (experimental)

Abstract:Counterfactual learning to rank (CLTR) can be risky and, in various circumstances, can produce sub-optimal models that hurt performance when deployed. Safe CLTR was introduced to mitigate these risks when using inverse propensity scoring to correct for position bias. However, the existing safety measure for CLTR is not applicable to state-of-the-art CLTR methods, cannot handle trust bias, and relies on specific assumptions about user behavior. We propose a novel approach, proximal ranking policy optimization (PRPO), that provides safety in deployment without assumptions about user behavior. PRPO removes incentives for learning ranking behavior that is too dissimilar to a safe ranking model. Thereby, PRPO imposes a limit on how much learned models can degrade performance metrics, without relying on any specific user assumptions. Our experiments show that PRPO provides higher performance than the existing safe inverse propensity scoring approach. PRPO always maintains safety, even in maximally adversarial situations. By avoiding assumptions, PRPO is the first method with unconditional safety in deployment that translates to robust safety for real-world applications.

Comments:	Accepted at the CONSEQUENCES 2024 workshop, co-located with ACM RecSys 2024
Subjects:	Machine Learning (cs.LG); Information Retrieval (cs.IR)
Cite as:	arXiv:2409.09881 [cs.LG]
	(or arXiv:2409.09881v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2409.09881

Submission history

From: Shashank Gupta [view email]
[v1] Sun, 15 Sep 2024 22:22:27 UTC (1,150 KB)

Computer Science > Machine Learning

Title:Proximal Ranking Policy Optimization for Practical Safety in Counterfactual Learning to Rank

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Proximal Ranking Policy Optimization for Practical Safety in Counterfactual Learning to Rank

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators