On Corrigibility and Alignment in Multi Agent Games

Dable-Heath, Edmund; Vodenicharski, Boyko; Bishop, James

Computer Science > Computer Science and Game Theory

arXiv:2501.05360 (cs)

[Submitted on 9 Jan 2025]

Title:On Corrigibility and Alignment in Multi Agent Games

Authors:Edmund Dable-Heath, Boyko Vodenicharski, James Bishop

View PDF HTML (experimental)

Abstract:Corrigibility of autonomous agents is an under explored part of system design, with previous work focusing on single agent systems. It has been suggested that uncertainty over the human preferences acts to keep the agents corrigible, even in the face of human irrationality. We present a general framework for modelling corrigibility in a multi-agent setting as a 2 player game in which the agents always have a move in which they can ask the human for supervision. This is formulated as a Bayesian game for the purpose of introducing uncertainty over the human beliefs. We further analyse two specific cases. First, a two player corrigibility game, in which we want corrigibility displayed in both agents for both common payoff (monotone) games and harmonic games. Then we investigate an adversary setting, in which one agent is considered to be a `defending' agent and the other an `adversary'. A general result is provided for what belief over the games and human rationality the defending agent is required to have to induce corrigibility.

Subjects:	Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2501.05360 [cs.GT]
	(or arXiv:2501.05360v1 [cs.GT] for this version)
	https://doi.org/10.48550/arXiv.2501.05360

Submission history

From: Edmund Dable-Heath [view email]
[v1] Thu, 9 Jan 2025 16:44:38 UTC (3,612 KB)

Computer Science > Computer Science and Game Theory

Title:On Corrigibility and Alignment in Multi Agent Games

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Science and Game Theory

Title:On Corrigibility and Alignment in Multi Agent Games

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators