Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control
The paper gives a formal test for when a panel of imperfect AI reviewers can safely approve another agent’s actions.
It defines “k-robust coalitional alignment,” a condition under which a threshold approval rule remains safe even if up to k reviewers disapprove. The guarantee is that the principal does at least as well in expectation as under a baseline policy. The result extends to sequential control in discounted MDPs, where state-by-state safety is necessary and sufficient. The authors also warn that strategic voting can make permissive thresholds unsafe, even when reviewers are individually aligned. ArXiv · AI/CL/LG's note
It defines “k-robust coalitional alignment,” a condition under which a threshold approval rule remains safe even if up to k reviewers disapprove. The guarantee is that the principal does at least as well in expectation as under a baseline policy. The result extends to sequential control in discounted MDPs, where state-by-state safety is necessary and sufficient. The authors also warn that strategic voting can make permissive thresholds unsafe, even when reviewers are individually aligned. ArXiv · AI/CL/LG's note
score 4