LOPT: Learning Optimal Pigovian Tax in Sequential Social Dilemmas

Yun Hua (Shanghai Jiaotong University) · Shang Gao (Deakin University) · Wenhao Li (Shandong University) · Haosheng Chen (Chongqing University of Post and Telecommunications) · Bo Jin (Tongji University) · Xiangfeng Wang (East China Normal University) · Jun Luo (Nanyang Technological University) · Hongyuan Zha (The Chinese University of Hong Kong, Shenzhen)
agent coordinationauxiliary tax agentcollective goalsexternalitiesindividual objectivesmarl benchmarksmulti-agent reinforcement learningnegative societal impactsnumerical representationoptimal tax policypigovian taxrational self-interested behaviorssocial dilemmassocial welfaretheoretical analysisunaccounted-for impact

Multi-agent reinforcement learning (MARL) has emerged as a powerful framework for modeling autonomous agents that independently optimize their individual objectives. However, in mixed-motive MARL environments, rational self-interested behaviors often lead to collectively suboptimal outcomes situations commonly referred to as social dilemmas. A key challenge in addressing social dilemmas lies in accurately quantifying and representing them in a numerical form that captures how self-interested agent behaviors impact social welfare. To address this challenge, \textit{externalities} in the economic concept is adopted and extended to denote the unaccounted-for impact of one agent's actions on others, as a means to rigorously quantify social dilemmas. Based on this measurement, a novel method, \textbf{L}earning \textbf{O}ptimal \textbf{P}igovian \textbf{T}ax (\textbf{LOPT}) is proposed. Inspired by Pigovian taxes, which are designed to internalize externalities by imposing cost on negative societal impacts, LOPT employs an auxiliary tax agent that learns an optimal Pigovian tax policy to reshape individual rewards aligned with social welfare, thereby promoting agent coordination and mitigating social dilemmas. We support LOPT with theoretical analysis and validate it on standard MARL benchmarks, including Escape Room and Cleanup. Results show that by effectively internalizing externalities that quantify social dilemmas, LOPT aligns individual objectives with collective goals, significantly improving social welfare over state-of-the-art baselines.