Understanding Action Effects through Instrumental Empowerment in Multi-Agent Reinforcement Learning

  • 2025-08-21 15:35:59
  • Ardian Selmonaj, Miroslav Strupl, Oleg Szehr, Alessandro Antonucci
  • 0

Abstract

To reliably deploy Multi-Agent Reinforcement Learning (MARL) systems, it iscrucial to understand individual agent behaviors within a team. While priorwork typically evaluates overall team performance based on explicit rewardsignals or learned value functions, it is unclear how to infer agentcontributions in the absence of any value feedback. In this work, weinvestigate whether meaningful insights into agent behaviors can be extractedthat are consistent with the underlying value functions, solely by analyzingthe policy distribution. Inspired by the phenomenon that intelligent agentstend to pursue convergent instrumental values, which generally increase thelikelihood of task success, we introduce Intended Cooperation Values (ICVs), amethod based on information-theoretic Shapley values for quantifying eachagent's causal influence on their co-players' instrumental empowerment.Specifically, ICVs measure an agent's action effect on its teammates' policiesby assessing their decision uncertainty and preference alignment. The analysisacross cooperative and competitive MARL environments reveals the extent towhich agents adopt similar or diverse strategies. By comparing action effectsbetween policies and value functions, our method identifies which agentbehaviors are beneficial to team success, either by fostering deterministicdecisions or by preserving flexibility for future action choices. Our proposedmethod offers novel insights into cooperation dynamics and enhancesexplainability in MARL systems.

 

Quick Read (beta)

loading the full paper ...