Abstract
Post-hoc model-agnostic local attribution (LA) methods have been widely adopted to explain opaque AI models by quantifying feature-wise contributions. However, many existing methods rely on heuristic or only partially justified attribution mechanisms, while the quality of attribution itself is often shaped by downstream objectives without universally accepted standards. In this work, we propose Taylor exPansion-Originated aDaptive Attribution (TaylorPODA), a new post-hoc model-agnostic LA method grounded in the Taylor expansion framework. We first introduce a set of postulates, which formalize principled requirements for explicitly and exhaustively attributing Taylor terms to the corresponding features. Based on these postulates, we analyze existing post-hoc model-agnostic LA methods and identify a fundamental tension between principled attribution and adaptation toward user-defined utilities. To address this challenge, TaylorPODA introduces a controllable allocation mechanism for Taylor interaction effects, enabling attribution results to adapt to downstream objectives while preserving the proposed postulates. Furthermore, although developed from a Taylor-expansion perspective, TaylorPODA also admits a Harsanyi-dividend interpretation, allowing the attribution mechanism to extend beyond model differentiability. Theoretical analysis demonstrates that TaylorPODA satisfies all the proposed postulates together with an additional adaptation property. Empirical results across multiple datasets and both differentiable and non-differentiable models further show that TaylorPODA achieves consistently improved alignment with user-defined utilities while maintaining the communicability of the resulting explanations. Overall, this work provides a starting point toward more trustworthy XAI systems for the deployment of increasingly powerful yet opaque task models.