Can counterfactual explanations of AI systems' predictions skew lay users' causal intuitions about the world? If so, can we correct for that?

Abstract

Counterfactual (CF) explanations have been employed as one of the modes ofexplainability in explainable AI-both to increase the transparency of AIsystems and to provide recourse. Cognitive science and psychology, however,have pointed out that people regularly use CFs to express causal relationships.Most AI systems are only able to capture associations or correlations in dataso interpreting them as casual would not be justified. In this paper, wepresent two experiment (total N = 364) exploring the effects of CF explanationsof AI system's predictions on lay people's causal beliefs about the real world.In Experiment 1 we found that providing CF explanations of an AI system'spredictions does indeed (unjustifiably) affect people's causal beliefsregarding factors/features the AI uses and that people are more likely to viewthem as causal factors in the real world. Inspired by the literature onmisinformation and health warning messaging, Experiment 2 tested whether we cancorrect for the unjustified change in causal beliefs. We found that pointingout that AI systems capture correlations and not necessarily causalrelationships can attenuate the effects of CF explanations on people's causalbeliefs.

Quick Read (beta)

loading the full paper ...