Rarely-switching linear bandits: optimization of causal effects for the real world

Abstract

Exploring the effect of policies in many real world scenarios is difficult,unethical, or expensive. After all, doctor guidelines, tax codes, and pricelists can only be reprinted so often. We may thus want to only change a policywhen it is probable that the change is beneficial. Fortunately, thresholdsallow us to estimate treatment effects. Such estimates allows us to optimizethe threshold. Here, based on the theory of linear contextual bandits, wepresent a conservative policy updating procedure which updates a deterministicpolicy only when needed. We extend the theory of linear bandits to thisrarely-switching case, proving such procedures share the same regret, up toconstant scaling, as the common LinUCB algorithm. However the algorithm makesfar fewer changes to its policy. We provide simulations and an analysis of aninfant health well-being causal inference dataset, showing the algorithmefficiently learns a good policy with few changes. Our approach allowsefficiently solving problems where changes are to be avoided, with potentialapplications in economics, medicine and beyond.

Quick Read (beta)

loading the full paper ...