Disentangling Recognition and Decision Regrets in Image-Based Reinforcement Learning

Abstract

In image-based reinforcement learning (RL), policies usually operate in twosteps: first extracting lower-dimensional features from raw images (the"recognition" step), and then taking actions based on the extracted features(the "decision" step). Extracting features that are spuriously correlated withperformance or irrelevant for decision-making can lead to poor generalizationperformance, known as observational overfitting in image-based RL. In suchcases, it can be hard to quantify how much of the error can be attributed topoor feature extraction vs. poor decision-making. To disentangle the twosources of error, we introduce the notions of recognition regret and decisionregret. Using these notions, we characterize and disambiguate the two distinctcauses behind observational overfitting: over-specific representations, whichinclude features that are not needed for optimal decision-making (leading tohigh decision regret), vs. under-specific representations, which only include alimited set of features that were spuriously correlated with performance duringtraining (leading to high recognition regret). Finally, we provide illustrativeexamples of observational overfitting due to both over-specific andunder-specific representations in maze environments and the Atari game Pong.

Quick Read (beta)

loading the full paper ...