Return-Based Contrastive Representation Learning for Reinforcement Learning

Abstract

Recently, various auxiliary tasks have been proposed to acceleraterepresentation learning and improve sample efficiency in deep reinforcementlearning (RL). However, existing auxiliary tasks do not take thecharacteristics of RL problems into consideration and are unsupervised. Byleveraging returns, the most important feedback signals in RL, we propose anovel auxiliary task that forces the learnt representations to discriminatestate-action pairs with different returns. Our auxiliary loss is theoreticallyjustified to learn representations that capture the structure of a new form ofstate-action abstraction, under which state-action pairs with similar returndistributions are aggregated together. In low data regime, our algorithmoutperforms strong baselines on complex tasks in Atari games and DeepMindControl suite, and achieves even better performance when combined with existingauxiliary tasks.

Quick Read (beta)

loading the full paper ...