AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Abstract

Reinforcement learning (RL) provides an appealing formalism for learningcontrol policies from experience. However, the classic active formulation of RLnecessitates a lengthy active exploration process for each behavior, making itdifficult to apply in real-world settings such as robotic control. If we caninstead allow RL algorithms to effectively use previously collected data to aidthe online learning process, such applications could be made substantially morepractical: the prior data would provide a starting point that mitigateschallenges due to exploration and sample complexity, while the online trainingenables the agent to perfect the desired skill. Such prior data could eitherconstitute expert demonstrations or sub-optimal prior data that illustratespotentially useful transitions. While a number of prior methods have eitherused optimal demonstrations to bootstrap RL, or have used sub-optimal data totrain purely offline, it remains exceptionally difficult to train a policy withoffline data and actually continue to improve it further with online RL. Inthis paper we analyze why this problem is so challenging, and propose analgorithm that combines sample efficient dynamic programming with maximumlikelihood policy updates, providing a simple and effective framework that isable to leverage large amounts of offline data and then quickly perform onlinefine-tuning of RL policies. We show that our method, advantage weighted actorcritic (AWAC), enables rapid learning of skills with a combination of priordemonstration data and online experience. We demonstrate these benefits onsimulated and real-world robotics domains, including dexterous manipulationwith a real multi-fingered hand, drawer opening with a robotic arm, androtating a valve. Our results show that incorporating prior data can reduce thetime required to learn a range of robotic skills to practical time-scales.

Quick Read (beta)

loading the full paper ...