Reward Shaping for Reinforcement Learning with Omega-Regular Objectives

  • 2020-01-16 18:22:50
  • E. M. Hahn, M. Perez, S. Schewe, F. Somenzi, A. Trivedi, D. Wojtczak
  • 9

Abstract

Recently, successful approaches have been made to exploit good-for-MDPsautomata (B\"uchi automata with a restricted form of nondeterminism) for modelfree reinforcement learning, a class of automata that subsumes good for gamesautomata and the most widespread class of limit deterministic automata. Thefoundation of using these B\"uchi automata is that the B\"uchi condition can,for good-for-MDP automata, be translated to reachability. The drawback of this translation is that the rewards are, on average, reapedvery late, which requires long episodes during the learning process. We devisea new reward shaping approach that overcomes this issue. We show that theresulting model is equivalent to a discounted payoff objective with a biaseddiscount that simplifies and improves on prior work in this direction.

 

Quick Read (beta)

loading the full paper ...