Learning Time-Invariant Reward Functions through Model-Based Inverse Reinforcement Learning

Abstract

Inverse reinforcement learning is a paradigm motivated by the goal oflearning general reward functions from demonstrated behaviours. Yet the notionof generality for learnt costs is often evaluated in terms of robustness tovarious spatial perturbations only, assuming deployment at fixed speeds ofexecution. However, this is impractical in the context of robotics andbuilding, time-invariant solutions is of crucial importance. In this work, wepropose a formulation that allows us to 1) vary the length of execution bylearning time-invariant costs, and 2) relax the temporal alignment requirementsfor learning from demonstration. We apply our method to two different types ofcost formulations and evaluate their performance in the context of learningreward functions for simulated placement and peg in hole tasks executed on a7DoF Kuka IIWA arm. Our results show that our approach enables learningtemporally invariant rewards from misaligned demonstration that can alsogeneralise spatially to out of distribution tasks.

Quick Read (beta)

loading the full paper ...