LGR2: Language Guided Reward Relabeling for Accelerating Hierarchical Reinforcement Learning

  • 2025-08-27 17:57:18
  • Utsav Singh, Pramit Bhattacharyya, Vinay P. Namboodiri
  • 0

Abstract

Large language models (LLMs) have shown remarkable abilities in logicalreasoning, in-context learning, and code generation. However, translatingnatural language instructions into effective robotic control policies remains asignificant challenge, especially for tasks requiring long-horizon planning andoperating under sparse reward conditions. Hierarchical Reinforcement Learning(HRL) provides a natural framework to address this challenge in robotics;however, it typically suffers from non-stationarity caused by the changingbehavior of the lower-level policy during training, destabilizing higher-levelpolicy learning. We introduce LGR2, a novel HRL framework that leverages LLMsto generate language-guided reward functions for the higher-level policy. Bydecoupling high-level reward generation from low-level policy changes, LGR2fundamentally mitigates the non-stationarity problem in off-policy HRL,enabling stable and efficient learning. To further enhance sample efficiency insparse environments, we integrate goal-conditioned hindsight experiencerelabeling. Extensive experiments across simulated and real-world roboticnavigation and manipulation tasks demonstrate LGR2 outperforms bothhierarchical and non-hierarchical baselines, achieving over 55% success rateson challenging tasks and robust transfer to real robots, without additionalfine-tuning.

 

Quick Read (beta)

loading the full paper ...