A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning

  • 2025-09-19 12:44:29
  • Shaopeng Zhai, Qi Zhang, Tianyi Zhang, Fuxian Huang, Haoran Zhang, Ming Zhou, Shengzhe Zhang, Litao Liu, Sixu Lin, Jiangmiao Pang
  • 0

Abstract

Robotic real-world reinforcement learning (RL) with vision-language-action(VLA) models is bottlenecked by sparse, handcrafted rewards and inefficientexploration. We introduce VLAC, a general process reward model built uponInternVL and trained on large scale heterogeneous datasets. Given pairwiseobservations and a language goal, it outputs dense progress delta and donesignal, eliminating task-specific reward engineering, and supports one-shotin-context transfer to unseen tasks and environments. VLAC is trained onvision-language datasets to strengthen perception, dialogic and reasoningcapabilities, together with robot and human trajectories data that groundaction generation and progress estimation, and additionally strengthened toreject irrelevant prompts as well as detect regression or stagnation byconstructing large numbers of negative and semantically mismatched samples.With prompt control, a single VLAC model alternately generating reward andaction tokens, unifying critic and policy. Deployed inside an asynchronousreal-world RL loop, we layer a graded human-in-the-loop protocol (offlinedemonstration replay, return and explore, human guided explore) thataccelerates exploration and stabilizes early learning. Across four distinctreal-world manipulation tasks, VLAC lifts success rates from about 30\% toabout 90\% within 200 real-world interaction episodes; incorporatinghuman-in-the-loop interventions yields a further 50% improvement in sampleefficiency and achieves up to 100% final success.

 

Quick Read (beta)

loading the full paper ...