Neuro-Symbolic World Models for Adapting to Open World Novelty

Abstract

Open-world novelty--a sudden change in the mechanics or properties of anenvironment--is a common occurrence in the real world. Novelty adaptation is anagent's ability to improve its policy performance post-novelty. Mostreinforcement learning (RL) methods assume that the world is a closed, fixedprocess. Consequentially, RL policies adapt inefficiently to novelties. Toaddress this, we introduce WorldCloner, an end-to-end trainable neuro-symbolicworld model for rapid novelty adaptation. WorldCloner learns an efficientsymbolic representation of the pre-novelty environment transitions, and usesthis transition model to detect novelty and efficiently adapt to novelty in asingle-shot fashion. Additionally, WorldCloner augments the policy learningprocess using imagination-based adaptation, where the world model simulatestransitions of the post-novelty environment to help the policy adapt. Byblending ''imagined'' transitions with interactions in the post-noveltyenvironment, performance can be recovered with fewer total environmentinteractions. Using environments designed for studying novelty in sequentialdecision-making problems, we show that the symbolic world model helps itsneural policy adapt more efficiently than model-based and model-basedneural-only reinforcement learning methods.

Quick Read (beta)

loading the full paper ...