arXiv · 2410.07096
Rejecting Hallucinated State Targets during Planning
Abstract
In planning processes of computational decision-making agents, generative or predictive models are often used as "generators" to propose "targets" representing sets of expected or desirable states. Unfortunately, learned models inevitably hallucinate infeasible targets that can cause delusional behaviors and safety concerns. We first investigate the kinds of infeasible targets that generators can hallucinate. Then, we devise a strategy to identify and reject infeasible targets by learning a target feasibility evaluator. To ensure that the evaluator is robust and non-delusional, we adopted a design choice combining off-policy compatible learning rule, distributional architecture, and data augmentation based on hindsight relabeling. Attaching to a planning agent, the designed evaluator learns by observing the agent's interactions with the environment and the targets produced by its generator, without the need to change the agent or its generator. Our controlled experiments show significant reductions in delusional behaviors and performance improvements for various kinds of existing agents.
Explore related subjects
Keep this discovery
Mingde Zhao, Tristan Sylvain, Romain Laroche, Doina Precup, Yoshua Bengio. 2024-10-09. Rejecting Hallucinated State Targets during Planning. https://arxiv.org/abs/2410.07096
Cite the original work for its findings. Save a collection to share your selection of sources.