arXiv · 2609.34575
Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning
Abstract
Offline goal-conditioned reinforcement learning (GCRL) learns goal-directed policies from reward-free data, but in long-horizon tasks, goal-conditioned value functions often provide unstable guidance due to sparse rewards and discounting. Hierarchical methods partially mitigate this issue via subgoal decomposition; however, high-level decision-making still relies on noise-sensitive value estimates, leading to unstable behavior in complex environments. We address this limitation by proposing \textbf{D}iffusion \textbf{S}ubgoal \textbf{P}lanning (\textbf{DSP}), a diffusion-based framework for high-level subgoal generation. DSP casts high-level planning as guided generative inference over goal-conditioned subgoals and learns both conditional and unconditional flows, enabling classifier-free guidance to introduce a goal-directed bias at inference time. By removing explicit value-based guidance from high-level planning, DSP generates reachable and goal-directed subgoals through a generative model while retaining hierarchical execution. Experiments on offline GCRL benchmarks demonstrate that DSP outperforms prior methods on a range of navigation and manipulation tasks, with particularly strong performance in maze environments that require multi-step subgoal planning.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hengrui Zhang, Yuhu Cheng, C. L. Philip Chen, Xuesong Wang. 2026-09-28. Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning. https://arxiv.org/abs/2609.34575
Cite the original work for its findings. Save a collection to share your selection of sources.