arXiv · 2609.35012
Amortized Feedback Planning: Turning Model-Based Rollouts into Executable Policies
Abstract
Closed-loop planning accounts for future observation-dependent actions but can be expensive to repeat at deployment. Bellman-gradient (BG) refinement differentiates conditional rollouts through an existing actor, corrects the current action, and stores the result in an executable policy. A backward sweep reuses deployed future feedback; full-horizon rollouts in an identified Gaussian belief model need no learned value critic. A nonlinear error recursion links quadratic continuation error, local correction, and policy storage. A controlled non-LQG example exhibits second-order policy accuracy with consistent storage, while ideal affine LQG admits exact backward recovery. In nonquadratic thrust, BG reduces actor cost by 2.47-3.20% and remains within 0.13-0.53% of the tested feedback MPC; original policies execute in 5-6 microseconds in a seed-0 native audit. With matched storage, BG attains competitive plant costs at about one ninth (arm) and one sixtieth (docking) of feedback-teacher-plus-student construction time on one GPU, excluding shared learning and node preparation. Richer common maps substantially narrow some student-BG gaps. These results expose the roles of learned feedback, local correction, and storage in executable control.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jeonggyu Huh. 2026-09-28. Amortized Feedback Planning: Turning Model-Based Rollouts into Executable Policies. https://arxiv.org/abs/2609.35012
Cite the original work for its findings. Save a collection to share your selection of sources.